‘A.I.-Washing’ Layoffs? + Why L.L.M.s Can’t Write Well + Tokenmaxxing
Summary
The layoff wave is a real labor signal, but the episode does not treat every AI explanation as causal. Atlassian cut 10%, or about 1,600 jobs; Block cut roughly 40%, or 4,000; and Reuters reported that Meta could eliminate 20% or more, as many as 16,000, though Meta called that “speculative reporting” and no cuts had been announced at recording. Casey Newton’s synthesis: companies keep saying AI matters, and “sooner or later, I do think we’re going to have to believe them,” even where overhiring, weak stocks, or dysfunction offer competing explanations.
For investors, AI is functioning as both an operating model and a valuation story. Block rose 17% the day after announcing its layoffs, while Meta is pairing possible cuts with $135 billion in planned capital expenditure; Kevin Roose’s framing is that companies are not necessarily reducing total costs, but “shifting the cost from human labor to AI.” The wager is that payroll can become compute spend—and that markets will reward management for telling that story before the productivity gains are proven.
Meta’s proposed labor-for-compute swap remains especially speculative because its own AI execution has been uneven. Zuckerberg says projects once requiring big teams can now be done by “a single, very talented person,” yet the hosts note that Meta abandoned Behemoth, reportedly delayed Avocado after missed targets, and reorganized its AI teams again. The frontier labs—OpenAI, Anthropic, and Google—are not themselves conducting comparable mass layoffs, though Kevin notes that OpenAI and Anthropic are much smaller and that some lagging companies may be using AI to catch up.
AI adoption has become a no-win signaling problem for workers and a potential tool of managerial discipline. Employees fear that heavy usage proves they are adaptable, but also that their jobs can be automated; Casey cautiously observes that repeated layoffs have made Meta workers quieter even if workforce control is not their stated purpose. Kevin sees a possible opening for tech unionization, while Casey offers the sharper organizing test: “I cannot think of anything that would make Mark Zuckerberg more mad than a union of software engineers at Meta.”
Modern chatbots are more useful for many tasks than GPT-2 or GPT-3, yet post-training has traded away surprise, voice, and stylistic range. Jasmine Sun argues that early models were “nutty” and unreliable but lacked today’s em dashes, tripartite lists, and “it’s not this but that” cadence; GPT-3 could also imitate writers more convincingly than ChatGPT 5.4 Thinking in her tests. RLHF, scripted dialogues, word restrictions, and human preference ratings transformed the “nut job concussed models” into safe corporate assistants.
The deeper creative-writing bottleneck is that artistic quality is neither cleanly verifiable nor grounded in a model’s own life. A writing evaluator described rubrics that penalized three exclamation marks or graded fan fiction for factuality, illustrating how methods suited to code—which runs or does not—fail on subjective art. LLMs can produce polished metaphors, but Jasmine argues that human writers’ work has stakes when it comes from experience, observation, or community; Casey’s pushback is that models can still discuss music evocatively despite never hearing it.
Token usage is becoming a costly status metric whose incentives may outrun its economic value. OpenAI’s seven-day leader reportedly consumed 210 billion tokens—about 33 Wikipedias—while Kevin heard that Anthropic’s top individual Claude Code user spent more than $150,000 in one month; some firms now incorporate consumption into performance reviews. The hosts invoke Goodhart’s law and the old warning that measuring programming by lines of code is “like measuring aircraft building progress by weight”: some tokenmaxxing creates real output, but leaderboards invite waste, side projects, budget blowouts, and artificial employee lock-in.
Deep dive
1. AI-linked layoffs are an early warning, not a clean causal experiment
The immediate numbers are substantial: Atlassian eliminated 10% of staff, about 1,600 jobs, while Block cut roughly 40%, or 4,000. Reuters reported that Meta was preparing to cut 20% or more—potentially 16,000 positions—but Meta called the report “speculative,” and the hosts stressed that nothing had happened as of recording.
Atlassian said the reductions would fund AI and enterprise sales; Block described a shift toward smaller, flatter teams; and Meta has publicly embraced a new AI-intensive way of working. Casey’s high-level judgment: the circumstances differ, but executives repeatedly identify AI as significant, and “sooner or later, I do think we’re going to have to believe them.”
Kevin sees tech workers as likely early casualties because their employers both build and rapidly adopt these tools. Yet the evidence does not isolate substitution: falling share prices, pandemic-era hiring, strategic resets, and organizational dysfunction all overlap with AI adoption.
Casey’s worker-centered pushback cuts through the attribution debate: “Does it actually matter if the effect on workers is the same?” Whether the proximate cause is automation, overstaffing, or investor pressure, thousands of people still lose their jobs.
2. Atlassian gets an AI pass; Block looks more like management cleanup
Atlassian CEO Mike Cannon-Brookes said AI was not replacing people, but that it would be “disingenuous” to deny changes in the skills mix or number of roles required. Casey considered that relatively candid and declined to call it AI-washing without clearer information about which functions were cut.
The harder problem for Atlassian is the hosts’ “SaaSpocalypse”: its business tools encode structured workflows that customers may eventually reproduce cheaply. With the stock battered, layoffs offer a new market narrative—fewer employees, higher productivity—even if customers continue buying Atlassian products at lower prices.
Block’s history makes its AI explanation less persuasive. Headcount had tripled from about 3,800 in 2019, and five months before the cuts the company spent $68 million flying 8,000 people to an event with Jay-Z; Casey’s verdict was that AI might justify the cleanup “if you squint,” but chronic mismanagement explains plenty.
The market nevertheless rewarded the message: Block shares jumped 17% the following day. Kevin compared AI’s narrative power to the crypto boom, when merely adopting fashionable language could lift a stock; Casey’s blunt conclusion was that “the public markets actually can just be tricked that easily.”
3. Meta is financing an unproven labor-for-compute substitution
Meta’s possible cuts sit beside $135 billion in planned capital expenditure this year. Casey reads the combination as reassurance to investors: management can pursue “the biggest bet in the company’s history” while signaling that it has not “completely” lost control of expenses.
Zuckerberg supplied the productivity premise: “Projects that used to require big teams now can be accomplished by a single, very talented person.” Kevin’s sharper accounting interpretation is that these firms are not necessarily saving money in aggregate—they are moving it from salaries into data centers, models, tools, and tokens.
A venture capitalist told Kevin that some highly AI-native startups already spend more on AI tools than on payroll. Kevin called that potentially an outlier, but also a picture of the destination executives imagine: most operating expense eventually buys machine labor rather than human labor.
Casey’s caveat is that Meta has not earned the productivity claim company-wide. It abandoned Behemoth, reportedly delayed Avocado because it barely outperformed Gemini 2.5, and reorganized its AI teams again; meanwhile, OpenAI, Anthropic, and Google—the frontier builders—are not laying off workers en masse. Kevin also noted that OpenAI and Anthropic are much smaller, and that some companies making cuts may be trying to catch up with competitors through AI.
4. Adoption anxiety may discipline workers before AI replaces them
One tech employee described an impossible choice: use AI aggressively to show alignment with management, or avoid demonstrating that the job can be automated. Kevin heard “jostling and fear and anxiety,” intensified by the knowledge that executives are actively planning reductions.
Casey would not claim that recurring layoffs are deliberately meant to keep Meta’s workforce in line, but he observed that they have had that effect. After employees feared they might genuinely lose their jobs, internal protests diminished and a workforce once willing to challenge management became “a lot more quiet.”
Kevin wonders whether this pressure could finally trigger mass tech unionization. Unlike nonunionized software workers, manufacturing unions historically negotiated redeployment and retraining when jobs were automated; Casey’s provocation was that nothing would anger Zuckerberg more than “a union of software engineers at Meta.”
5. Post-training made chatbots useful by sanding away their voices
Jasmine’s claim is narrower than “humans write better”: most writing is bad, and LLMs outperform most people at ordinary language tasks. Her puzzle is why leaders promise superhuman coding and scientific discovery while Sam Altman cautiously imagines only “a real poet’s okay poem.”
Looking back through model generations, Jasmine preferred the prose of GPT-2 and especially GPT-3. Those systems may have lied and wandered—the hosts likened GPT-2 to someone who had fallen down stairs—but their tone varied, they surprised the reader, and GPT-3 could emulate figures such as Paul Graham or Richard Dawkins.
Reusing an old GPT-3 style prompt with ChatGPT 5.4 Thinking produced something Jasmine called “God awful.” Modern outputs instead share recognizable tics: em dashes, three-part lists, “it’s not this but that,” and a chirpy competence optimized for corporate assistance.
Her mechanism is post-training: labs give base models scripted dialogues, behavioral restrictions, approved vocabulary, and RLHF ratings from human graders. Those layers tame what she called “crazy unpredictable” and “nut job concussed models,” but also trap them inside a generalized helpful-assistant persona.
6. Creative quality breaks the machinery of verifiable rewards
Labs recognize that AI researchers may understand good code better than good prose, so firms such as Mercor and xAI advertise “creative writing expert” roles around $45 an hour, sometimes requesting a New York Times bestseller or starred Kirkus review. Yet expertise does not rescue a nonsensical evaluation system.
A Scale AI contractor working as a writing evaluator recalled penalizing responses for having three exclamation marks. The evaluator was also asked to grade fan fiction for factuality—Jasmine’s specimen of well-resourced companies trying to turn aesthetic judgment into a mechanical checklist.
Kevin connected the failure to verifiable rewards: generated code can be tested because it runs or it does not, while no evaluator can consistently prove why Shakespeare is Shakespeare or one Neruda poem succeeds. User demand reinforces the outcome because the dominant request is not literature but “Write this email for me,” where blandness works.
Jasmine’s second explanation is grounding. The model can produce striking phrases such as “the liminal day that tastes of almost-Friday,” but it has no life supplying stakes, observation, or point of view; Casey countered that models still discuss the sensory qualities of music surprisingly well, perhaps by pattern-matching writing from people who have listened.
7. Text generation is automatable; the rest of writing remains stubborn
Blind tests may show readers preferring AI prose until its source is revealed, and Jasmine accepts that people resist obvious machine writing. Her quibble is task definition: she estimates text generation occupies only 25% of her working hours; interviewing, finding ideas, selecting sources, reporting, and deciding what deserves to be written constitute the rest.
Kevin raised the “cope” objection—that writers are repeating engineers’ mistake of defining their value around whatever models cannot yet do. Jasmine’s answer is empirical and hedged: she has tried for three years to automate herself with Claude and failed, though style could improve, fine-tuning could help, and she does not say “never.”
Genre fiction shows both progress and constraint. Sudowrite co-founder James Yu and other practitioners described immense engineering work to undo post-training’s chirpy, sycophantic, PG-13 tendencies; Jasmine calls the resulting author-model workflow a “centaur model,” because humans must prompt and bully the system toward weirdness and sensuality.
Her conditional forecast: given interview transcripts, models might eventually write strong features or literary fiction if labs invested as heavily as they do in coding agents. She doubts that will be financially attractive compared with “automating 23-year-old software engineers,” while remaining less bullish on models independently reporting from the world.
8. A personalized Claude editor works because it learns one writer’s taste
Jasmine’s productive workflow does not ask Claude to write for her. In a Claude project, she loaded her Substack archive, freelance work, post-publication retrospectives, audience, beat, and goals, then co-developed criteria based on her own aspirations rather than a generic standard of “good writing.”
The resulting rubric identifies traits such as her “insider anthropologist position in Silicon Valley,” movement between startup jargon and internet slang, and shifts from policy analysis to personal scenes. She divided review into ideation, structure, prose, and final fact-checking phases.
Instead of inventing material, Claude might flag a summary conclusion as boring, recall that an earlier essay ended more powerfully on a scene, and ask what Jasmine felt as a plane took off or whether a conversation could animate dry policy. She retains judgment, but the tool pushes her toward “the best version of myself as a writer.”
9. Tokenmaxxing confuses AI adoption with measurable productivity
A token is a fragment of a word and the unit by which model providers meter consumption; roughly 10,000 tokens can generate 7,500 words. Agentic coding now burns hundreds of thousands or millions in one session because developers run longer, concurrent processes rather than exchanging a single prompt and response.
OpenAI’s top seven-day employee reportedly consumed 210 billion tokens, about 33 Wikipedias, though some were cached. Kevin heard that Anthropic’s top individual Claude Code user spent more than $150,000 in one month; a Swedish engineer told Kevin that he spends more than his salary on Claude.
Companies use leaderboards to motivate experimentation and monitor whether engineers have embraced agentic programming; some now include token consumption in performance reviews. Employees at the labs may receive access for free, while workers elsewhere can outstrip their employers’ budgets. But Goodhart’s law applies immediately: make usage a target and workers can inflate it with worthless projects—or, as one person speculated, use the company’s compute for side projects.
Casey’s historical analogy lands cleanly: measuring programming through token volume resembles measuring aircraft-building progress by weight. Kevin still resists calling all tokenmaxxing theater—some power users may genuinely be more productive—but managers should ask what the spend produced, especially as AI-use scoring reaches at least some marketing reviews and engineers ask prospective employers, “What’s my token budget?”