Pioneers Insight Method Research Author
A Big Ruling on LLM Training and Midsummer Mail on NBA Salaries in Tech, Starting from Scratch in 2025, and More
Back to Episodes

A Big Ruling on LLM Training and Midsummer Mail on NBA Salaries in Tech, Starting from Scratch in 2025, and More

Summary

  • The Anthropic ruling gives model developers a major fair-use win for training while allowing claims over allegedly pirated inputs to proceed. Ben Thompson backed Judge William Alsup’s distinction between transformative use and unlawful acquisition: digitizing used books Anthropic purchased for training was fair use, but “just because your usage of it is fair use doesn’t mean your acquisition is now suddenly legal.” Andrew Sharp accepted the legal result while warning that courts, rather than Congress, keep deciding technology’s largest policy questions.

  • Copyright liability should turn on infringing output, not merely on a model having absorbed copyrighted work. Ben argued that no individual book is material to a corpus of roughly seven million books and that an LLM’s output is “unquestionably generating original work”; Andrew’s pushback was that machine-scale ingestion differs meaningfully from human reading. Ben’s answer: “If you think that’s a problem, then make a law about it. Don’t try to retrofit copyright, which is about copying.”

  • Making model vendors responsible for every possible reproduction would structurally favor closed incumbents over open models. Claude’s aggressive filters helped Anthropic because the plaintiffs did not allege that its product reproduced their books, but Ben said those filters should not be a legal prerequisite. Hold a user liable for deliberately copying an article, he argued, while recognizing that a general-purpose LLM has “far, far, far more” lawful uses—the Betamax principle, not Napster.

  • Congress could turn lawful training-data access from an incumbent moat into a U.S. competitive advantage. Ben proposed directing the Library of Congress to digitize its holdings and provide a lawful training corpus to U.S. AI companies; otherwise, startups must scour the world for used books and build scanning operations while foreign developers or rule-breakers proceed more cheaply. Andrew was partly persuaded that “you have to prioritize AI,” but still wanted nationally consistent rules.

  • Elite AI researchers now operate in the closest thing tech has seen to an NBA labor market. Their specialized work is unusually portable across Google, OpenAI, Anthropic, and Meta, the pool of proven talent is small, and outputs are somewhat measurable—so “there’s a bank in every city.” Headline figures such as $6 billion around Jony Ive, $14 billion around Alexandr Wang, and perhaps $1 billion around Nat Friedman and Daniel Gross dwarf athlete pay nominally, yet remain smaller relative to the value of their companies.

  • The clearest near-term LLM return may come from bespoke software for obscure workflows that conventional software companies never discovered. A disability day program replaced manually counted Medicaid TXT-file spacing with a spreadsheet workflow built first through ChatGPT and then Gemini; checking a box now produces the billing output. It still required “a few hours of back and forth” and personal agency, but the endpoint is an ambient assistant that notices repetitive work and builds the automation—“Clippy’s final redemption.”

  • ChatGPT could become a child’s primary computing environment before social media or even a smartphone. One seven-year-old already uses ChatGPT, Roblox, and Minecraft, while her parent expects to withhold social accounts until age 14 or 15 in 2032; Ben saw that as a real platform risk for Meta because parents may prefer AI interactions to social feeds. The unresolved strategic question is whether AI products should avoid social features so they do not inherit social media’s parental distrust.

  • Starting an independent media business in 2025 likely means owning the product, accepting a smaller audience, and charging dramatically more. Ben would avoid dependence on TikTok or YouTube and potentially charge a few hundred or thousand committed readers $1,000 or even $10,000 a month, echoing research analysts who charge $100,000 annually; Andrew suspected he would have tried media briefly, then gone to law school. AI can help build a unique product, but using it as the content generator makes the result “completely non-differentiated.”

Deep dive

1. Copyright is a policy trade-off, not a natural entitlement

  • The immediate ruling: Judge William Alsup found Anthropic’s use of copyrighted books to train Claude lawful, while allowing claims over pirated copies to proceed. Ben called it “great news” and the clearest AI fair-use decision so far.

  • Ben’s starting principle was that intellectual property is a “government-granted monopoly,” not a God-given feature of the universe. Copyright exists to incentivize creation, but that makes it “definitionally a trade-off”—a point even two people who sell copyrighted analysis should remember.

  • Fair use therefore requires balancing purpose and transformation, the nature of the original work, the amount used, and market effects. Ben’s concrete contrast: redistributing a printer manual may readily qualify, while a novel receives stronger protection because society deliberately makes different trade-offs for each.

2. Training on whole books can still be transformational

  • The difficult factor is substantiality: LLMs ingest entire books, whereas Ben normally quotes only a few paragraphs from Bloomberg or The Wall Street Journal. His defense flips the factor—within a corpus of roughly seven million books, any single book is “substantially unimportant” and could disappear without changing the model.

  • On transformation, Alsup repeatedly emphasized that Claude did not reproduce the plaintiffs’ books, and the authors were not alleging that it did. Andrew highlighted the opinion’s word “orthogonal”: the model’s generated output is fundamentally different from receiving the original book.

  • Andrew’s pushback — worth keeping: a human may read many books, but not seven million, and machine-scale ingestion could depress the long-term market for human books, music, images, and video. Ben conceded the competitive disruption but rejected copyright as the remedy: “Copyright law didn’t exist so that you’d have less competition from other authors.”

  • Ben also argued that abundant synthetic content may increase the value of trusted, differentiated brands such as The New York Times. The policy question, in his formulation, is whether the U.S. should handicap AI to protect undifferentiated content rather than police actual copying.

3. Liability should follow the copier, not the general-purpose tool

  • The New York Times examples involved ChatGPT-3, Ben recalled with a hedge, continuing an article after a user pasted its first two paragraphs. He compared that to making 1,000 illegal copies on a photocopier: the user’s action may infringe without making the machine itself unlawful.

  • His governing analogy was Sony’s Betamax victory: time-shifting was fair use, and a machine with substantial lawful uses was not liable merely because it could also be abused. Napster represented the opposite extreme—a service judged to be used for essentially nothing but illegality.

  • Andrew initially dismissed the Times prompts as extreme edge cases “clearly generated for the purpose of bringing a lawsuit.” Ben noted, however, that some Llama 1 examples can apparently regenerate a fair amount of material with relative ease, so output-level infringement remains a genuine issue rather than a wholly manufactured one.

  • Claude’s restrictive filters helped Anthropic establish clean output, but Ben said “filter or no filter” should not determine whether an LLM is legal. Requiring every model to carry an expensive compliance wrapper would impose massive liability risk and “de facto” reserve the market for closed products such as Claude and ChatGPT.

4. Piracy remains illegal—and lawful data access favors incumbents

  • Alsup separated transformation from acquisition. Anthropic could buy used books, cut off their bindings, scan them, discard the physical copies, and retain the digital versions for training; downloading books from archives that had “fallen off a truck on the internet” did not receive the same protection.

  • Ben accepted that piracy should not get a free pass, yet called the economic distinction awkward: used-book purchases usually send no revenue to authors, and scanning produces the same training input as downloading. Andrew supplied the concise description—buying and destroying the books can look like a legal “fig leaf.”

  • That fig leaf still creates a meaningful market structure. If lawful entry requires sourcing millions of physical books and operating industrial scanners, incumbents can comply while startups cannot; companies outside U.S. jurisdiction, including Chinese developers, gain an advantage by ignoring the hurdle.

  • Ben’s congressional proposal was specific: direct the Library of Congress to digitize its holdings and make a lawful corpus available to U.S. AI companies. Andrew said Ben had convinced him “to some degree” that AI must be prioritized, while maintaining that Congress should establish consistent rules rather than leave each district court to improvise.

5. AI researchers finally have an NBA-style labor market

  • Sports and technology share a timing mismatch: seniority means stars often receive their largest contracts after peak performance, while younger contributors are underpaid. Asked whether an engineer’s prime might be the late 20s rather than early 40s, Ben’s exact hedge was “Probably.”

  • AI changes employee bargaining power because the work is both specialized and unusually portable. What a researcher does at Google resembles the work at OpenAI, Anthropic, or Meta; unlike company-specific expertise, that portability lets talent capture more of the value it creates.

  • The market is unusually sports-like because the elite pool is small, identifiable, and judged through somewhat measurable outputs. Andrew’s NBA analogy carried the point: “There’s a bank in every city”—Menlo Park, Mountain View, and San Francisco included.

  • Andrew contrasted $6 billion around Jony Ive, $14 billion around Alexandr Wang, and perhaps $1 billion around Nat Friedman and Daniel Gross with NBA careers worth $500 million to $600 million. Nominal AI numbers are larger, but as a percentage of company value they may be less generous; Ben’s conclusion was that tech employees have historically been underpaid.

6. LLMs turn invisible clerical pain into bespoke software

  • Alex’s wife had been preparing Medicaid billing TXT files by manually counting spaces between rigid fields. ChatGPT wrote a Google Apps Script that translated her Google Sheet into the required format; Gemini later reduced the workflow to checking a box and receiving the correct billing output.

  • Ben focused on the agency requirement: she still invested “a few hours of back and forth with Gemini,” plus patience and executive function, before receiving the downstream savings. The tools are already powerful, but “most people are not even remotely taking advantage of it.”

  • Alex’s larger claim was that LLMs let nontechnical people create custom software for problems no software founder knew existed. His envisioned assistant would watch the billing process for a week, notice the repetition, and proactively say, “I built you a piece of software to automate that.”

  • Ben said such intervention must feel familiar and nonthreatening, then proposed a paperclip character: “Clippy’s final redemption.” The joke carried a serious historical point—many failed dot-com and assistant concepts were sound ideas launched 15 years too early, though he stressed that fully ambient automation still has “a long way to go.”

7. AI may organize childhood—and eventually society

  • A seven-year-old already uses ChatGPT, while her parent expects another seven or eight years before allowing Instagram or whatever social service matters in 2032. ChatGPT, Roblox, and Minecraft may therefore shape her computing expectations before she chooses a phone or social network.

  • Andrew said ChatGPT feels safer than social media or today’s “absolute wasteland” of a general internet. Ben connected that comfort to technology companies’ long campaign to enter schools: “Once you get kids hooked, you have them for life,” spanning “Philip Morris to Google and Apple.”

  • That creates a plausible risk to Meta and other social platforms as parents grow more reluctant to grant early access. Ben left one strategic question unresolved: ChatGPT and its peers may need to be wary of becoming social products, lest they enter the same distrusted category they are currently bypassing.

  • At the institutional extreme, Ben expects AI to continue the internet’s pattern: a few overarching organizational structures alongside greater devolution into city-states, fiefdoms, or other small entities. His analogy was pre-printing-press Europe, with the Catholic Church above fragmented local powers.

8. Starting from scratch means owning the niche and charging more

  • Ben would still insist on owning his site, archive, and customer relationship, even though discovery is much harder in 2025. A creator built entirely on TikTok or YouTube remains trapped by the platform’s incentives and rule changes.

  • His economic adjustment would be “more exclusionary, not less”: assume a smaller audience, make the content dramatically more expensive, and concentrate on the few hundred or few thousand readers willing to pay. He floated $1,000 or $10,000 monthly and noted that independent Wall Street research can cost $100,000 a year.

  • Andrew’s honest counterfactual was less triumphant. The VC-funded ladder from SB Nation and Vox to Grantland gave his generation upward mobility that no longer looks repeatable; at 23 or 24 today, he might try content briefly, become impatient by 24 or 25, and attend law school.

  • Their self-indictment was that they are “the boomers of online content creation”—beneficiaries of unusually favorable entry conditions telling younger creators to try harder. AI does not eliminate differentiation: used as a tool it can realize a unique vision, but used as the generator it produces something “completely non-differentiated.”

9. Content owners must choose what to unbundle

  • A subscriber proposed a “Ben LLM” trained on the Stratechery archive, replacing repeated searches with a conversational interface. Ben has considered it, but the user experience, query costs, and Automattic or WordPress implementation all require work.

  • His hesitation is fundamentally about product definition. Stratechery sells facts and arguments, but also essays in Ben’s voice and a continuing stream of future work; an archive chatbot could separate the informational substrate from the authored bundle.

  • Ben nevertheless said such features would “almost certainly” be common in three or four years and conceded, “Probably should do it. Stay tuned.” Publishers must decide what to expose because “sometimes you do wanna make things more inconvenient for your customers.”

  • Taylor Swift supplied an adjacent lesson in controlling content value. Ben praised her for mobilizing fans around Taylor’s Versions, depressing the value of her masters, and eventually reacquiring them—“She got all her fans to listen to inferior versions of her music so that she could buy it back at a cheaper price”—even while rejecting her original ownership complaints.

10. China benefits too much from Apple to treat it as Huawei

  • Ben rejected the premise that China is merely waiting for Huawei’s operating system and chips to catch up before ejecting Apple. His question was simpler: Apple remains deeply useful to China, so “Why would they want to wreck a good thing?”

  • The benefit is industrial, not merely symbolic. Apple has spent billions training Chinese manufacturers, and each hardware innovation diffuses through the broader electronics ecosystem; Ben said Apple employs roughly a million people in China, if not more, while Andrew believed Foxconn is the country’s largest private employer.

  • Ben also noted Apple’s accommodation of Chinese App Store restrictions. He said people they speak with “strongly believe” Apple handed over encryption keys, while carefully adding that Patrick McGee did not state this explicitly and it has never been officially confirmed.

  • Andrew preserved the downside case: government-office restrictions, consumer nationalism, and a more complicated Apple-China relationship could eventually change the calculus. Ben’s rebuttal was that Huawei’s progress matters only if Apple ceases innovating; betting that its hardware usefulness will disappear is “sort of betting against history.”

11. Addictive video lacks the external harm that drives vice regulation

  • A gambling-industry listener wanted the ability to “self-exclude” from YouTube Shorts. Ben distinguished compulsive video from gambling, drunk driving, or drugs: those generate family destruction, debt, crime, healthcare costs, or harm to bystanders; wasting one’s own time is largely “a you problem.”

  • Andrew raised cigarettes and drugs as counterexamples, but Ben pointed back to secondhand smoke, health-care costs, and related crime as the political justifications. He was reluctant to invite government into behavior that chiefly affects people unable to control their own attention: “Where do you draw the line?”

  • Andrew wanted better user controls. Ben called Shorts pollution in YouTube search “extremely irritating,” while Andrew could not recall a 90-second clip he was glad to have watched; Ben granted that TikTok sometimes produces good specimens, even if X’s versions are “terrible.”

12. Mechanical watches win by doing less, for much longer

  • Ben’s warning was that “mechanical watches are like short-form video for very rich men”—addictive and potentially ruinous. Yet he prefers an analog mechanism that is always readable, requires no charging, supplies no notifications, and does not turn health metrics into another source of stress.

  • Experience narrowed his requirements to a date complication and a GMT function. His Tudor Black Bay GMT keeps a 24-hour reference while letting him change the local time without stopping the watch; he sets the bezel to Taiwan, though daylight-saving rules still make a phone easier for complicated comparisons.

  • Tudor occupies what Ben sees as Rolex’s former position: attainable, durable everyday watches rather than luxury collectibles. His model cost about $3,000 when purchased and roughly $4,500 now; compared with replacing a $1,500 iPhone every year or two, a watch worn daily for years can be “trivially worth it.”

  • His practical advice was to discover what actually matters—GMT over chronograph, perhaps titanium for lightness, rubber for comfort, and a case proportionate to the wrist—before chasing movements. Andrew’s Garmin compromise follows the same logic: step counting and seven-day battery life, with the stressful training metrics ignored.