Pioneers Insight Method Research Author
It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast
Back to Episodes

It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast

Summary

  • Cotra’s modal expectation is for top-human-expert-level AI in the early 2030s, after which software automation could spill rapidly into robotics, chipmaking, and the full physical production loop. If AI and robots can perform everything required to make more AI and robots, today’s roughly 2% growth norm ceases to be a persuasive ceiling. Her outside possibility is that 2050 differs from today as radically as today differs from hunter-gatherer society: “10,000 years of progress rather than 25 years of progress.”

  • The investable disagreement is not whether AI matters, but whether it adds 0.3 percentage points to growth or pushes peak growth toward 1,000% annually. Slow-growth thinkers extrapolate 150 years of stubbornly stable frontier growth and assume hidden bottlenecks; fast-growth thinkers extrapolate the longer historical acceleration created by larger populations, more ideas, and positive feedback. The resulting gap is “100 or 1,000 or a 10,000-fold disagreement,” large enough to reverse views about work, policy, and risk.

  • Benchmark headlines are weak warning signals because every benchmark follows an S-curve, saturates, and is replaced before it establishes real-world danger. Cotra instead wants fixed-cadence disclosure of labs’ strongest internal results, actual productivity gains, internal AI usage, and the share of pull requests mostly written and reviewed by AI. The decisive indicator is whether AI has begun accelerating the entire AI stack—not merely code, but chip design, fabrication, equipment, maintenance, and raw-material supply.

  • “Crunch time” is the narrow interval when AI may be powerful enough to transform safety work but not yet powerful enough to escape human control. Once AI R&D is substantially automated, progress previously expected over 10, 20, or 30 years might arrive within six months to two years. Cotra’s prescription is to redirect as much AI labor as possible away from recursive capability improvement and toward alignment, cyber defense, biodefense, monitoring, negotiation, and collective decision-making.

  • Major frontier labs’ convergent safety plan is to use each generation of AI to understand and secure its successors, but the binding risk may be institutional commitment rather than technical impossibility. A lab with 100,000 smart-human equivalents could still assign only 100 to safety while competition consumes the rest. Cotra is “reasonably bullish” that control techniques can extract useful work from early, non-“galaxy brain” systems, yet warns that insufficient checking hands power to the models while exhaustive checking destroys the speed advantage.

  • Compute and model access become strategic assets if external safety organizations must mobilize during crunch time. The leading lab might withhold its best internal system, or price inference near the opportunity cost of using that compute for further self-improvement. Cotra therefore entertains owning GPUs, securing model-access agreements, or hedging compute inflation through NVIDIA and other AI-exposed public equities—though a super-exponential feedback loop could still turn today’s competitive market into winner-take-all concentration.

  • Cotra’s organizational lesson is that AI adoption must begin before the emergency, because neither governments nor philanthropies can improvise a new operating model in six months. Open Philanthropy might eventually spend more on API credits and GPU time than on human salaries, but its multilayer grant process is poorly shaped for billion-dollar, time-critical deployments. Her broader warning is vivid: without aggressive adoption, industry’s “fast cars” will be overseen by regulators using “horses and buggies.”

  • The episode’s post-publication framing says the timetable may already be shortening. The cross-post notes that on March 5, after making forecasts in January 2026, Cotra wrote “I Underestimated AI Capabilities Again” because several expectations were already beginning to be met in the first couple of months of the year; it also cites Anthropic’s Mythos model and reported zero-day discoveries across major operating systems and browsers. Its conclusion is deliberately conditional but urgent: “crunch time is arguably here now.”

Deep dive

1. The cross-post says crunch time may already have begun

  • The Cognitive Revolution introduces Cotra as an unusually strong forecaster: third among more than 400 participants in the 2025 AI Digest forecasting survey. Its framing is that even accelerationists may experience “future shock” if compounding automation encounters no insurmountable bottleneck.

  • The postscript says Cotra’s March 5 article, “I Underestimated AI Capabilities Again,” reported that predictions made in January 2026 were already beginning to be met in the first couple of months of the year.

  • The introduction also points to Anthropic’s Mythos, citing major benchmark gains and reported zero-day discoveries in every major operating system and browser. Its bottom line is not that superintelligence has arrived, but that “crunch time is arguably here now.”

2. AGI has been watered down until it predicts almost nothing

  • Cotra’s DealBook example exposes the semantic drift: seven or eight of roughly ten panelists expected, by 2030, AI able to do everything humans can do, yet eight expected AI to create more jobs than it destroyed over the following decade.

  • Cotra herself did not endorse the by-2030 forecast because her timelines were somewhat longer. Her confusion was conceptual: “Why is it that you think we will have AI that can do absolutely everything that the best human experts can do in five years,” yet employment remains conventionally expansionary?

  • When pressed, panelists retreated to milder definitions—“we kind of already have AGI,” perhaps GPT-5—then treated the absence of immediate transformation as evidence that AGI was never momentous. Cotra sees that as importing evidence from one definition into a radically stronger one.

3. Cotra’s 2050 is closer to prehistory than to 2000

  • The mainstream picture is manageable continuity: a somewhat larger population, better medicine, modestly longer lives, and technological change comparable to 2000–2025. Even people expecting AGI in 2030 often retain that baseline.

  • At the opposite pole, the worldview represented by If Anyone Builds It, Everyone Dies has a sudden jump from systems like GPT-5 or GPT-6 to intelligence against which humans resemble “cats or like mice or ants,” followed by immediate physical leverage through nanotechnology or near-light-speed probes.

  • Cotra sees a broad spectrum between those extremes: intermediate stages still lead toward technologies near physical limits, including useful self-replicating systems. Her striking comparison is that 2050 might embody “10,000 years of progress rather than 25 years of progress.”

4. Remote-work superintelligence could close the physical loop

  • Cotra’s modal forecast is top-human-expert-level AI in the early 2030s: systems better than the best virologist, software engineer, or other specialist at tasks performed remotely through a computer. Narrower systems may already have changed the world substantially before then.

  • From that point, superhuman cognitive systems could direct human labor to build robotic actuators, then automate more physical work. Cotra’s uncertainty is wide, but she considers perhaps one or a few years enough to make robotics materially more autonomous because large models, scale, data, and imitation are already advancing it.

  • Tom Davidson’s “Three Types of Intelligence Explosion” supplies the missing frame: better AI software is only one loop. Full self-reproduction requires chip design, fabs, the machines that build fab equipment, repairs, raw materials, and every upstream dependency needed to produce more chips and models.

5. Growth forecasts differ by three or four orders of magnitude

  • The conversation puts the live range starkly: some serious analysts expect AGI to add only 0.3 percentage points to economic growth, perhaps roughly a 15% increase over current rates; others envisage peak growth of 1,000% annually or even thousands of percent.

  • This is not a disagreement among isolated camps that never exchanged arguments. People have discussed the mechanisms repeatedly and still differ by factors of 100, 1,000, or 10,000 in the technology’s likely economic impact.

  • The policy consequences reverse with the forecast. An accelerationist who moved from 0.3% to 1,000% would likely become far more concerned; an x-risk researcher moving the other way might regard the new view as decisive evidence against much of their current work.

6. Competing outside views prevent convergence

  • Slow-growth thinkers begin with 100–150 years of roughly 2% frontier growth. Electrification, washing machines, radio, television, computers, and the internet transformed life without producing an obvious permanent break in measured growth, so AI may merely sustain the same trend.

  • Their supporting intuition is Hofstadter’s law: “It always takes longer than you think even when you take Hofstadter’s law into account.” Cotra’s favorite variant is the programmer’s credo: “We do these things not because they are easy, but because we thought they would be easy.”

  • Fast-growth thinkers zoom out 10,000 years and see acceleration, from perhaps 0.1% growth around 3000 BC to post-Industrial-Revolution rates near 2%. More people produced more ideas, raising food output, supporting more people, and repeating the feedback loop.

  • Each side has an error theory for the other. One hears every explosive scenario as another story omitting drag; the other hears “there will be bottlenecks” as an ungrounded blanket assumption whose specific examples might reduce 1,000% growth without establishing ceilings of 2% or even 10%.

7. Bottlenecks need observable tests, not another round of stories

  • Cotra wants near-term observations that can adjudicate between those priors, beginning with actual AI uplift in software and AI R&D. METR’s randomized controlled trial split developers into AI-allowed and AI-disallowed groups and, in that setting, found AI slowed completion of real work.

  • She does not expect that result to remain true, but values having the measurement before speedups become overwhelming. Benchmarks should be cross-checked against internal rollouts, company RCTs, self-reports, and observable output rather than treated as self-validating capability measures.

  • Her formula is to find domains with the deepest adoption and measure actual production: not just whether a model answers questions about solar panels, but whether an AI-heavy factory manufactures them faster or improves their performance more rapidly.

8. Agent benchmarks improved, but the real world remains the test

  • Cotra’s late-2023 requests for proposals distinguished difficult, realistic agent benchmarks from other evidence. She wanted models to book flights, repair software, write and run tests, and iterate until success—not merely answer multiple-choice questions.

  • Researchers responded strongly to the benchmark arm, producing projects including Cyber, a cyber-offense benchmark used in standard evaluations. Interest was much weaker in surveys and RCTs because benchmarks are “clean and contained,” while reality is “messy and open-ended.”

  • Cotra nevertheless stresses that benchmarks routinely overestimate real-world performance. The companion evidence program was meant to capture deployment, adoption, and outcomes that automated scoring cannot reproduce.

  • The Forecasting Research Institute’s LEAP—Longitudinal Experts on AI Panel—now asks roughly 100–200 AI experts, economists, and superforecasters about six-month, one-year, and five-year indicators, then connects granular predictions to longer-run worldviews and checks who was right.

9. A private intelligence explosion could outrun public awareness

  • AI 2027 provides the scenario: a leading company, OpenMind in the story, becomes so far ahead that it retains its best models internally and releases only systems slightly better than competitors’ public frontier.

  • Such a company need not monetize its best product if internal AI R&D is worth more than external revenue. The public could therefore see only incremental product improvement while the lab experiences explosive productivity behind closed doors.

  • Cotra’s concern is the lost warning interval. Society might have known six or twelve months earlier which growth regime had begun, yet governments and outside experts would remain unable to act because the decisive evidence stayed internal.

10. Transparency should measure delegation, not just benchmark scores

  • Cotra would like labs to report their strongest internal benchmark results on a calendar cadence—perhaps every three months—rather than only when a product launches. Public model cards for Claude Opus 4 or GPT-5 are helpful but miss purely internal deployment.

  • She is more interested in how systems are actually used. CEOs sometimes boast that AI writes 90% of code, but line count is weak evidence; a sharper measure is the fraction of internal pull requests mostly written and mostly reviewed by AI.

  • That measure captures both capability and organizational deference. If humans continue handling management, approval, and review, progress can accelerate only so far; for “crazy fast” growth, AI eventually has to perform those higher-level functions too.

  • Cotra also wants internal uplift studies, subjective speedup estimates, algorithmic progress, and serious misalignment incidents—for example, whether a deployed model lied about something important and concealed the logs. Those disclosures are also the most competitively and reputationally sensitive.

11. Public evidence matters because an alarm must become common knowledge

  • Benchmarks alone cannot trigger the alarm because they repeatedly trace an S-curve, saturate, and give way to harder successors. Even 100% on today’s evaluations would not convince Cotra that a model could take over the world.

  • The late but clear signal is observed productivity: internal discovery occurring far faster than before. That is the point at which “they should definitely sound the alarm,” because theoretical progress is visibly feeding back into more progress.

  • Rob suggests confidential reporting to technically capable government agencies, which routinely receive commercially sensitive information. Cotra accepts that as better than silence but doubts that 10 or 50 understaffed officials can interpret a shifting evidentiary frontier fast enough.

  • Her preferred alarm resembles COVID or the public reassessment after Joe Biden’s disastrous debate: a society-wide conversation. A prominent skeptic such as Arvind Narayanan publicly changing his mind would create common knowledge that rumors at San Francisco parties cannot.

12. Transparency legislation should target the highest-value signals

  • Cotra does not rank an all-or-nothing disclosure package as the single top policy fight. She favors identifying what evidence would reveal an intelligence explosion, then securing the highest-value, biggest-bang-for-buck elements first.

  • Public disclosure creates difficult trade-offs: unusually rapid algorithmic progress attracts competitors, unusually slow progress discourages investors, and misalignment incidents are embarrassing. Industry aggregation might soften those costs, but with only a few frontier labs observers could often infer the source.

  • Existing efforts such as New York’s RAISE Act and California’s SB 53 fit the broader approach because they emphasize transparency and whistleblower protection. Cotra treats whistleblowers as an important policy plank supporting transparency.

  • Leaks remain insufficient. Social proximity between Bay Area safety researchers and lab employees can reveal what is coming, but “rumors in San Francisco tech-bro parties” cannot justify costly action in Washington, London, or Brussels.

13. The alarm should redirect AI labor before capabilities compound

  • If AI substantially automates frontier AI R&D, progress previously expected over 10, 20, or 30 years might arrive within a year or two, perhaps six months. Systems may be manageable at the alarm point yet rapidly approach “godlike abilities.”

  • Cotra’s response is to redirect as much AI labor as possible from making stronger models toward work that protects society from successor systems. The target includes takeover risk, misuse, infrastructure vulnerability, geopolitical conflict, and other disruptions created by increasingly powerful AI.

  • A leading company cannot easily redirect unilaterally because rivals may catch it. The warning therefore creates a coordination window: if society can credibly see six, 12, or 18 months to radical superintelligence, competitors may coordinate to redirect resources toward protective work.

14. AI-powered safety is not circular, but trust is the bottleneck

  • Critics describe the strategy as “flying by the seat of your pants”—using the technology creating the problem to solve it. Cotra replies that civilization repeatedly used general-purpose technologies to contain their own externalities.

  • Cars enabled drive-by shootings and carjackings, but also equipped law enforcement; computers enabled hacking, but also automated monitoring and vulnerability discovery. Forecasts often imagine a new technology’s harms in detail while failing to imagine equally technology-enabled defenses.

  • Nathan initially suggests misalignment is uniquely awkward because an AI asked to solve alignment might sabotage its own constraint. Cotra’s pushback is broader: a power-seeking model would also undermine biodefense, epistemic improvement, and civilizational defenses that might expose or block it.

  • The central technical problem is therefore building enough confidence through control, alignment, interpretability, and related techniques to rely on AI outputs. Excessive checking bottlenecks progress; inadequate checking means “we hand the AIs the power to take over.”

15. The defensive portfolio runs from alignment to civilization-wide resilience

  • Alignment is foundational: each system and its successors must remain motivated to help humans, honest, steerable, and responsive to instruction. A failure anywhere in that chain compromises every later protective application.

  • Cyber defense is a natural dual-use race. The same models capable of finding vulnerabilities in power grids, weapons, and critical systems could identify and patch them before malicious actors obtain comparable capability.

  • Biodefense could combine rapid detection of novel pathogens, faster medical countermeasures, and large-scale production of PPE, clean rooms, or related infrastructure. If robotics has matured too, AI could help execute manufacturing rather than merely recommend it.

  • More speculative work targets collective cognition: truth-seeking, negotiation, compromise, policy design, avoiding US–China war, space governance, value lock-in, and healthier political discourse. Cotra groups much of this as AI for “coordination, compromise, negotiation, truth-seeking.”

16. Major frontier labs are betting that one generation can secure the next

  • Cotra sees the same architecture in public plans from OpenAI, Anthropic, and Google DeepMind: as models improve, the companies expect to incorporate AI itself more heavily into alignment, control, interpretability, and safety work.

  • The plan requires a useful interval before systems become uncontrollably powerful—ideally observable in advance and lasting at least six months or a year even without an extraordinary slowdown.

  • If a generality threshold instead produces extreme superintelligence within days or weeks, the strategy fails. Society may not notice the transition, mobilize resources, validate the models’ work, or coordinate before control is already lost.

17. Capability ordering can make the window useless

  • A particularly bad ordering produces AI that is extraordinary at AI R&D but weak at everything protective, even closely related safety research. It might improve successor models’ sample efficiency for months without being useful for alignment, biodefense, or governance.

  • Nathan notes the apparent contradiction: such a narrow system is not itself broadly dangerous. Cotra’s resolution is an AlphaFold-like savant or blind search that discovers an architecture or training method whose successor can suddenly “go foom.”

  • Current skill profiles already look jagged. Cotra’s ML-researcher friends receive larger gains than she does in weird open-ended thinking, podcasts, and grant-related emails; control and alignment benefit because much of that work resembles software engineering and ML research.

  • Moral philosophy, negotiation, policy, and institution-building may lag badly. The more distant a task is from code and ML research, the larger Cotra expects the capability penalty to be.

18. Tight feedback loops favor code over institutions

  • AI excels where success is rapidly legible: code either runs or fails. That “very hard to fake signal” supports training, evaluation, and repeated improvement in a way that year-long management or political projects do not.

  • A model may produce a persuasive white paper and receive a thumbs-up without being able to lead thousands of humans and robots through execution. Nathan’s caricature captures the gap: “Here’s the widget…go and make 10 billion of them.” The model replies, “Good luck.”

  • Defensive planning should therefore assume that ideas arrive before execution capacity. Long-lead infrastructure—PPE, vaccines, pathogen detection, clean rooms, or manufacturing capacity—may need to be built before systems become capable of designing threats.

  • Cotra treats AI defensive labor as a forecast to plan around, “not a guarantee.” Humans should specialize now in tasks where future models may remain comparatively disadvantaged, especially physical deployment and social coordination.

19. The intelligence explosion changes both urgency and capability

  • Safety work should use AI as soon as it becomes useful; crunch time is not permission to wait. Its special significance is the clock: the default trajectory might leave only 12 months to uncontrollable superintelligence.

  • Crunch time also supplies a new lever. Because AI is, by definition, highly capable at AI R&D then, some of that work can be redirected toward filling its own skill gaps—fine-tuning or scaffolding systems for biodefense, philosophy, negotiation, or other priorities.

  • Cotra extends the logic to philanthropy: today more than 80% of grant money goes to salaries for humans; within a few years, the better allocation could be API credits or rented GPU time supporting a similar distribution of research and policy work.

20. A pause should stretch the transition, not freeze and jump

  • Cotra regards her proposal as compatible with pausing at the brink of an intelligence explosion. If the default is 12 months, she would vote to make the transition “10 times longer or even longer”—10, 15, or 20 years rather than one.

  • She reframes a pause as a continuum. Instead of choosing between 100% of AI labor going to better models and 0%, society can repeatedly slow, redirect, test, and selectively advance capabilities that improve defenses without making systems uncontrollable.

  • Her quibble is with “pause, hang out for 10 years, then unpause,” which preserves the later jump. She prefers slowly inching through a sweet spot where models are powerful enough to help but not powerful enough that “we’ve already lost the game.”

21. The likeliest lab failure is underinvestment, not impossible alignment

  • Cotra’s leading failure mode is that labs never execute the promised redirection. Competitive pressure could leave 100 of 100,000 smart-human equivalents working on safety—more labor than before, but tiny relative to the pace of capability improvement.

  • Rob proposes legal requirements or inter-company commitments to devote perhaps 50% of compute to safety. Cotra sees the enforcement problem: what counts as safety, who audits every team, and how can regulators distinguish protective research from capability work without deep technical competence?

  • She considers severe early misalignment plausible but not the most likely failure. If controls cannot obtain honest work, models might selectively advance capabilities while sabotaging alignment, biodefense, and other safeguards; even a slowdown would then leave humans “on our own.”

  • A discontinuous corporate pivot is intrinsically difficult. Cotra wants labs to increase human labor, inference compute, and fine-tuning devoted to safety gradually, so the final transition follows an established schedule instead of requiring an abrupt institutional reinvention.

22. Philanthropy may have to replace salary budgets with inference budgets

  • Open Philanthropy’s immediate task is to monitor how useful AI is across its own work and that of grantees such as Forethought, Redwood Research, Apollo, and policy organizations. Once a funded activity shows “signs of life” from automation, it may warrant rapid scaling.

  • Cotra suggests explicitly separating conventional grants from spending that buys AI labor. Tracking ChatGPT Pro subscriptions, API credits, and inference compute would reveal whether AI’s share of giving is rising in line with capabilities and expected crunch-time timing.

  • Existing approval chains are poorly shaped for a billion-dollar decision: a junior staffer develops a case, then information climbs through two, three, or four managerial layers. That process cannot simply scale to a time-critical commitment that no junior employee could reasonably authorize.

  • In the easy scenario, Open Philanthropy itself is largely automated and has perhaps 1,000 AI-team workers rather than 45. In the harder jagged scenario, AI masters a few fundable verticals while the institution remains human-bottlenecked and lacks a visceral sense that the moment has arrived.

23. Model access and compute ownership become strategic bottlenecks

  • An external safety organization may possess billions of dollars yet be unable to buy the leading model. A frontier lab could retain its best systems internally and sell only less-capable systems that remain marginally ahead of competitors’ public products.

  • Even willing sellers may charge prohibitive prices because the opportunity cost of inference is training or operating a stronger successor. During crunch time, compute prices could therefore reflect recursive-improvement value rather than ordinary API economics.

  • Cotra’s hedge ranges from owning GPUs—renting them commercially in normal times, then redirecting them—to holding NVIDIA or other liquid AI-exposed stocks. Rising compute costs would then increase portfolio value, though the organization would still need rights to run frontier software on its hardware.

24. Competition initially opens access, then may create a winner-take-all race

  • Cotra expects early crunch time to resemble today: several leading companies within perhaps a month of one another, with different capability spikes, limited moats, and strong incentives to sell API access.

  • A super-exponential loop changes the market structure. Under ordinary exponential growth, competitors growing at the same rate preserve relative shares; under accelerating growth, the actor reaching each milestone first can compound its lead into disproportionate wealth and power.

  • A distant leader has reason to conceal its system, forgo quarterly revenue, and recurse toward intelligence capable of rivaling nation-states or “decisively” shaping the future. Close competitors and impatient investors instead force commercialization and narrow the room for secrecy.

  • Safety also competes with two more attractive uses: private power and ordinary goods, services, and media content that consumers will pay for. As with today’s limited spending on biodefense, cyber defense, or moral philosophy, market demand does not automatically allocate transformative capacity toward public protection.

25. Sequential bottlenecks make preparation more valuable than raw compute

  • Physical defenses contain irreducible calendar time: experiments must run, factories must be built, and countermeasures must be manufactured. Doubling compute does not necessarily halve any of those schedules.

  • Social systems are similarly sequential. AI might identify a mutually beneficial and enforceable US–China agreement, but officials still have to meet, deliberate, come to a decision, and ratify terms.

  • Even theoretical research is not perfectly parallel. A hundred models can search for an insight, but if the solution requires three or four dependent conceptual leaps, the field must discover each foundation before work on the next can begin.

  • The best pre-crunch contributions are therefore long-lead assets: physical biosecurity infrastructure and social consensus around possibilities such as slowing jointly, redirecting compute, or negotiating treaties. Ideas need years to enter policymakers’ “toolkit” before an emergency.

26. Governments risk regulating fast cars with horses and buggies

  • Cotra especially wants government entities to adopt AI now. Random types of red tape could widen the gap until regulated companies have “fast cars” while regulators retain “horses and buggies.”

  • Defensive organizations should not wait for a universal model. Each needs a team repeatedly asking where AI has become genuinely useful, how workflows should change, and which outputs can be verified reliably.

  • Her advice is aggressive but conditional: maximize adoption for the organization’s actual use case, measure performance, and watch where automation first becomes dependable. Familiarity with current limitations is part of preparedness, not merely a productivity initiative.

27. Cotra’s grantmaking career began as emergency triage

  • Cotra joined Open Philanthropy in 2016 but made no grants for more than six years. The FTX collapse changed that: hundreds of promised recipients faced cancellation or clawbacks, creating an emergency call that required surge capacity.

  • In roughly six weeks, she went from no grants to about 50. The experience showed her that she liked parts of grantmaking, but rapid triage prevented the depth she naturally sought.

  • Technical proposals in interpretability or adversarial robustness often looked reasonable without a complete causal story. Cotra wanted to specify how a research direction produced a capability or technique, how that technique fit a takeover-prevention plan, and what success would concretely mean.

28. Inside-view funding produced a narrow, expensive, opinionated bet

  • Open Philanthropy’s older heuristic—support a strong researcher with a good record who cares about AI safety—was defensible, but emotionally unsatisfying to Cotra. She wanted enough understanding to answer both more-doom and less-doom critics in sustained detail.

  • Her temporary compromise was a barbell: process renewals quickly using conventional heuristics while developing a smaller number of deeply justified programs. The main inside-view bet became realistic agent benchmarks and non-benchmark evidence about AI’s effects.

  • The agent-benchmark RFP ultimately made about $25 million in grants, with another $2–3 million through its broader RCT-and-survey companion. It specified hard, realistic tasks and repeatedly warned applicants that apparently difficult tasks were probably “not hard enough.”

  • Cotra acknowledges the opportunity cost: the same time could have funded more low-hanging opportunities across ten fields. Her defense is that deep homework changes details, reveals stronger versions of researchers’ ideas, and enables grantmaker and grantee to co-create better projects.

29. Burnout revealed the cost of operating without an institutional thought partner

  • When Holden left, Cotra lost a manager who would debate timelines, takeoff, and threat models at the object level. Remaining leadership had less AI context and less bandwidth, turning “Is this strategy right?” into “You can do it that way if you want.”

  • Cotra felt alone while trying to build an understanding-oriented program and discovered she was not naturally entrepreneurial. Hiring could not substitute for a shared central brain because the vision remained nebulous and candidates had to resonate with it before it was fully defined.

  • Perfectionism made management expensive. Writers rarely captured her ideas as she wanted, editing often took longer than writing, and delegating grant work could be slower than doing it herself—the familiar new-manager problem arriving just as guidance from above diminished.

  • Her four-month sabbatical mixed practical recovery, a new group house, exercise, involvement in the Curve conference, unpublished writing, and career reflection. She concluded that she wants to advise and help an organization’s center more than run an isolated internal startup.

30. EA’s appeal combined impartial care, intellectual depth, and radical integrity

  • Cotra entered the effective-altruism “rabbit hole” at 13. Its first attraction was moral scope: helping distant people, animals, future generations, and potentially conscious artificial systems rather than privileging those nearby or familiar.

  • The second was methodological seriousness: stop, research, quantify, and compare interventions that may differ by orders of magnitude. Her analogy is how intensely someone would investigate treatments if they or a spouse had cancer.

  • The third was an exacting integrity beyond what donors demanded. GiveWell maintained a public mistakes page and refused donation matching because it is usually a scam, even though using it might have raised more money.

  • As EA shifted from persuading analytical donors toward deploying money, talent, and political influence, radical openness became strategically costly. Donors sought privacy, campaigns could not reveal tactics, and adversaries began searching Open Philanthropy’s publications for material that could damage it.

31. Cotra wanted EA to be more religion-shaped, not less

  • Cotra regards EA as far more truth-seeking than religion, but accepts the analogy because it offers a “map of the good life,” a community, a worldview, and guidance that crosses personal conduct, politics, and one’s place in history.

  • What it supplied professionally was not quite what she needed emotionally. Her corner of EA idealized having an impactful job and working extremely hard; she wanted more deliberate existential reflection on morality and a world that might become utopian or dystopian within a decade or two.

  • Her imagined “EA church” would convene thoughtful discussion each week, reconnecting daily documents and emails to the larger stakes. Rob says he personally likes the more professional, limited aspect because he wants to go home and not think about the work all the time; Nathan, by contrast, says he wants to think about his work in a more spiritual way.

  • She recognizes the danger of cultishness and why a professional community should welcome good safety researchers without demanding philosophical conversion. Still, Joe Carlsmith’s popularity among committed EAs suggests she is not alone in wanting that form of nourishment.

32. Better local fit made the same work feel lighter

  • Cotra initially planned to start a Substack after sabbatical, admitting the choice was motivated partly by desire rather than a confident highest-impact case. She stayed when Open Philanthropy searched for new global-catastrophic-risk leadership and both leading candidates seemed strong.

  • Working with Emily Olsen restored the integrated role Cotra wanted: Emily had a larger project, needed Cotra’s answers, and would use them in strategy. Cotra found herself working more hours than before sabbatical while experiencing the work as less difficult.

  • At recording, she was also exploring Redwood Research and a work trial with METR; the cross-post introduction says she is now working on risk assessment at METR. Both offered narrower missions—AI control or early warning—along with the opportunity for deeper work.

  • Her least glamorous conclusion may be the most generalizable: “the literal person you’re reporting to matters a huge amount.” Leadership changes should trigger a fresh assessment of role, rhythm, collaborators, and fit even when the organization’s nominal mission is unchanged.

33. EA’s AI-era niche may be speculative research others avoid

  • Cotra thinks AI safety no longer requires the whole EA moral package; misalignment and misuse concern people with many values. EA remains distinctive where altruistic motivation, unconventionality, and tolerance for philosophical uncertainty are prerequisites.

  • Digital sentience is the clearest example: asking whether AI systems can suffer, deserve protections, or create trade-offs with safety is neither lucrative nor conventionally prestigious. EA can incubate the field before it becomes institutionally respectable.

  • Value lock-in need not require a dictator. Nathan imagines “social media plus plus”: personalized, superintelligent information bubbles that help each person defend existing beliefs indefinitely, leaving power distributed while society becomes increasingly incapable of reflection or revision.

  • The broader comparative advantage is “informed speculation” between storytelling and measurement. Nathan wonders whether some EAs steered into operations or policy, while privately wanting to be “a weird truth-teller,” should return to research precisely because the rest of the world undersupplies it.