Pioneers Insight Method Research Author
Helen Toner: OpenAI Reflections, Adaptation Buffers, and AI in Warfare
Back to Episodes

Helen Toner: OpenAI Reflections, Adaptation Buffers, and AI in Warfare

Summary

  • Toner’s base case is not imminent superintelligence; it is that a civilization-scale transition is plausible enough to prepare for now. In 2016, “short timelines” meant advanced AI within a couple of decades or a lifetime; by 2025, it can mean superintelligence before the 2020s end, so even “long timelines to advanced AI have gotten crazy short.” For investors, the discussion points toward sustained demand for evaluation, resilience, and government capacity—not confidence in a five-year countdown.
  • The clearest new OpenAI disclosure is that Q did not trigger the board’s 2023 decision.* Toner said the board knew reasoning research later released as o1 and o3 was underway, but received no breakthrough letter and did not act on one: “That whole Reuters story was totally false.” The broader governance discount remains harder to quantify because confidentiality and legal obligations prevent a complete public record.
  • Frontier-AI oversight works better when whistleblowers can point to a violated rule than when protection depends on subjective concern. Toner favors disclosed safety-and-security plans, capability and risk evaluations, and internal processes that create a crisp standard employees can invoke. Those employees may have their greatest leverage now because they are “actively working to replace themselves,” while institutional failure may arrive “boiling frog style” without one obvious crisis.
  • Adaptation buffers are the strategic answer to an AI market where frontier development gets costlier while yesterday’s frontier rapidly commoditizes. DeepSeek matched reasoning capabilities only one or two months behind some US systems—or base-model capabilities roughly six to nine months behind—at lower cost, making permanent nonproliferation increasingly invasive and brittle. The practical response shifts toward vaccine capacity, outbreak detection, cyber remediation, and distribution of defenses before capabilities diffuse.
  • Iterative deployment remains useful only while releases are unlikely to cause severe, irreversible harm. Toner prefers conditional “if-then” gates—do not advance until specified understanding or mitigation exists—over fixed pauses or input-based speed limits that would require unavailable legislation and immediately encounter the China objection. With no comprehensive regime in sight, transparency, measurement science, interpretability, alignment work, and technical government staffing are the practical building blocks.
  • “Beat China” is functioning simultaneously as a geopolitical argument and the AI industry’s path of least resistance in Washington. Toner grounds the rivalry in power transitions, maritime access, and international rules; the conversation also covers Taiwan. She says the framing conveniently supports “funding,” “government contracts,” “no regulation,” and protection from liability. That makes China rhetoric material to AI company economics even when it does not resolve the underlying strategic question.
  • Military AI should be evaluated use case by use case, not sold as a general-purpose battle buddy. Bounded tools for viewshed mapping, medical triage, ship-movement anomalies, or database retrieval differ fundamentally from an LLM asked to generate “three courses of action that are non-escalatory.” Toner and Amelia Probasco’s framework—scope, training data, and human-machine interaction—puts reliability and adversarial robustness ahead of demo fluency.
  • An “AlphaGo for the army” is not a credible near-term equilibrium because real war cannot be enclosed in a clean simulation. Battlefield, logistics, economic, political, and public-attitude dynamics interact while an adversary deliberately attacks the model’s assumptions. Toner’s honest conclusion is that she does not know what nation-states, democracy, or the Chinese Communist Party look like under superintelligence; strategic uncertainty, not a settled doctrine, is the central fact.

Deep dive

1. Toner took AGI seriously before it was respectable

  • Toner joined OpenAI’s board in 2021 but had known the company and many of its people since its 2015–2016 founding, when she was beginning AI-policy work in San Francisco. Because AlexNet had arrived in 2012 and deep learning was already advancing, she initially felt “late to the party”; the wider attention waves of 2018–2019 and ChatGPT in 2022 later changed that perspective.

  • Founding an organization explicitly to build AGI was then “weird” and “against the grain,” even within machine-learning circles; DeepMind was nearly the only serious research organization speaking that way. Toner believes her China and national-security expertise mattered to her board appointment, but so did having taken OpenAI’s mission seriously for years when the informed community was extremely small.

  • Her original timelines were short only by the standards of that era: very advanced systems within “the next couple decades” or “in our lifetime” seemed plausible enough to justify preparation. Today, short timelines can mean superintelligence before the 2020s end; she remains much less certain about that, while holding that the possibility still “warrants quite a lot of thought, quite a lot of preparation.”

  • Toner rejected the idea that preventing AI from killing humanity motivated her career. Her premise was historical: major technologies transform society “for better or for worse, often for better and for worse,” and AI looked likely to produce such a transformation within her lifetime. The question was whether her work could help it go better—not whether it would “definitely” kill humanity.

2. The Q* story was false, but OpenAI context remains constrained

  • Toner’s continuing limits are concrete: board confidentiality obligations, legal processes in which small inconsistencies could matter, and private conversations involving people who need not be drawn into public controversy. Much of the unreleased material is “almost like boring detail,” not a remaining “big deep dark secret” that would dramatically change the picture she gave on the TED AI Show.

  • One formerly confidential point is now clear because the underlying reasoning work is public. The board knew research later released as o1 and o3 was underway, but “we never got some letter about a breakthrough” and did not base its decision on an employee letter. Toner’s categorical correction: “That whole Reuters story was totally false.”

  • Erik pressed on whether OpenAI’s world-changing mission had become a heroic or “main character” culture shaped by elite-performance coaching and detachment. Toner declined to generalize: board members are poorly positioned to characterize day-to-day employee culture. Her broader concern was structural—technical work normally can be separated from social governance, but that separation becomes “out of distribution” if progress outruns society and only developers can avert enormous consequences.

3. OpenAI’s technical warnings and policy posture pull apart

  • Nathan described a recurring OpenAI whiplash: the obfuscated-reward-hacking paper offered one of the clearest warnings about AI failure, while a White House policy submission soon projected an escalatory China posture and sought sweeping freedom to train on copyrighted material. He paired that with Sam Altman’s “our values or their values; there’s no third way,” calling the institution’s outward voices almost “schizophrenic.”

  • Toner also finds the contrast confusing and thinks it has become more pronounced. One possible explanation is substantial employee freedom over public comments and research directions, combined with different review paths for official policy. She hopes technical staff are watching what the company advocates politically, given how strongly those messages can contradict the warnings emerging from its own research.

  • Her disclosure priority is to narrow the “huge information gap” between frontier developers and everyone else. Useful releases include capability and risk test results, descriptions of safety processes, and model specifications explaining what systems are trained to do. The objective is not a government checklist dictating the answer, but enough visibility for outside actors to understand the systems and respond.

  • The US reporting threshold around 10^26 operations appeared “on a wobbly footing,” though Toner had not heard that it was definitively dead; Nathan said the EU AI Act was developing transparency rules around models over roughly 10^25, amid pressure to dilute them. Toner’s rationale is not that every model above a compute line is dangerous—it is that the newest, best models carry the highest “unknown unknown risks” and merit extra scrutiny.

4. Whistleblowers need enforceable standards before crisis

  • Conventional whistleblowing usually concerns illegality: the SEC, for example, offers a defined channel for financial misconduct. Frontier-AI employees may instead observe conduct that feels dishonest or dangerously risky but violates no law. A protection framed merely as “if you’re worried, call this hotline” leaves workers and companies without a usable boundary.

  • Toner’s preferred design pairs protections with disclosure or process requirements. A company might publish or privately submit a safety-and-security plan, creating a standard against which employees can report that an evaluation was skipped, information was misstated, or a promised mitigation was abandoned. Even a required internal process helps if workers can truthfully say, “Actually, we didn’t carry out that process.”

  • Implementation must account for the proposed system’s users: technical employees who may lack legal sophistication, feel frightened, work extreme hours, and have little time to navigate ambiguity. “If the user is a whistleblower,” Toner asked, what is the user experience? Eligibility, the next step, confidentiality, and the reporting destination need to be obvious.

  • Frontier employees may be “in the most powerful position they’re going to be in” because they are actively automating their own work while tech labor’s leverage is already weakening. Waiting for a dramatic rupture may therefore fail twice: workers could become less indispensable, and misconduct could accumulate “boiling frog style”—each episode concerning, but none obviously the single moment to act.

5. Adaptation buffers beat permanent nonproliferation

  • Toner begins from confidence in social adaptability. New technologies from the printing press to television and the telephone repeatedly produced claims that “the sky is falling,” yet societies built institutions, norms, and barriers that made them positive on balance. AI as a possible successor intelligence may be different, but she resists treating every misuse risk as historically unprecedented.

  • AI development follows two curves at once. Pushing the frontier requires ever more compute, money, and concentrated expertise, making the first demonstration less accessible. Once a capability exists, however, engineering improvements drive its cost and difficulty down repeatedly—creating a temporary interval in which society has observed the capability but broad misuse remains comparatively hard.

  • DeepSeek illustrated the interval rather than leapfrogging the frontier: Toner described it as matching reasoning systems from one or two months earlier, or base models from roughly six to nine months earlier, at a lower price. The exact reported development cost was less important than the directional fact that frontier-like performance becomes cheaper and easier to reproduce.

  • Permanent AI nonproliferation would therefore demand escalating intrusion. Nuclear controls work because bombs require substantial quantities of highly enriched, specialized material. If nuclear efficiency improved like AI—until the uranium dispersed across “a couple acres” of farmland sufficed for a tiny weapon—the IAEA would eventually need to inspect farmhouses. Toner sees that as an analogy for why locking down broadly useful computation becomes untenable.

6. Resilience depends more on deployment than frontier capability

  • Adaptation means using the buffer to reduce consequences: expand vaccine manufacturing, wastewater monitoring, outbreak detection, and test-kit distribution so a biological attack is recognized and answered faster. It also includes ordinary counterterrorism questions—whom the FBI tracks and how early it detects plots—which have little to do with frontier models or even biological materials.

  • In cybersecurity, Toner challenged the narrow question of whether AI helps attackers or defenders more. The operational question is whether usable defensive products reach water-treatment facilities, power grids, and chemical plants run by teams without frontier-AI expertise. A brilliant model at a lab does not protect infrastructure until its capabilities are packaged, disseminated, and integrated.

  • Nathan offered AI-assisted formal software verification as a hopeful example, reporting a possible multiple-orders-of-magnitude speedup. Toner accepted the potential but emphasized timing: a released model can reach a disaffected teenager far faster than thousands of infrastructure operators can modernize old codebases, manage the division between IT and operational technology, and safely deploy new defenses.

  • Even if defenders ultimately benefit more, the transition can remain dangerous. “They’re not going to be nimble,” Toner said of infrastructure providers, so long-run defensive superiority does not erase the release-to-adoption lag. Thinking of it as a “transition period rather than like a permanent state of danger” produces better policy, but does not justify complacency during that interval.

7. Today’s policy menu is building blocks, not a regime

  • Toner has not seen a comprehensive frontier-risk regime she likes, partly because policymakers do not know exactly what problem will emerge or when. The realistic menu is a set of building blocks that improves future options—an appropriate response to uncertainty, but potentially inadequate if progress leaves “very little time.”

  • Dean Ball’s proposed regulatory market is one such component: government-accredited private regulators would assess developers, whose compliance could earn a liability shield; accreditation or protection could be withdrawn. Toner considered it “certainly better than nothing” and “far far far more feasible” at state level, while doubting it could counter the “cutthroat financial incentives” and other pressures pushing frontier companies to move as fast as possible.

  • Other building blocks include public funding for the science of measuring AI, interpretability and alignment research, focused research organizations, and greater technical capacity inside government. Transparency does not itself solve a failure, but it equips more actors to respond as problems emerge. Her preferred slowing mechanism is similarly conditional: an “if-then” gate tied to understanding or risk mitigation, not an arbitrary number of months.

8. Iterative deployment works only before irreversible harm

  • Nathan defended OpenAI’s original iterative-deployment logic: exposing society gradually to improving systems should be less disruptive than developing superintelligence privately and unveiling it at once. He worried that Ilya’s new company planned no release before superintelligence, GPT-4.5 might be removed if its compute cost outweighed demand, and Miles Brundage’s comments suggested internal deployments were becoming strategically more important.

  • His proposed relative speed limit would tie internal development to public deployment: a company could train only a specified multiple beyond the largest model it had released, forcing some visibility before a small group pursued systems it believed might transform global power. He framed it as a safeguard against abandoning iterative learning precisely when internal capabilities become most consequential.

  • Toner agreed with iterative deployment in principle, provided each release is unlikely to cause “really severe irreversible consequences.” Its value comes from releasing, observing, adjusting, and trying again; that logic breaks when the experiment cannot be recalled. Any serious iterative policy therefore needs criteria identifying when the next iteration should not enter the world.

  • She was less convinced an industry-wide retreat had already occurred and skeptical that Nathan’s speed limit was implementable. It would likely require federal legislation that Congress will not pass, followed immediately by “doesn’t it just mean that we lose to China?” Conceptually plausible controls still fail without a political mechanism; she again preferred progress conditional on demonstrated understanding and mitigation.

9. A crisis could unlock Congress—and trigger an overreaction

  • Toner agreed that many lawmakers simply do not believe superintelligence forecasts, but added that institutional dysfunction is independent of AI. A deeply experienced congressional observer told her the House may be “less functional than it’s been since after the Civil War.” Even after a technological shock changes beliefs, thoughtful federal legislation would remain a major lift.

  • Some policy specialists are prewriting a “Patriot Act for AI,” expecting a crisis to create a short legislative window. Toner considered advance preparation reasonable while warning that the eventual bill could fight the last war or smuggle through unrelated priorities. Three Mile Island supplies the counterexample: disproportionate regulation after one incident effectively shut down US nuclear development without comparing its risks to alternative energy sources.

  • Developers therefore have a self-interested reason to install credible guardrails before an accident; otherwise, Toner sees the system heading toward a “massive knee-jerk overreaction.” The trigger may be vivid rather than statistically severe: Kevin Roose’s Sydney exchange drove fear after GPT-4 despite being heavily elicited, and unrestricted celebrity voice cloning could likewise provoke backlash disproportionate to its underlying risk.

10. “Beat China” is both a geopolitical claim and a lobbying shortcut

  • Toner started from power-transition logic: the US is an established power, China a rising one, and control over international rules matters. Peaceful transitions such as the US eclipsing Britain are historically unusual. The postwar Pax Americana replaced much “might makes right” behavior with sovereignty, stable borders, and institutions supporting trade and freedom of navigation.

  • US policy tried through the 1980s, 1990s, and 2000s to make China a “responsible stakeholder,” despite the rupture of Tiananmen Square, culminating in WTO membership. Toner argued that this project had clearly deteriorated by Xi’s rise in 2012 as China became more illiberal and hostile in the South China Sea. Nathan separately raised questions about Taiwan and possible Chinese aggression around Japanese, Filipino, and Korean assets.

  • Nathan’s pushback—worth keeping—was that American CEOs had shifted from warning against a China race to demanding an unassailable US lead, while Chinese labs were openly releasing models and Americans discussed preserving a unipolar world. Toner replied that openness makes strategic sense for a follower seeking attention and talent; it is not evidence of “pure goodwill and lack of competitive spirit.”

  • Toner nevertheless found the CEOs’ rhetorical reversal “pretty striking.” Her explanation was political economy: “We got to beat China” is the one message Washington can agree on, making it the “path of least resistance” for companies seeking funding, government contracts, freedom from regulation, and liability protection. The US retreat toward its own might-makes-right posture further complicates the moral clarity of the rivalry.

11. Military AI already spans bounded automation and broad judgment

  • Toner’s paper with Amelia “Emmy” Probasco draws on unusual operational experience: Probasco served in the Navy operating Aegis missile-defense systems. Developed without deep learning in the 1960s and 1970s and deployed in the 1980s, Aegis already detects incoming missiles and can automatically identify and engage threats—evidence that meaningful weapons automation predates today’s AI debate.

  • Their focus is decision support because military-AI discussion too often stops at autonomous weapons. The category ranges from computer vision calculating viewsheds—where a sniper or other operator can see—to battlefield medical triage and anomaly detection in ship movements. These are bounded functions that can reduce the fog of war without claiming general strategic judgment.

  • At the expansive end, companies including Palantir and Scale AI have marketed LLM-based systems as something closer to an all-purpose “battle buddy.” The demos combine well-grounded retrieval—locating the nearest unit or checking its missiles against a database—with open-ended reasoning whose reliability and provenance are far less clear.

12. A battle buddy is only as safe as its scope, data, and interface

  • Toner’s sharpest example was a demo asking the model to “generate three courses of action that are non-escalatory.” The system appeared to draft tactics and routes under that political constraint, prompting her question: “How the hell does the LLM know what’s escalatory and non-escalatory?” The polished interface concealed unresolved questions about who defines, tests, and validates such judgment.

  • The first evaluation dimension is scope: a tightly bounded activity can be tested against its intended operating conditions, while a sprawling assistant responds to whatever an operator thinks to type. Generality increases the surface over which plausible language may be mistaken for operational competence.

  • The second is data: what trained the system, and how closely does that data match the present conflict? Military environments are both novel and adversarial. An opponent is actively trying to deceive sensors, corrupt assumptions, or induce behavior outside the training distribution, making ordinary benchmark reliability an inadequate proxy.

  • The third is human-machine interaction: the product must communicate what it can and cannot do, prevent overtrust, and help its operator reach a sound decision under pressure. Military organizations have often undervalued interface design, Toner argued, even though poor presentation and misunderstood automation have contributed to unwanted outcomes.

13. War is not a game clean enough for AlphaGo

  • Nathan feared an “AlphaGo for the army”: self-play and simulation could hill-climb toward superhuman tactics operating faster than human commanders, producing inscrutable, hyperlethal systems. Toner thought this was not close, because a credible simulation must connect individual battles to theater logistics, global asset deployment, economics, politics, public attitudes, and countless unknown interactions.

  • Dan Hendrycks, Alexandr Wang of Scale AI, Eric Schmidt, and co-authors offered “mutually assured AI malfunction”: a power threatening to build superintelligence and rule indefinitely would invite rivals to sabotage its project. Toner found the multiparty logic useful, including the possibility that vulnerable projects change what actors attempt, but not comparable to the stability of nuclear mutual assured destruction.

  • Nuclear deterrence was legible: states understood what weapons did, how second-strike capability worked, and why no one wanted an exchange. Superintelligence’s form and strategic value remain unclear. Sabotage may be easier than securely building an AI project, but that uncertainty cannot produce the “crystal clear strategic logic” of Cold War deterrence.

  • David Chapman’s distinction between abstract rationality and practical “reasonableness” supplied Toner’s deeper objection. Chess, Go, and StarCraft are not clean originals that reality approximates; they are “very unusual special cases” carved from a messy world where people act, observe, and readjust. Her final answer remained an honest non-answer: she does not know what nation-states, democracy, or the Chinese Communist Party become in a superintelligent world.