Pioneers Insight Method Research Author
A.I. Safety Is So Back + Mythos Mayhem with Nikesh Arora + Hot Mess Express
Back to Episodes

A.I. Safety Is So Back + Mythos Mayhem with Nikesh Arora + Hot Mess Express

Summary

  • Claude Mythos has forced a rapid Washington safety U-turn by making dangerous cyber capability concrete rather than hypothetical. A rumored executive order would create Biden-like prerelease model reviews that Trump canceled on his first day back in office, after Republicans had attacked such testing as anti-innovation. Casey Newton’s verdict: the administration’s worldview “did not survive contact with reality.”

  • The government still lacks a coherent model-access strategy, creating policy risk across chips, contractors, China, and allied cybersecurity. The Pentagon is simultaneously fighting to designate Anthropic a supply-chain risk and installing Mythos to scan for vulnerabilities; Trump is exploring Nvidia chip access for China while the administration has not resolved whether China should get Mythos. The result is an administration “installing and uninstalling Anthropic at the same time.”

  • Cyber defense is being repriced from a days-long response problem into a minutes-long infrastructure race. Palo Alto Networks found 26 critical exploits covering 75 issues, versus a typical baseline below five, while Mozilla reported 423 fixes in April against a 2025 monthly average near 22. Nikesh Arora says legacy defenses were “designed for days,” so enterprises must overhaul them to “fight AI with AI.”

  • Mythos is powerful because sustained compute lets it chain weaknesses together, but it is neither automatic nor infallible. Arora reported roughly 30% false positives and said performance improved only after Palo Alto supplied code purpose, expected behavior, and threat intelligence from 10,000 attacks over five years. Mythos and GPT-5.5 Cyber found different issues, while Mythos’s compute-intensive “ultra mode” made persistent experimentation and “daisy-chaining vulnerabilities” more effective.

  • The near-term economics favor attackers, while a large remediation cycle requires scaled cybersecurity vendors and integrators. Defenders must be right 100% of the time while an attacker needs one working vulnerability; firewalls can provide “temporary scaffolding,” but open-source dependencies and unmanaged endpoints remain slow to patch. Arora expects enterprises to undergo a three-to-six-month “cleansing of the vulnerability backlog,” supported by firms including IBM, PwC, Deloitte, and Accenture.

  • The most exposed organizations are technology-dependent businesses whose core competency is somewhere else. Arora is less worried about well-resourced financial institutions than hospitals, small businesses, industrial operators, and medical practices—the “95% something else” companies that lack engineers. He cited the Change Healthcare breach as an example of how an incident can halt a physician ecosystem. Consumer email and telecom providers also need stronger gatekeeping before AI makes phishing materially more convincing.

  • AI agents enlarge the attack surface precisely by becoming useful enough to hold credentials and act autonomously. Arora called OpenClaw “a scary thing from a security perspective” because it can be given permissions and credentials to act across accounts; his segregated installation is “totally useless” because it cannot reach his calendar or email. That tradeoff—capability requiring access, access creating risk—will shape enterprise agent deployment.

  • Arora expects AI productivity to expand engineering output before it eliminates engineering demand. Feature backlogs already extend six to 12 months, so gains of 30% to 60% can fund more development; announced workforce reductions of 7%, 15%, or 20% may instead create room for people with newer skills. His broader call is a “decade-long transformation of business,” with functional efficiencies paying for tokens and additional AI capacity.

Deep dive

1. Mythos overturned Washington’s “let them cook” consensus

  • Kevin Roose’s opening evidence was a rumored executive order creating an AI working group and potentially requiring government review before frontier models ship. The resemblance to Biden’s revoked framework is striking: similar testing was previously derided as “communist,” anti-innovation, and a route to losing the AI race against China.

  • Casey’s causal account was blunt: Mythos can apparently discover novel, exploitable weaknesses across many programs, so the administration’s laissez-faire posture “did not survive contact with reality.” Once officials saw a model that could cause “vast amounts of harm” if broadly released, the practical question became how government could prevent that harm.

  • The institutional fight pits CAISI—the renamed U.S. AI Safety Institute—against intelligence agencies such as the NSA. It also pits former AI czar David Sacks’s “let them cook” philosophy against Republican officials newly willing to treat model capability as a national-security threat.

  • Casey argued that Trump’s position was always narrower than Republican opinion: Republicans and Democrats were both deeply skeptical of AI, and Republican state legislators were already racing to regulate it. Mythos simply made the bill come due for a federal faction pursuing “all gas, no brakes.”

2. A safety regime still has an enforcement gap—and a censorship risk

  • CAISI has researchers specifically hired to evaluate increasingly dangerous models, including people who joined despite reservations about serving the administration. Casey nevertheless identified the missing mechanism: what happens when evaluators deem a model too dangerous, but its developer invokes “business imperatives” and releases it anyway?

  • Kevin’s pushback—worth keeping—is that regulation can become political coercion, as social-media oversight did. Casey agreed that prerelease review might operate as prior restraint, with officials blocking a model “not because it’s actually dangerous, but just because it seems woke and gay”; litigation may eventually be necessary.

  • Their provisional balance favored intervention despite that risk. Casey would currently rather hear, “The crazy cyber model, don’t give that to everyone,” while Kevin welcomed government finally admitting the technology might require action: “I’ll take the little wins where I can get them.”

3. Chips, models, and allies collide in an incoherent China strategy

  • Trump traveled to China with Jensen Huang, Elon Musk, Tim Cook, and Meta’s Dina Powell McCormick while AI talks with Xi Jinping were reportedly on the agenda. Kevin saw an inherent contradiction: sell China the Nvidia hardware needed to build Mythos-caliber systems while trying to deny it today’s Mythos.

  • The same contradiction appears inside the Pentagon. It is defending Anthropic’s supply-chain-risk designation—imposed after Anthropic rejected contractual permission for any “lawful use”—while also deploying Mythos during the period when Anthropic technology is supposedly being removed.

  • A Chinese think-tank representative approached Anthropic in Singapore seeking access, while Germany proposed its own CAISI-like institution and demanded state-of-the-art models. Casey favored greater Western cooperation: fixing internet-wide vulnerabilities may require “all the help we can get,” after the U.S. had effectively told allies it was winning and they could “like it or learn to live with it.”

  • Kevin said the “AI is just a normal technology” position becomes untenable once models find zero-days and military and intelligence agencies alter their behavior around them. His darker forecast was continued contradiction until “some big event” forces officials to sit up straight.

4. Cyber compromise moved from days to minutes

  • Arora’s seven-year comparison defined the inflection point: attackers once needed days after entry to extract an organization’s “crown jewels”; with AI, that interval has compressed to minutes. Defenses built for human-paced response now need automated detection and action on the same timescale.

  • Palo Alto disclosed 26 critical exploits covering 75 issues, against a normal baseline below five—roughly five to seven times the usual discovery rate. Arora called the exercise a “great cleansing” of accumulated “tech debt or vulnerability debt,” conducted by hundreds of engineers across every product.

  • The pattern extended beyond Palo Alto. Mozilla pushed 423 security fixes in April versus about 22 per month during 2025; Google’s threat-intelligence group identified its first attacker using a zero-day it believed was AI-developed; and the Canvas attack forced an outage and negotiation over stolen data.

  • Arora cautioned against extrapolating the sevenfold spike forever: the concentrated audit should clear much of Palo Alto’s backlog. But every organization must now determine how much old code contains similar weaknesses, and open-source components will generally be remediated more slowly than proprietary software.

5. Mythos is a context-hungry force multiplier, not a magic scanner

  • Arora’s first impression was less dramatic because Mythos flags too much: about 30% of findings were false positives, each requiring verification. Its usefulness rose as engineers explained what code was meant to do and what normal behavior should look like.

  • Palo Alto then supplied its proprietary threat corpus—techniques drawn from roughly 10,000 attacks over five years—and asked whether known methods could apply in new contexts. Arora described that as giving the model “all the human training of the past” to build future defenses.

  • Mythos and GPT-5.5 Cyber discovered different weaknesses, suggesting their training and grounding produce complementary coverage rather than a single definitive answer. For Arora, that divergence means “there is still a lot that’s gonna get found.”

  • These systems also identify configuration mistakes, not merely defective code. Arora’s cleanest example was an internet-exposed product control panel left open for remote convenience: “If I can find it, other people can find it too.”

6. The patch cycle cannot keep pace with automated exploitation

  • Casey challenged the traditional 90-day responsible-disclosure window with Palo Alto’s finding that AI-assisted attackers could gain initial access and exfiltrate data within 25 minutes. Arora agreed the window will shrink, though “how much does it shrink” remains unsettled.

  • SaaS is the easier case: software can be investigated, patched, and deployed centrally, as Palo Alto did within two or three weeks. Laptops, servers, switches, and routers require organizations to act, so Arora expects three to six months of unusually frequent updates as the backlog is cleansed.

  • Integrators including IBM, PwC, Deloitte, and Accenture are mobilizing remediation resources. Where immediate fixes are impossible, Palo Alto can encode known vulnerable paths into perimeter-firewall signatures, creating “temporary scaffolding” that blocks exploitation while an organization repairs the code behind it.

  • The attacker-defender contest remains structurally asymmetric: “We have to be right 100% of the time. The bad guys are right once.” Four blocked vulnerabilities and one successful exploit still earn the defender “zero,” so Arora believes attackers currently capture more value from comparable model capability.

7. Restricted Mythos access buys defenders time, not a permanent advantage

  • Arora credited Anthropic and OpenAI with trying to expose the “art of the possible” responsibly while defenders still had time to react. “They partly got most of it right,” he said; both fumbled portions of the rollout, but there is no easy distribution policy.

  • Mythos’s distinguishing property is its compute-heavy “ultra mode,” which can persist much longer than the “flash mode” typical of released models. Persistence lets it try multiple techniques and chain successful steps, making the compute cost—and the risk—about sustained search as much as base capability.

  • Arora therefore favored giving companies time to fix systems before equivalent access becomes general. If attackers obtain Mythos-level tools, ransomware and nation-state economic harm remain familiar outcomes; what changes is “the pace and the volume” of attacks, not their fundamental nature.

  • The four-to-six-week evaluation window allowed defenders to study model behavior and build AI-powered sensors before what Arora called a “tsunami of AI-based attacks.” The race is whether those protections and patches arrive before nation-states, open-source actors, or third parties reproduce the capability.

8. Under-resourced industries and consumers form the weak layer

  • Arora worries most about organizations whose business is “95% something else” and only 5% technology: hospitals, small businesses, industrial manufacturers, infrastructure operators, and medical practices. Financial institutions have engineering depth; a doctor’s office can be paralyzed by an upstream incident like the Change Healthcare breach.

  • Consumers receive weaker protection than enterprises because they lack an effective universal gatekeeper. Corporate defenses can observe a phishing sender at one customer and block it elsewhere, whereas personal email and telecom providers need stronger controls against impersonation attempts that should be easy classifiers for companies “building AI.”

  • Kevin framed the familiar consumer playbook as strong passwords and multifactor authentication, while Arora emphasized provider-level controls and urged people to install software updates. Kevin’s daily fake X password-reset emails illustrated the looming problem: within six months or a year, he expects the same lure to become far more convincing.

9. Autonomous agents turn useful permissions into security liabilities

  • Arora called OpenClaw “a scary thing from a security perspective” because it can be given credentials and permissions, then act across a user’s accounts. At dinner, one enthusiastic adopter showed off an agent named Zara while the person beside him reacted: “Holy shit, that’s a security nightmare.”

  • His own OpenClaw runs on a segregated device disconnected from his calendar and email, which makes it “totally useless.” The mechanism is straightforward: an agent without access cannot book meetings or answer email, while one with access can act on the user’s behalf.

  • Engineers are simultaneously excited, overworked, and fearful. Among Palo Alto’s 9,000-plus technical staff, Arora sees every emotion because the immediate tools are promising while their implications over the next two or three years remain radically uncertain.

10. AI expands the backlog before it shrinks technical employment

  • Arora rejected the inference that 30%, 40%, 50%, or 60% productivity gains automatically mean fewer engineers: “I need more.” Product roadmaps already stretch six to 12 months because teams lack capacity, so the first gains should flow into long-deferred features and testing.

  • Companies announcing headcount reductions of 7%, 15%, or 20% may be “reshaping” rather than permanently shrinking, creating capacity to hire people with newer skills. Across the business, efficiencies in finance, HR, and other functions are the likely source of money to pay for tokens.

  • His macro framing was a “tsunami of a desire to transform” and a decade-long business transition. A CFO or HR leader does not want AI in the abstract; each wants a more efficient operation, whether through automated assessment, interviewing, or internal workflows.

  • Kevin resisted the “AI assessor” as a bad candidate experience, but Arora argued it could evaluate domain skills better than conversation. His preferred test is demonstrable output: when an applicant claims AI fluency, “Show me”—a simplistic recipe-to-shopping-list agent does not establish serious capability.

11. The Hot Mess Express shows adoption colliding with trust

  • Venmo began testing a friend-only default for new users, addressing a long-running privacy failure whose public transactions helped reporters identify accounts or payments involving Joe Biden, J.D. Vance, and Matt Gaetz. The hosts classified it as belated cleanup: it “used to be a very hot mess,” but the easy investigative trail is finally closing.

  • At Amazon, employees reportedly generated unnecessary Meshclaw activity to increase token consumption and look better to managers. Casey invoked Goodhart’s law—“When a measure becomes a target, it ceases to be a good measure”—and dubbed the result a “hot mesh.”

  • University of Central Florida arts and humanities graduates booed the claim that AI is “the next industrial revolution”; Kevin sees roughly 80% of students he meets saying, “I hate this.” Casey defended the audience, while Kevin demanded that anyone who used ChatGPT academically disclose that history before booing.

  • Grindr’s Madonna ad played “Hi Grindr, it’s Mother” even when users’ phone volume was off, potentially outing users near family; Casey called it a dangerous mess, not merely a marketing failure. Elsewhere, Dua Lipa sought $15 million over Samsung packaging, eBay dismissed GameStop’s apparently underfunded $55 billion offer as “neither credible nor attractive,” and testimony suggested Elon Musk had considered passing OpenAI control to his children.