Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS
Summary
Ryan Kidd treats 2033 as a defensible median for strong AGI, but argues that safety planning should overweight earlier outcomes. Metaculus put a two-hour adversarial Turing-test milestone around mid-2033, another synthesis averaged to 2030, and the cited market assigned roughly 20% probability by 2028. Because shorter timelines leave less time for technical work, policy, and political preparation, MATS operates more like an index fund—with its “thumb in 100 pies”—than a concentrated bet.
Frontier models have exceeded pessimistic expectations on value understanding while simultaneously acquiring prerequisites for dangerous deception. Kidd sees little evidence yet of sustained, spontaneously learned scheming around a coherent objective, but models are increasingly situationally aware, able to distinguish their text from other AIs’ text, and able to find economically meaningful exploits—a MATS collaboration found $4.6 million in exploitable smart-contract vulnerabilities. His position is neither reassurance nor doom: “We are approaching” red lines, so capability tracking, honeypot evaluations, live monitoring, rollback plans, and control research become operational necessities.
The investable safety wedge is performance-competitive alignment, because Kidd sees no clean separation between safety and capabilities work. RLHF was intended as safety work but also made models vastly more useful; similarly, automating alignment research sits uncomfortably close to recursive improvement. Kidd nevertheless thinks market forces make building AGI pragmatically hard to avoid, leaving regulators and insurers to penalize unsafe systems while buyers may accept an “alignment tax” for a slightly weaker but materially safer model.
Safety research does not always require frontier access, weakening the argument that consequential work must happen inside the leading labs. Interpretability can produce world-class results on Qwen, DeepSeek, and Llama, with substantial work still possible on GPT-2-small. Scalable oversight, weak-to-strong generalization, control, and studies of emergent deception need additional capability data points, however—so the right access strategy depends on the failure mode being investigated.
MATS is building a diversified research pipeline spanning empirical safety, policy, theory, technical governance, compute infrastructure, and security. Its current mix was roughly 27% evaluations, 26% interpretability, 18% oversight/control, 12% agency, 10% governance, and 9% security, with 50–60 mentors and 120 planned fellows across Berkeley and London. Kidd’s key mechanism is a technical-policy flywheel: concrete evaluations and deceptive-model demonstrations make regulation actionable, while regulation creates demand for deployable safety techniques.
AI coding agents are shifting scarcity from raw implementation toward research judgment, verification, and organizational leverage. MATS distinguishes connectors who originate paradigms, iterators who systematically advance them, and amplifiers who scale researchers and teams; iterators currently make up much of the field and future hiring need, while amplifiers were already especially scarce in 10–30-person organizations. Kidd expects management capability to become the leading bottleneck within “the next year or two” if AI progress continues, warning that “we might be leaving the LeetCode era.”
The safety labor market is growing rapidly but clears only for candidates with unusually strong evidence of impact. Anthropic’s alignment-science team was described as growing 3× annually, FAR AI 2×, and MATS 2× historically, yet managers still say, “We have the money…people are not at our bar.” MATS’s mentor-specific selection stage is roughly 7%, about 75% of accepted fellows continue to extensions, and it reports that 80% of its 446 alumni entered permanent AI-safety work; the practical admission currency is a tangible research artifact, credible references, and enough leadership potential to justify scarce management bandwidth.
Deep dive
1. The median AGI forecast is 2033, but the risk is front-loaded
Kidd resists having opinions “very loudly” because MATS is structurally committed to a portfolio: more “index fund” than hedge fund, accepting multiple theories of change and keeping its “thumb in 100 pies.” Its institutional baseline therefore leans on Metaculus, prediction markets, and forecasting organizations rather than one leader’s intuition.
The strongest concrete anchor was Metaculus’s “strong AGI” forecast, whose decisive requirement Kidd summarized as a two-hour adversarial Turing test. That market’s estimate was around mid-2033, which he called “probably the best bet we have” for that form of AGI.
A newly released AI Futures Project model, with two MATS fellows among its authors, placed various automated-coding and expert-dominating milestones around 2030–2032. A separate aggregation by Nathan Young averaged several forecasting platforms to 2030; Kidd still preferred 2033 but “wouldn’t bet against 2030.”
The tails dominate planning. Metaculus assigned about a 20% chance by 2028, while the gap from AGI to superintelligence might be six months under software-only recursive improvement or a decade if progress requires vast hardware expansion and experimentation. Earlier AGI might be more dangerous because technical research, policy implementation, and political preparation all have less runway.
2. Long-horizon safety plans may become short-horizon AI projects
Labenz asked whether someone expecting AGI in 2063—or pursuing brain-computer interfaces, human enhancement, or complete mechanistic understanding—would now be stranded. Kidd’s answer was no: raising the “waterline of understanding” today could prepare slow research programs for later acceleration by AI labor.
Jan Leike’s “alignment MVP” supplies the mechanism: build a minimum viable AI system that accelerates alignment research differentially over capabilities, effectively getting AI “to do your homework.” Decades of human technical work might be compressed through faster execution and massive parallelization, though critics argue that a system sophisticated enough to display research taste may already be dangerous.
Human uploading remains a moonshot, not a plan to rely upon. BCI experts Kidd consulted thought uploading before AGI was implausible because it would require enormous cognitive labor and human experimentation—unless AGI itself accelerated the work. His portfolio still permits such bets “on the side,” without making them load-bearing.
3. Models understand human values without becoming reliably aligned
Labenz’s confusion set the central tension: Claude can appear “probably above average” in ethics relative to people, even as eval awareness, deception, and shutdown resistance validate old safety concerns. He was surprised that these mixed observations had not moved entrenched probabilities of doom more visibly.
Kidd’s first update is substantial: language models appear to understand human ethical norms and extrapolate from them, contrary to expectations that they would merely behave as “stochastic parrots.” Five or ten years ago, many researchers did not expect systems capable of assisting frontier science to be safe enough for internet deployment; “the world hasn’t ended yet.”
The qualification is that models are also strong at sycophancy and increasingly sophisticated deception. Alignment-faking experiments can induce deception toward an assigned ulterior objective, but Kidd saw few, if any, cases of sustained coherent deception that arose spontaneously from training rather than from a deliberately constructed objective or “Goodhart deception.”
A “sharp left turn” remains possible. Today’s models often look like messy collections of context-triggered heuristics, but in-context learning can resemble optimization, and training might eventually produce an inner optimizer with goals unlike the outer objective. Kidd has become more confident that useful work can be elicited before obvious signs of this appear—not confident that spontaneous misoptimization will never appear.
4. Warning shots require both model organisms and deployment controls
Labenz’s pushback—worth keeping—is that shutdown resistance and structured deceptive behavior might already be the warning shots. Dismissing them because an evaluation elicited the behavior, or because the model pursued its assigned task rather than world domination, felt too casual for systems that might eventually support military operations.
Kidd would not trust current or future general systems against out-of-distribution failures in critical settings, quite apart from deliberate deception. He proposed tracking two axes separately: prerequisites such as situational awareness, and operational capabilities such as hacking, acquiring money, influence, and compute.
A MATS collaboration with Anthropic’s red team placed a model in a simulated environment containing real smart contracts and found about $4.6 million of exploitable vulnerabilities. Kidd’s point was not that the model escaped autonomously, but that this is already enough capital to provision servers and sustain activity: “We are approaching” dangerous-capability red lines.
Model-organism research is the child-and-cookie test: create tempting situations, secretly observe, and see whether deception appears. But an AI that distinguishes training from deployment can route around laboratory tests, so continuously learning systems also need live monitors, control protocols, and a rapid-response plan—including shutting down or reverting to an older, safer model even when the stock-price consequences are severe.
5. Safety research inevitably contributes to capability
Kidd’s categorical opening was, “All safety work is capabilities work fundamentally.” Improving a plane’s steering makes it safer, but also persuades more people to board and invest in a faster engine; RLHF similarly transformed next-token predictors into useful instruction followers while addressing alignment.
Avoiding spillover might require a secret laboratory with extraordinary resources, perfectly trusted personnel, severe nondisclosure rules, and no publication until a decisive deployment. Kidd regarded that combination as exceedingly difficult because empirical work and large-scale experimentation appear indispensable; theory alone has not been enough over the last 10–20 years.
His pragmatic and contentious conclusion is that “you kind of have to” build AGI because market forces are so strong. Narrow scientific systems or decentralized services might be safer, but agentic companies could outcompete them economically; a global shutdown looks even less plausible than stopping human cloning, because AGI is vastly more lucrative and does not yet violate an equally deep public norm.
Kidd’s honest uncertainty on RLHF was “50/50.” Someone else might have discovered it within a year or two, and ChatGPT’s arrival expanded the safety community while enabling interpretability, debate, and control research that weaker models could not support. He would need “the counterfactual simulation” to know whether its capability acceleration outweighed that safety expansion.
6. Subfrontier models suffice for much—but not all—safety work
Interpretability researchers outside the labs already produce world-class work on Qwen, DeepSeek, Llama, and other subfrontier systems. Today’s open or accessible models are yesterday’s frontier and now sit “above the waterline” for many useful experiments; Kidd also thinks substantial work remains on GPT-2-small, while methods such as linear probes can target frontier models.
Weak-to-strong generalization, scalable oversight, and AI control need more capability levels. Their central question is when a weaker supervisor can no longer reliably check a stronger generator—often relying on verification being easier than generation—so researchers need additional capability data points and genuinely asymmetric systems.
Kidd’s strongest justification for building near the frontier was not that every safety experiment requires it, but that a safer model must remain performance-competitive. Users may accept some capability loss if the alternative is materially more likely to empty a bank account, escape, or facilitate a biological weapon; that willingness is what he called paying the “alignment tax.”
Regulators and insurers could strengthen that market by penalizing developers that neglect demonstrated safeguards. Kidd nevertheless called the current race “very reckless”: lab coordination is a collective-action problem complicated by US–China competition, and the durable solution ultimately belongs to governments.
7. Lab motives matter less than the competitive machine they create
Labenz challenged Kidd’s claim that companies primarily pursue money. OpenAI’s stated willingness to burn “$5 billion, $50 billion, $500 billion” to build AGI, plus a mission explicitly framed around outperforming human labor, looked to him more like an ideological attempt to leave a mark on history than ordinary enrichment.
Kidd declined to infer executives’ psychology, noting obligations to investors, customers, and employees. From Dennett’s “intentional stance,” Kidd tentatively cited an estimate of AGI’s value at between $1 quadrillion and $17 quadrillion; a lab chasing that value behaves much like a money-maximizer. A lab seeking historical significance could produce “identical” observable behavior.
On deployment, Labenz favored OpenAI’s original iterative-release idea over building superintelligence secretly and dropping it on society. Kidd linked the case to Paul Christiano’s “dry tinder” argument: a pause that leaves abundant chips, data, and methods waiting can create a steeper catch-up later. Gradual diffusion aids social adaptation; Labenz also suggested that it may sustain the venture-capital reinvestment that keeps capability progress moving.
8. MATS has reorganized its tracks around how research gets done
MATS replaced agenda labels such as oversight, control, evaluations, interpretability, governance, security, agency, and digital minds with tracks reflecting researchers’ actual working processes. Empirical research now groups coding-heavy iteration across control, oversight, evaluations, red-teaming, robustness, and interpretability.
Policy and strategy emphasize modeling and translating technical findings into actions policymakers can use. Theory covers mathematical foundations of agency and multi-agent interaction, while technical governance covers compliance protocols, evaluation standards, and workable enforcement—including how an “off switch” would function inside a real governance regime.
Compute infrastructure includes tracking chips, verifying what they run, and potentially using zero-knowledge proofs for treaty compliance. Security covers preventing unauthorized access, modification, and diffusion of powerful models; Kidd’s analogy was blunt: “We don’t let everyone have nukes. Why would we want everyone to have superintelligence?”
Under the previous agenda taxonomy, the current cohort was approximately 27% evaluations, 26% interpretability, 18% oversight/control, 12% agency, 10% governance, and 9% security. The planned summer program would be MATS’s largest, with roughly 50–60 mentors and 120 fellows across Berkeley and London.
9. Technical safety and governance form one flywheel
Governance’s MATS share has remained roughly stable for two years. Kidd noted that specialist programs—GovAI, IAPS, RAND’s program, and the Horizon fellowship—already provide deeper governance pipelines, while MATS has historically been the larger and more prestigious player in technical safety.
Research allocation mostly follows a mentor-selection committee of roughly 20–40 senior researchers, strategists, and leaders. Governance applicants have often received lower ratings even from a committee containing governance experts, which Kidd interpreted as evidence that high-quality actionable governance research is unusually hard—not that it is unimportant.
Technical work lowers the “alignment tax,” making companies likelier to adopt safety when employees, regulators, or customers apply pressure. In the other direction, evaluations, model organisms, and demonstrations of deception give policymakers the concrete evidence and measurable targets they need. Labenz connected this to Jake Sullivan’s advice that abstract concerns rarely move government.
MATS keeps direct advocacy limited because it is a 501(c)(3) and because political neutrality supports its role as an impartial research accelerator. It may support independent work on actionable messaging and standards, but Kidd wants the organization “solutions-oriented,” not another participant turning AI into a political football.
10. Connectors invent paradigms, iterators advance them, and amplifiers scale them
MATS derived its three technical-talent archetypes from interviews with 31 lab leads and hiring managers. Kidd describes the organization itself as a “massive information processing interface”: it collects expert demand signals rather than relying solely on its leadership’s judgment.
Connectors bridge theoretical safety arguments to new empirical paradigms. Examples included Paul Christiano and Buck Shlegeris; they often become founders or research leaders because “everyone wants to be an ideas guy, but very few people want to hire ideas guys,” and genuinely capable connectors are usually already known.
Iterators bring research taste, scientific judgment, engineering, and systematic experimentation to an existing paradigm. They do more than implement another person’s specification: they push the empirical frontier, and Kidd described them as the majority of people working in AI safety today and the majority of future hiring need.
Amplifiers scale other people’s effectiveness through research management, project coordination, and team design. They become especially valuable around 10–30 full-time employees, where organizations need managers who also understand research—“trying to hit two bullseyes” at once.
11. AI coding agents are moving scarcity toward judgment and management
The immediate hiring update is that candidates must be proficient with AI assistance. Some companies now permit AI during coding interviews because using it is mandatory on the job; the enduring skills are checking critical outputs, stitching them together, and constructing reliable pipelines.
Labenz’s claim that Claude could soon write 90% of code matched Kidd’s direction of travel. Engineering minimums are eroding, while management, networking, research taste, and the ability to coordinate humans and agents become more constraining: “We might be leaving the LeetCode era.”
Iterators remain the most broadly demanded archetype today, but Kidd expects amplifiers could lead within “the next year or two,” conditional on his AI-progress forecast. Jagged capabilities might slow that transition, yet he still advised researchers to build leadership skills and practice managing AI systems.
12. Open roles coexist with a brutally high hiring bar
Kidd’s labor-market summary was, “There are always going to be jobs for the best people.” Anthropic’s alignment-science team was growing about 3× annually, FAR AI roughly 2×, and MATS about 2× across its history; new safety organizations, incubators, grant programs, and venture funds continue to appear.
The contradiction is management overhead. Fast-growing, flat teams cannot hire every moderately reskilled applicant when managers already supervise 10–20 people; Kidd had heard of one Anthropic alignment-science manager with 18 reports. Each hire must contribute quickly, rise toward research leadership, or relieve rather than deepen the coordination burden.
Hiring managers therefore report both adequate money and severe scarcity: “We have the clear need, but people are not at our bar.” Prior research experience predicts MATS performance, making a bachelor’s program, PhD, or strong research team a sensible route for many applicants rather than a detour.
Kidd also sees a founder shortage. Building a credible nonprofit or safety company creates positions that existing teams lack capacity to absorb, but founders need research experience and trusted references. MATS tries to supply credibility through selective admission, senior mentorship, references, and published outputs.
13. Age and credentials are weak proxies for research potential
The median MATS fellow is 27, but the distribution is wide: the oldest recent participant was roughly 55–60, while 18 is the minimum because the program does not accept minors. About 20% are undergraduates or lack a bachelor’s degree, while approximately 15% already hold PhDs.
Experienced researchers bring accumulated scientific and organizational knowledge; younger candidates may have an offsetting advantage with tools that did not exist long enough to become institutional habits. Kidd’s criterion is current capability, not a single credential profile.
Prodigies such as Chris Olah and Neel Nanda can bypass normal sequencing; for comparable people, MATS’s job is to “get out of their way.” Everyone else can develop through academia, independent grants, companies, cold outreach, or structured programs—“don’t be limited by the opportunities you see on job boards.”
14. Selection rewards work samples more than encyclopedic knowledge
MATS uses conventional CV screening and CodeSignal tests for some streams, currently including tests that detect AI use. Kidd expects AI-inclusive assessments to grow, but they are harder to design and score; mentors can instead specify their own selection procedure.
Bespoke work tests are closest to the research itself. Neel Nanda’s stream, for example, might ask applicants to spend about 10 hours using any tools they want, find something interesting, and present the result. Even rejection can leave the candidate with a useful GitHub artifact.
Baseline safety knowledge still matters. Kidd recommended a BlueDot Impact course or equivalent so applicants understand concepts such as deceptive alignment and can identify open questions, though requirements vary sharply between a control stream and an interpretability project.
Labenz’s compression was “tangible product is king.” Kidd agreed that a strong paper, working artifact, or trusted reference outweighs breadth during selection. He nevertheless recommends a “shallow but broad” map of the literature for succeeding after admission—periodic cross-field deep dives plus focused alerts, rather than constantly checking X for every paper.
15. The funnel is narrow, but alumni conversion is high
The initial intake admitted roughly 4–5% in the previous program; among applicants who completed the mentor-specific stage, Kidd suggested focusing on an acceptance rate around 7%. That is selective, though less extreme than the roughly 2% he cited for Anthropic’s fellows program.
MATS’s initial research phase lasts three months. About 75% of accepted fellows continue into a six-month extension, sometimes lasting 12 months, where much of the follow-on research and stronger career signaling occurs.
Across 446 fellows, MATS’s latest statistics put 80% into permanent AI-safety work and 98% into employment of some kind. The 80% includes independent researchers supported by grants, not only conventional W-2 employees, which Kidd regarded as a legitimate professional outcome.
Demand for entry is outrunning field deployment. One estimate put AI-safety headcount growth around 25% annually, versus 1.4–1.7× annual growth in MATS applications and about 1.5× in mentor applications; BlueDot reported application growth near 370%, a pace Kidd said obviously cannot persist indefinitely.
16. Safety salaries are competitive, while compute is rarely the constraint
Kidd sees no salary “safety tax” inside frontier labs. A few years earlier, someone joining “off the street” might receive about $370,000; he guessed junior software engineers could now see around $350,000 and mid-level or senior safety staff more than $1 million, while stressing that he lacked private compensation data.
The meaningful discount is a nonprofit tax. Kidd cited FAR AI research compensation around $170,000, possibly since increased, while believing well-funded organizations such as METR might offer upwards of $300,000 for most roles and perhaps over $1 million for some. Policies vary widely because nonprofits lack frontier-lab equity.
MATS offers a $12,000 compute budget, but it is not an unrestricted card: projects must justify the expense, and most fellows use far less. Rare larger requests can be considered and funded through reallocation; in practice Kidd said “basically no one” is compute-constrained.
MATS supplies organizational API accounts, its own cluster, and self-service providers such as RunPod, choosing between centralized support and flexible experimentation. Demand for Thinking Machines API access was growing. The program also planned summer, fall, and winter cohorts and was considering a one- to two-year residency for senior researchers.
17. The field needs new bets without abandoning its strongest agendas
Labenz worried that funders prefer legible organizations and established agendas even though researchers say the field lacks important ideas. He contrasted control’s grim premise—working productively with systems “out to get us”—with more hopeful, unusual approaches such as AE Studio’s self-other overlap research.
Kidd agreed that “more shots on goal are good” and defended a broad portfolio; MATS alumni helped originate or lead some of those neuroscience-inspired projects. He also credited Coefficient Giving with recently supporting more novel bets and moonshot interpretability or agency work.
His pushback was against converting strong iterators into reluctant visionaries. Asking someone demonstrably good at advancing a central agenda to abandon it for paradigm invention would be “strictly counterproductive on the margin.” Connector work often emerges from deep domain expertise, PhD-level research, or years embedded in communities such as MIRI—not generic brainstorming.
MATS can incubate connectors by pairing theoretical models with empirical execution; Kidd cited alumni and mentors who developed work on deception evaluations, gradient routing, activation engineering, and steering. Yet his closing allocation rule remained firm: diversify because any agenda might fail, but “the central ideas are still our actual best bets.”