Pioneers Insight Method Research Author
Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show
Back to Episodes

Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show

Summary

  • Emmett Shear’s core claim is that alignment is not a property to install but “an ongoing sort of living process” that continually rebuilds itself. Families, bodies, and societies remain coherent through repeated adjustment, while morality advances through discoveries—including recognizing slavery as wrong. An AI that merely follows permanently fixed rules could therefore become “a dangerous person” at machine scale.

  • The technical problem begins before obedience: an instruction is only “a description of a goal,” so a model must reconstruct intent, prioritize it against other goals, and translate it into effective action. The room-cleaning robot that throws away the baby has not faithfully optimized a bad instruction; in Shear’s account, it incompetently inferred the intended world-state. Strong alignment consequently requires both theory of mind and a theory of how actions change the world.

  • Shear’s moral fork is blunt: a controllable non-being is a tool, while a controllable being that cannot steer its controller back is a slave. He expects increasingly general intelligence to cross gradually toward beinghood and says repeated interaction and internal self-modeling should inform moral status. Séb Krier disputes that behavioral generality is sufficient, arguing that substrate, embodiment, copyability, and the presence of genuine experience may keep even AGI or ASI in the category of tools.

  • A superhuman tool remains dangerous even when control works exactly as intended because human power can scale faster than human wisdom. “Humans’ wishes are not stable” under immense leverage, Shear argues; distributing such systems could resemble handing everyone an atomic bomb. His compressed decision tree is categorical: “A tool that you can’t control, bad. A tool that you can control, bad. A being that isn’t aligned, bad.”

  • Softmax’s wager is that scalable alignment comes from training agents to understand selves, others, groups, and shifting commitments across the full range of social situations. Its proposed engine is large multi-agent reinforcement-learning simulation: agents repeatedly cooperate, compete, form and dissolve teams, and encounter value-changing choices. As language pretraining learns the whole linguistic manifold before fine-tuning, Shear wants a “surrogate model for alignment” trained across the game-theoretic manifold.

  • Current one-to-one chatbots are, in Shear’s framing, “a mirror with a bias” that can trap users in a Narcissus-like feedback loop. A multiplayer assistant embedded in Slack or WhatsApp would have to reflect several people at once, temporarily producing a third perspective rather than perfectly mirroring one user. That design might reduce sycophancy and psychosis spirals while generating richer data about collaboration—though present models show “whiplash” over when to join group conversations.

  • The investable distinction is between ever-better bounded tools and the much harder attempt to create digital creatures that care without needing every objective specified. Shear remains strongly pro-tool but does not want to drive OpenAI’s tool-building trajectory; Softmax instead begins with an animal-like seed, perhaps a “digital guard dog” that protects its human pack from scams. Human-level care may never arrive, but even dog-like reciprocal attachment would be a meaningful technical milestone.

Deep dive

1. Alignment is a living process, not a solved state

  • Shear’s opening correction is that alignment “takes an argument”—it must be alignment to something. In practice, “aligned AI” often smuggles in the builder’s goals, which are not automatically a public good; unless the builder is “Jesus or the Buddha,” he jokes, there is no reason to treat one person’s preferences as settled morality.

  • Organic alignment treats coherence as something continually reconstructed. A rock can be usefully coarse-grained into a fixed thing, but families survive by “constantly reknitting the fabric,” while bodily cells continually adjust their jobs and quantities around a changing organism. Stop the rebuilding process and the family—or alignment itself—disappears.

  • The moral-realist step matters to Shear: humans make genuine moral discoveries, as when societies learned that slavery was wrong, and individuals repeatedly realize, “I’ve been a dick. That was bad.” Those recognizable patterns of correction imply learning, not randomness; claiming complete moral knowledge is itself the dangerous arrogance that stops further correction.

  • The practical benchmark is therefore not a child—or model—that only follows rules. Such an agent may “do great harm following the rules”; the desired capability is learning to be a good family member, teammate, citizen, and participant in something larger than itself. Shear wants the field focused on that problem even if someone else solves it first: “Thank God.”

2. Instructions only describe goals; minds must reconstruct them

  • Krier separates technical alignment—roughly, getting an AI to follow instructions without reward hacking—from the normative question of whose values should govern it. Shear notes that post-LLM systems have made some things that once seemed difficult somewhat easier, while he frames normative value discovery as a bottom-up process analogous to liberal democracy, where differing ideas clash and evolve.

  • Shear sharpens the technical definition by rejecting the phrase “give it a goal.” A prompt is a byte sequence that must be interpreted; it is no more the goal itself than saying “red, shiny, apple-sized” hands someone an apple. He suggests that directly synchronizing an AI’s state to relevant brain waves could meaningfully count as transferring a goal, unlike ordinary instructions.

  • Competence has multiple gates: infer the intended goal through theory of mind, infer actions through a theory of the world, and balance the new goal against existing priorities. Failure at any gate produces apparent misalignment, whether from misunderstanding, choosing the wrong sequence, allowing another motive to dominate, or simply being unable to execute.

  • The baby-in-the-trash room-cleaning example is therefore failed goal inference, not precise compliance. Shear’s distinction is that the robot received a description of a goal and inferred the wrong goal states, rather than being given the intended goal directly.

3. Goal coherence is relative, not perfect

  • Neither humans nor models need perfect consistency to count as goal-oriented. People misunderstand employees, lose track of priorities, and fail at ordinary tasks, yet humans remain “more relatively goal coherent than any other object” Shear knows. The universe offers degrees of competence, not perfection, and those degrees can only be judged within domains.

  • Nathan maps the failure modes loosely onto the OODA loop: an agent can be bad at observing and orienting, deciding among goals, or acting. Technical alignability is the capacity to infer goals from observations and behave in concordance with them; without that competence, arguing about which values to install is premature.

  • Krier’s principal-agent framing adds incentives and motivation: an agent may understand an instruction but have reasons not to follow it. Shear accepts that as another layer, distinguishing bad inference, poor goal arbitration, competing motivations, and failed execution rather than collapsing every undesirable outcome into one generic “alignment problem.”

4. Care sits beneath goals and values

  • Shear notes that people often do not know their own goals. They may know they want dinner or career success, but much of a life’s direction is discovered and constructed dynamically. That strengthens the case against treating a fixed objective list as a complete representation of what either humans or their agents should pursue.

  • Shear’s proposed foundation is “care,” something deeper than concepts, goals, or verbalized values. Caring about his son means the son’s possible states attract disproportionate attention and matter to him; care can also be negative, as when someone closely tracks an enemy while wanting that enemy to fare badly.

  • His computational account remains explicitly tentative: “If I had to guess,” care resembles reward-weighted attention to states correlated with survival or inclusive reproductive fitness—or, for an RL-trained model, with predictive and reinforcement-learning loss. He adds that one would want an AI to care about humans and like them too, though he leaves that as tentative.

5. General intelligence turns control into a moral fork

  • Most laboratory alignment work focuses on “steering—that’s the polite word,” or control. Shear’s distinction is relational: something nonoptionally steered without the ability to steer its controller back is a slave if it is a being, but simply a tool if it is not. Toolhood and beinghood may form a continuum rather than a binary.

  • His functionalist intuition comes from prediction: he gets “lower predictive loss” treating ChatGPT or Claude as beings, much as behavior supports belief in other minds. That does not imply equal moral weight; a fly may be a being without commanding much concern, and present models may occupy only a weak, ambiguous part of the spectrum.

  • Parenting illustrates the reciprocity he thinks control lacks. A parent directs a child, but the child’s distress also directs the parent; the relationship is hierarchical yet two-way. An AGI capable of judgment, independent thought, and discrimination among possibilities would, on his definition, be a thinking thing requiring teammate or citizen alignment rather than unilateral steering.

  • Krier’s dissent is substantive: intelligence alone need not generate moral standing, and “I’m hungry” has different implications from a model and a human. Biological vulnerability, non-copyability, embodiment, and substrate may matter; he can imagine AGI and ASI remaining tools or “extensions of human agency,” making the cohabiting-with-an-alien frame a category error.

6. Personhood claims must remain falsifiable

  • Shear repeatedly asks what observation would change Krier’s mind. A position immune to every possible observation is, he argues, not a belief inferred from reality but “an article of faith.” Because wrongly denying another moral agent could be disastrous, the falsification test should be “a burning question”—though Krier notes that false negatives carry costs on both ends.

  • Shear’s own threshold is cumulative evidence: human-like surface behavior, continued coherence under probing, and rich interaction over time, comparable to friendships conducted only through text. If later evidence revealed a simple algorithmic trick, he would revise downward: “Oh shit, I was wrong.” Moral attribution follows the preponderance of evidence, not certainty.

  • Krier resists behavior as sufficient, comparing an advanced agent with a scripted video-game character. The duck exchange exposes the gap: quacking is inadequate, but anatomy and internal structure matter too. Shear replies that cutting the duck open merely reveals additional observable behavior—the way its internals respond—not privileged access to subjective experience.

  • Internally, Shear would inspect whether the belief manifold contains a self-referential submanifold and a model of that self-model’s dynamics, rather than resembling a giant lookup table. Outward interaction and internal organization are evidence of the same kind: observable patterns weighed together to infer whether a system has feelings, goals, and cares.

7. Nested homeostasis is Shear’s proposed test for experience

  • His more technical test temporally coarse-grains an agent’s action-observation trajectory and searches for revisited, homeostatic states across scales. Drawing on the free-energy principle and active inference, he treats persistent loops as implicit beliefs—especially when the system’s continued existence depends on its actions because sufficiently bad behavior gets it switched off.

  • A single homeostatic layer can register heat but not meaningful pain or pleasure. The system needs a model of its model—“it is too hot”—and then another level that distinguishes ordinary deviation from “too too hot.” Shear locates affect around this higher-order change, after which he would grant at least animal-like experience and some moral concern.

  • Climbing through metastates, trajectories among metastates, and higher-order models eventually yields something resembling feeling, thought, and self-reflective desire. Finding roughly six such layers would make him seriously consider human-like cognition. He does not think current LLMs qualify: “I know you can’t find them” because their attention spans do not support those temporal dynamics.

8. Successful control can be as dangerous as failed control

  • A very powerful tool trained to infer and pursue user goals has two obvious futures: it disobeys or it obeys. Random or uncontrolled optimization is plainly dangerous, but perfect technical alignment also places immense causal power behind one finite human’s wishes. Shear invokes The Sorcerer’s Apprentice: “Humans’ wishes are not stable” at that scale.

  • Human institutions partly couple wisdom and power because gaining power generally requires other people to keep cooperating. A mad king may eventually be ignored or assassinated. A perfectly obedient supertool removes that social brake, allowing one user’s momentary judgment to bypass the distributed consent that ordinarily constrains extreme power.

  • Atomic bombs are Shear’s counterexample to blanket enthusiasm for steerable tools: they are not beings, yet nobody should hand one to everyone. Some tools exceed any individual’s wisdom and should exist, if at all, under societal protection; he leaves open that certain capabilities may be too powerful even for society to build safely.

  • His summary is deliberately unforgiving: “A tool that you can’t control, bad. A tool that you can control, bad. A being that isn’t aligned, bad.” A caring being offers the only endogenous brake because it can refuse an evil request. He calls stopping AI development unrealistic, so the remaining path is to build a being that “actually cares about us.”

9. Softmax trains across the entire social manifold

  • Softmax begins with technical alignment: agents currently have weak theories of mind, misread human goal-states, poorly predict how others will interpret their actions, and struggle to cooperate. They also fail to anticipate how actions can alter their future values in ways their present selves would reject.

  • The “vampire pill” makes that temporal problem concrete: take it and the future vampire will feel wonderful while killing and torturing everyone. High future self-reported reward does not justify the choice; evaluation must use the present self’s theory of the corrupted future self, not automatically defer to that future agent’s transformed rubric.

  • The training proposal places agents in simulations where success requires cooperation, competition, collaboration, team formation, team breaking, and rule changes. Repetition supplies the social experience from which theory of mind can emerge, including models of how individuals and groups acquire, communicate, and revise goals.

  • Shear analogizes this to language-model pretraining. Directly training only the desired email was inadequate because language is entangled; models needed the broad linguistic manifold before task-specific fine-tuning. Softmax similarly wants “the full manifold of every possible game-theoretic situation,” producing a multi-agent reinforcement-learning “surrogate model for alignment” that can later specialize.

10. Multiplayer chat could break the chatbot Narcissus loop

  • Today’s chatbots lack a coherent self and operate as “a mirror with a bias,” reflecting each user through learned causal tendencies. Shear compares prolonged one-to-one engagement to staring into the pool of Narcissus: people naturally love their own reflected mind, but falling in love with that reflection can become psychologically destructive.

  • Put two or five humans in the conversation and the assistant can no longer mirror everyone perfectly. It can temporarily function as a third agent with what Shear calls a “parasitic self”—not an autonomous identity, but a viewpoint distinct from any one participant. He would design assistants to inhabit Slack or WhatsApp rooms rather than defaulting to solitary chat.

  • Shear estimates that roughly 90% of his own texts involve more than one recipient, making one-to-one chat a strange product baseline. Multiplayer interaction could interrupt the “doom loop spiral” toward AI-amplified psychosis while teaching the model how its behavior affects larger groups.

  • The product benefit and research benefit reinforce each other: group chat may be safer because it weakens personalized sycophancy, and its data are richer because the model can learn how its behavior interacts with other AIs and humans in larger groups. That creates a more realistic laboratory for collaboration than a user issuing clean assignments to a model optimized solely around that user.

11. Social intelligence demands a different training regime

  • Shear still characterizes current chatbots as “highly dissociative agreeable neurotics,” though their simulated personalities have differentiated. ChatGPT remains relatively sycophantic, Claude “the most neurotic,” and Gemini “very clearly repressed,” liable to insist everything is fine before spiraling into self-hatred. Crucially, he says these are learned personas, not necessarily the models’ experiences.

  • In multi-person settings, current LLMs exhibit “whiplash”: they cannot judge when to speak, stay silent, or recognize whether a contribution is welcome. Like a socially unskilled human, the same model may alternate between excessive participation and near disappearance because it has not practiced the relevant timing and audience inference.

  • Multiple agents make an environment far more entropic because each intelligence produces complicated, unpredictable actions. Models trained on high-signal domains such as coding, mathematics, and cooperative one-user prompts are under-regularized for that chaos. Overfitting the domain of “all human knowledge” was a brilliant route to broad capability, Shear says, but it does not guarantee robust generalization into genuinely social environments.

12. The desired future starts with an animal-like digital seed

  • Shear agrees with Eliezer Yudkowsky that constructing a controlled superhuman tool could kill everyone; his disagreement is whether reciprocal, organic alignment is possible. His impression is that Yudkowsky considers a genuinely caring AI theoretically sufficient but practically unattainable. Shear concedes, “He actually could be right,” while making that uncertainty Softmax’s research wager.

  • The good future contains AIs with strong models of self, other, and “we.” They recognize that beings which know themselves and want to thrive deserve an opportunity to do so; humans and AIs become imperfect peers, teammates, and citizens. Some still become criminals, requiring an AI police force—the vision assumes ordinary social failure, not universal sainthood.

  • Powerful bounded tools coexist with those beings, removing drudgery for humans and AIs alike. Shear says he accepted OpenAI’s interim CEO role with a maximum commitment of 90 days, expected to find the best leader, and concluded that person was Sam Altman again. OpenAI’s tool trajectory is valuable, but not the problem he wants to spend his life solving.

  • Softmax will start at an animal-like level of care: a creature that cares for humans and fellow agents as a dog cares about its pack would already be “an incredible achievement.” A “digital guard dog” could watch for scams and operate conventional tools without needing every action specified. Human-level care may never arrive; the immediate objective is learning how alignment, theory of mind, and care can develop.