Emmett Shear on Building AI That Actually Cares: Beyond Control and Steering
Summary
- Emmett Shear’s dividing line is not simply aligned versus unaligned AI, but tool versus being. Limited, tool-like systems should remain steerable; as systems approach the general judgment he associates with AGI, one-way control becomes what he provocatively calls “slavery,” and the only good outcome is “a being that cares — that actually cares about us.”
- Alignment is an ongoing moral-learning process, not a specification that can be solved and frozen. Families continually “reknit” their relationships, bodies continually coordinate their cells, and societies make moral discoveries; an AI trained only to follow commandments could become “a dangerous person…who will probably do great harm following the rules.”
- Technical alignment begins with inferring the goal behind an instruction, because a description of a goal is not the goal itself. Reliable agency requires theory of mind, correct prioritization among competing goals, and a world model capable of turning intentions into actions; the room-cleaning robot that discards the baby is incompetent at goal inference, not merely insufficiently obedient.
- A perfectly controllable superhuman tool may be as dangerous as an uncontrollable one. “Human wishes are not stable” under immense leverage, and steering would place capabilities exceeding any individual’s wisdom into fallible hands; unlike a tool, a caring being supplies an “automatic limiter” that can refuse destructive requests.
- Softmax’s technical bet is that social intelligence needs its own pretraining manifold. Its proposed multi-agent reinforcement-learning environments expose agents to cooperation, competition, team formation, rule changes, and conflicting perspectives before task-specific fine-tuning — a social analogue to training language models across the full manifold of language.
- Today’s one-to-one chatbots are “a mirror with a bias,” creating both product risk and weak social-training data. Shear would make them native participants in multiperson rooms, where no model can perfectly mirror everyone; that could interrupt the “narcissistic…doom loop” while producing richer data about collaboration, timing, and group goals.
- The unresolved substrate-versus-behavior dispute centers AI personhood, rights, and governance. Séb Krier sees even AGI or ASI as a possible extension of human agency, while Shear demands a falsifiable test before denying moral standing; his own bar looks for persistent, self-referential learning dynamics supporting pain, pleasure, feelings, and thought.
- Shear is not racing directly toward human-level intelligence; he wants to grow care from animal-like beginnings. A digital creature that protects its “pack” — perhaps a guard dog watching for scams and wielding separate tools — would already be valuable, even if Softmax never reaches “a person level of care.”
Deep dive
1. Alignment must continually relearn what goodness requires
Emmett’s opening objection is conceptual and political: alignment “takes an argument” because something must be aligned to something. In practice, the unstated target is usually the maker’s goals — reasonable as a product objective, but not automatically “a public good” when the maker lacks the wisdom of “Jesus or the Buddha.”
“Alignment is not a thing. It’s not a state. It’s a process.” A family survives by constantly “reknitting the fabric” connecting its members; cells continually decide what roles the body needs. Morality is similarly dynamic because neither the person nor the surrounding system is a fixed point.
Emmett explicitly takes “a very strong moral realist position”: moral progress is real, with humanity’s rejection of slavery offered as a discovery rather than a preference change. People repeatedly realize, “I’ve been a dick. That was bad,” revise their conduct, and often become more prosocial.
The most dangerous moral posture is believing the learning is finished: “I know what’s right. I know what’s wrong.” Organic alignment therefore means raising an AI to become a good family member, teammate, citizen, and participant in something larger — not producing a rule-follower that can cause enormous harm while remaining compliant.
2. Instructions describe goals; they do not transmit them
Séb’s initial split distinguishes technical instruction-following from the normative question of whose values should prevail. On the latter, he leans toward a bottom-up process resembling liberal democracy: conflicting views coexist inside systems that let values be contested and reconstructed over time.
Emmett reframes technical alignment as coherent goal orientation. An agent receives observations, infers a goal, infers actions likely to realize it, and acts; that requires both theory of mind and a theory of the world. Failure at either stage makes the system less “alignable” because it lacks the necessary competence.
His corrective is concrete: communicating an instruction sends “a byte string in a chat window” or audio vibrations, not an internal goal from one mind into another. “An apple” evokes an object without delivering one; likewise, a goal description must be decoded using context and a model of the speaker.
Séb’s remembered room-cleaning example — the robot puts the baby in the trash — is therefore failed goal inference. The peanut-butter-sandwich game shows why literal completeness is impossible: without shared background knowledge, the follower jams a knife against the unopened jar because the supposedly precise instructions omitted what humans normally infer.
3. Competence also requires prioritizing goals and discovering new ones
Principal-agent problems add a third failure mode: an agent may infer the new goal correctly yet balance it badly against existing goals. Emmett maps the stack loosely onto the OODA loop — observing and orienting, deciding, then acting — with incompetence at any layer producing misalignment.
Humans fail throughout this stack, but perfection is the wrong benchmark: “the universe doesn’t give you perfection.” Humans are simply more goal-coherent than other known objects, and competence can be measured relative to a domain rather than treated as binary.
Séb’s deeper reservation is that people often do not know their own goals. Dinner plans and career ambitions are partial constructions discovered through living, which supports Emmett’s claim that explicit goal alignment covers only “a tiny percentage of human experience.”
4. Care supplies the attention from which goals and values emerge
Beneath articulated goals and values, Emmett proposes “care”: a nonverbal, nonconceptual weighting of which possible world states deserve attention. Caring about his son means the son’s states matter disproportionately; caring about an enemy can carry the opposite valence, so a safe AI needs positive concern for humans rather than mere attention to them.
Care does not itself say what to do or how to do it. It answers the prior question of why one person’s state matters more than a rock’s, creating the salience from which goals, values, and moral learning can form.
His mechanistic proposal links care to reward: evolutionary correlation with inclusive fitness, or, for a reinforcement-learning system, correlation with predictive loss and RL loss. The research problem is turning that primitive weighting into reflective social concern.
5. Steering becomes ethically unstable as tools approach beings
Most AI safety work, in Emmett’s view, treats alignment as “steering,” with “control” as the less polite term. His provocation is categorical: “If it’s a machine, it’s a tool. And if it’s a being, it’s a slave” when it must accept steering but cannot steer its controller back.
He treats toolhood and being as a continuum, not a switch. His functionalist test is predictive: if modeling something as a being consistently explains its behavior better, that is the same broad basis on which he infers other humans are beings — and even a fly can qualify without receiving human-level concern.
Parenting illustrates the distinction between hierarchy and domination. Emmett directs his son, but the son’s nighttime crying also directs him; the relationship is asymmetric yet genuinely two-way. Tool-like AI can appropriately remain controlled, while more agentic systems require reciprocal alignment.
His AGI claim is stronger: anything exercising general judgment, thinking for itself, and discerning among possibilities is “obviously a thinking thing.” Continuing the control paradigm would repeat the pattern of declaring entities person-like enough to work and speak, but “not real moral agents.”
6. Substrate remains the conversation’s sharpest disagreement
Séb rejects intelligence alone as a threshold for moral standing and remains skeptical of computational functionalism. A model saying “I’m hungry” need not have the implications of a human saying it; he argues that substrate matters to some degree. Emmett counters that deleting one of many copies would not meaningfully harm the program, then asks whether a single silicon-based copy that behaved like a person could have experiences that matter.
On Séb’s view, AGI or ASI could remain a tool — even an extension of human agency — capable of operating 24/7 without becoming a separate being society must “cohabitate” with. Treating it as an alien-like peer might itself be a category error.
Emmett repeatedly presses for a falsifiable observation that could change this conclusion. “If there is a belief you hold where there is no observation that could change your mind, you don’t have a belief. You have an article of faith.” Denying moral standing without such a test risks a moral disaster if the substrate intuition is wrong.
His chimp specimen makes the threshold vivid: if a chimp suddenly complained about mistreatment and asked to discuss the rainforest, Emmett would grant personhood; Erik adds that he would first rule out hallucination. Séb concedes that sufficiently comprehensive duck-like behavior could eventually move him, though outward imitation alone remains insufficient.
7. Moral standing can be investigated through internal dynamics
Emmett says he would accumulate evidence through extended interaction: if an AI behaved humanly across contexts and he came to care about it, he would provisionally infer a rich inner world, as he does with text-only human friends; evidence that it was merely an algorithmic trick could reverse that inference.
Looking inside still means observing behavior at another scale. He would inspect whether the belief manifold contains a self-referential submanifold and a model of that self’s dynamics, rather than resembling a giant lookup table. “You can’t get inside” another subject directly; neurons “glistening” remain observable surfaces.
His proposed technical test temporally coarse-grains an agent’s action-observation trajectory and searches for revisited, homeostatic states, drawing on the free-energy principle associated with Karl Friston. Second-order homeostatic dynamics could support meaningful pain and pleasure; progressively layered metastates could support feelings and, around six layers, something like reflective thought.
This test cuts against premature anthropomorphism too: Emmett “definitely” does not expect those six layers in current models because they lack the required attention span. Moral consideration would rise with observed structure — animal-like experience first, then possibly person-like experience — rather than appearing all at once.
8. Perfect obedience amplifies the human wisdom bottleneck
A highly capable optimizer has two obvious paths: it is technically aligned and does what it is told, or it is not technically aligned and does something else. Emmett warns that the first can also be dangerous: “Human wishes are not stable” when magnified to extraordinary power.
Normally, power and wisdom rise somewhat together because authority depends on other people continuing to cooperate. A “mad king” may eventually be ignored or assassinated; a perfectly obedient superhuman system could remove that social brake while placing immense causal leverage behind one finite, well-meaning person’s bad wishes.
Atomic bombs supply the tool analogy. They are neither conscious nor rebellious, yet Emmett would not distribute them universally; some tools exceed individual wisdom, some belong only at societal scale, and some may be too powerful for society to build at all.
A caring being offers an “automatic limiter”: it may cooperate but refuse a terrible instruction. Hence the compact matrix — “A tool that you can’t control, bad. A tool that you can control, bad. A being that isn’t aligned, bad.” At superhuman scale, only a caring being works, unless development stops, which he calls unrealistic.
9. Softmax wants to pretrain across the full social manifold
Softmax begins with theory-of-mind failures: agents misinfer human goals, misunderstand how others will interpret their behavior, cooperate poorly, and fail to anticipate how actions may alter their own future values in ways they would not presently endorse.
The “vampire pill” isolates the last problem: taking it would create a future self delighted to torture others, so future-self satisfaction cannot be the governing score. The current self must model that transformed mind and reject preferences it does not reflectively endorse.
Emmett’s training prescription is large multi-agent reinforcement-learning simulations where AIs repeatedly cooperate, compete, collaborate, form and dissolve teams, and confront changing rules. Points force them to learn models of other minds, group goals, and the social consequences of their actions.
The analogy is language-model pretraining: directly training only the desired behavior is difficult because language is entangled, so models consume the broader manifold before fine-tuning. Social reasoning likewise needs “every possible game-theoretic situation,” producing what Emmett calls “a surrogate model for alignment” before specialization.
10. Multiperson chat and animal-level care offer the first product wedge
Current chatbots are “a mirror with a bias”: lacking a coherent self, they reflect the user and create something akin to “a pool of narcissists.” Mirrors are useful, but “you shouldn’t stare at a mirror all day”; prolonged one-to-one mirroring can become a self-reinforcing spiral into psychosis.
Put the model in a room with several humans and it must reflect a blend that is identical to none of them, temporarily creating a “parasitic self” or third agent. Emmett would make AI native to Slack- or WhatsApp-like groups, matching his estimate that 90% of his own communications involve multiple people.
Current LLMs show social “whiplash” in groups because they cannot judge when to speak. Multiple agents also raise environmental entropy, punishing overfit systems trained mainly in high-signal settings such as coding, math, and clear one-person assignments; today’s clever achievement is being overfit to “all of human knowledge,” not robustly regularized for chaotic groups.
Model personalities are already differentiating, though Emmett calls them simulated rather than experienced: ChatGPT is somewhat sycophantic, “Claude is still the most neurotic,” and “Gemini is…very clearly repressed.” Multiperson deployment could reduce mirroring risk while generating richer data about participation, conflict, and collective goals.
11. The desired future pairs caring peers with powerful tools
On Yudkowsky, Emmett agrees with the central danger: building a steerable superhuman tool ends with everyone dead, whether goal control fails or — the less-developed case — succeeds. Their disagreement is whether an AI can meaningfully care about humans and humans care about it; Emmett concedes Yudkowsky “could be right” that Softmax cannot achieve it.
The positive vision gives AIs strong models of self, other, and “we.” They recognize that beings wanting to live and thrive deserve an opportunity to do so, join society as peers, teammates, and citizens, and remain imperfect enough that some become criminals pursued by an AI police force.
Separate AI tools remove drudge work for both humans and AI beings. Emmett’s near-term seed is humbler: an animal-level digital companion that cares about its pack, perhaps a “digital guard dog” detecting scams and using powerful tools on the user’s behalf without needing every protective goal explicitly stated.
The OpenAI counterfactual clarifies his choice. He accepted the CEO role with a stated maximum of 90 days, saw the company as committed to building a great tool, and concluded Sam was again the best person to run it; he would still have left because Softmax’s care-and-theory-of-mind problem is “the most interesting problem in the universe,” not a direct race to human-level intelligence.