Pioneers Insight Method Research Author
Karl Friston - Why Intelligence Can't Get Too Large (Goldilocks principle)
Back to Episodes

Karl Friston - Why Intelligence Can't Get Too Large (Goldilocks principle)

Summary

  • Friston’s core call is that intelligence occupies a Goldilocks zone of scale: it can disappear when a system becomes either too small or too large. At tiny scales, randomness overwhelms the recurrent structure needed for agency; at planetary or astronomical scales, averaging erases constituent complexity, leaving motion without planning. An intergalactic civilization may therefore need federation rather than a single intelligence: evolution “doesn’t have to be intelligent to be beautiful.”

  • Bigger organizations do not become smarter merely by aggregating smart parts. Corporations can preserve large-scale intelligence only by imposing structures that balance dissipation with recurrent, conservative dynamics—the organizational equivalent of operating “on the edge of chaos.” Friston’s sharper institutional warning is that globalization and organizations may become “too big for their own good” because they lose the ability to plan.

  • Machine consciousness is possible in principle, but Friston doubts standard von Neumann architecture can deliver it efficiently—or perhaps at all. A conscious artifact would need agency, substantial counterfactual breadth and temporal depth, embodiment as a “beast machine,” and substrate-dependent or “mortal” computation. His directional hardware bet favors processing-in-memory, memristors and neuromorphic photonics—not necessarily spiking neural networks—because the memory substrate must participate and self-organize.

  • Agency and consciousness are orthogonal, although recursive causal structure is used to explain agency and some consciousness proposals. Friston locates agency in a system’s inability to directly observe its own actions, so it infers them as causes of sensation and acquires a “future on the inside” from which it selects paths. Consciousness may require additional precision control, attention and metacognition—“looking at the looking at the looking”—rather than merely having Bayesian beliefs.

  • After roughly 20 years, Friston says he has not encountered the decisive “that doesn’t work then” moment for the free energy principle. Its core is almost minimal: partitions create conditional probability distributions, whose dynamics follow a principle of least action. The difficulty is communicating something “almost tautologically simple” whose consequences are profound; Friston remains ambivalent because a little “magic and mysticism” also motivates inquiry.

  • The episode rejects both brain chauvinism and indiscriminate pan-intelligence. Mike Levin’s basal-cognition program treats intelligence in viruses, slime molds and xenobots as an empirical question, while Keith argues that a virus is essentially a “mousetrap” outsourcing its machinery to a host and doing no lifetime learning. The useful unit may instead be a colony or ecosystem—but Friston still requires the right causal structure, counterfactual depth and scale.

  • For autonomous systems, static perception is insufficient because meaningful boundaries are histories, not just pixels. Markov-blanket discovery must identify entities that persist and recur over time; image segmentation alone cannot supply understanding. That makes structural learning and dynamic world models central to robotics and autonomous vehicles, while Friston concedes that action at a distance still leaves him without “good maths” for choosing between competing causal structures.

Deep dive

1. Twenty years have strengthened, not settled, Friston’s conviction

  • Keith dates the free energy principle to around 2005 and asks for a twenty-year audit. Friston resists a victory lap: assessing it while developing it is like reviewing a meal while cooking, but every phenomenon that “slip[s] in quite neatly” produces a dopamine hit—and he has yet to reach the decisive “that doesn’t work then” moment.

  • Friston’s disappointment is communicative. He is ambivalent when the principle is called notoriously difficult: some “magic and mysticism” gives people a challenge worth mastering, yet the framework is meant to be “almost tautologically simple,” and he suspects his own explanations have not kept it sufficiently intuitive.

  • Keith compares it with probability theory: the sum and product rules are easy to write down, yet their consequences are profound and routinely misapplied. Friston agrees because conditional probability is the framework’s heart: partition states, condition one set upon another, then describe a principle of least action for the dynamics of those conditional densities—“before thermodynamics” and “before quantum mechanics.”

  • Friston sees generative AI and large language models primarily as engineering resources for understanding natural intelligence. That understanding could make life better, support sustainability and serve mental wellbeing through computational psychiatry; invoking Feynman—“That which I cannot create, I do not understand”—he adds the uncomfortable implication that genuine understanding may require creating “things that suffer.”

2. Markov blankets turn missing causal links into natural kinds

  • Friston begins with external, internal, sensory and active states. Different forbidden causal influences produce an ontology of “natural kinds”: an entity with neither internal nor active states is a causal black hole, quintessentially inert and effectively invisible because it never acts back upon its environment.

  • In an ordinary touch-bound system such as a cell, sensory states form the outside layer, active states sit beneath them, and internal states occupy the center. The arrangement enforces both directions of conditional independence: external states cannot directly reach action, while internal states cannot directly cause sensation.

  • Hierarchical organisms are different because their active states become sequestered from internal states. From the inside, unseen actions now appear as causes of sensation alongside the external world; the organism must therefore infer not only its environment but itself. That recursion is what makes these systems “strange things.”

  • Action at a distance complicates the picture. A bacterium inhabits a touch-local world, but vision, hearing, electromagnetic radiation and magnetic fields—as illustrated by the compass discussed—create nonlocal causal possibilities. Friston’s candid gap is structural learning: he knows of no good mathematics for deciding how best to disambiguate between competing causal organizations.

3. Self-modeling gives an agent a private future

  • Keith’s framing is that computational strange loops unlock an entirely different category of behavior. Friston agrees: once a system models itself as a cause of its environment, it must infer what it is doing and what those actions will produce; that is an “authentic kind of agency,” not merely reactive regulation.

  • Because an action’s consequences have not yet arrived, self-modeling is intrinsically future-pointing. The transition from thermostat, virus or single cell to agent occurs when a system has a “future on the inside,” represents divergent possible paths, and selects among them—arguably necessary, Friston says, for true agency and possibly elemental consciousness, though not a definition of either.

  • Tim presses the apparent conflation of agency with phenomenal experience: deeper reflexive layers plausibly strengthen agency, but why should they create consciousness? Friston concedes the point completely, aligning with Anil Seth’s distinction: intelligence, agency and consciousness are orthogonal, and “consciousness is not implied by being agentic.”

4. Consciousness may arise from precision, access and recursive attention

  • The free-energy treatment offers dual-aspect monism rather than a direct solution to consciousness. The same brain substrate carries a thermodynamic information geometry and a representational geometry in which internal states encode Bayesian beliefs about the world, including action; physical dynamics and inference are two readings of one process.

  • Merely encoding a posterior belief may not suffice. A thermostat can be interpreted as holding beliefs about temperature without thereby being conscious, so stronger accounts require beliefs to become dynamically “ignited”—sufficiently precise to affect updating elsewhere in a hierarchical or distributed system.

  • Friston invokes a “felt uncertainty” account associated in his remarks with Mark Solms: consciousness tracks representations of confidence, precision or uncertainty, not simply content. He cautiously recalls the claim that consciousness can be switched off through a brain-stem region roughly three millimeters cubed, containing the origins of ascending systems that encode and represent those precision signals.

  • More demanding theories add control and self-recognition: attend, recognize that you are attending, then recognize that recognition. Chris Fields’s inner-screen hypothesis recasts nested Markov blankets as holographic screens and posits an irreducible inner screen that knows itself only by acting on the rest of the brain—compatible, Friston thinks, with higher-order thought and precision-gated global neuronal workspace theories.

5. Machine consciousness requires long horizons and mortal hardware

  • Keith argues for a minimum complexity and causal structure, comparing consciousness with a Game of Life glider that cannot exist below some number of pixels. Friston agrees, while treating the boundary as technically vague—like asking how many grains make a pile—and nominates temporal depth as the most plausible dimension.

  • Counterfactual breadth is the number of available future paths; counterfactual or temporal depth is how far those paths extend. A thermostat has only the instantaneous future implicit in a differential equation. A machine with human-like agency would need a world model of its actions’ consequences with both wide alternatives and a substantially longer horizon.

  • Friston’s answer to whether machines could eventually understand or become conscious is “in principle, yes,” but mathematical simulation alone may not suffice. Following Seth’s “beast machine” framing and Geoffrey Hinton’s mortal computation, he argues for embodiment and substrate dependence: the computation must depend on its physical substrate rather than merely being implemented independently of it.

  • He therefore doubts von Neumann systems, where processors read and write separate memory, because that blanket structure makes memory difficult to self-organize. Friston says a conscious machine would have to pursue a path of least action thermodynamically and informationally; otherwise it could not be conscious for a nontrivial amount of time. He points instead toward processing-in-memory, memristors and neuromorphic photonics, while explicitly saying spiking neural networks are not required.

6. Basal cognition is empirical, but viruses expose the boundary problem

  • Tim asks whether plants and viruses imply “pan-intelligence,” challenging the privileged status given to brains. Friston directs the question toward Mike Levin and Chris Fields’s basal-cognition program: intelligence may be widespread, but experiments must be designed to reveal it rather than assumed from appearance.

  • By many adaptive-behavior yardsticks, Friston says, slime molds, viruses and xenobots can look “incredibly intelligent.” He is sympathetic to that attack on chauvinism, yet insists that a virus lacks the expressive agency of a strange thing; some causal structures are categorically present or absent, even if the broader intelligence threshold remains vague.

  • Keith’s pushback is biological outsourcing. A virus is mutated by radiation or transcription error rather than modifying itself, behaves like a “mousetrap” that injects genetic material, and leaves transcription and the rest of the work to the host cell. Over its own lifetime it neither learns nor adapts, making “intelligence” less useful than ordinary unfolding dynamics.

  • Scale partly reconciles them: the individual virus may be uninteresting while a colony or ecosystem becomes the relevant unit. Friston similarly calls natural selection Bayesian model selection—favoring phenotypes with the greatest model evidence—without equating adaptiveness with planning. Keith’s stricter rule assigns intelligence to the smallest sufficient boundary, not everything containing it.

7. Intelligence vanishes below and above a Goldilocks scale

  • Friston proposes evaluating each candidate at the scale of its Markov blanket: does it contain enough machinery for a world model with meaningful counterfactual depth and breadth? A virus may be too small; a biosphere may be too large because coarse observation averages away the rich causal organization of its constituents.

  • The physical mechanism begins with the Helmholtz decomposition of dynamics into dissipative and conservative parts. Dissipative fluctuations supply randomness and openness; conservative or solenoidal flows cycle through states, supporting routines, biorhythms, reproduction and Poincaré recurrence—the repeated return near a starting point that lets a system maintain a recognizable identity.

  • At very small scales, Friston argues, quantum randomness dominates and recurrent conservative organization becomes insufficient. Human-scale organisms mix an itinerant, changing world with stable revisitation of characteristic states. At astronomical scales, averaging removes the fluctuations, leaving predominantly conservative Newtonian motion: planets orbit, but the motion contains no evidence of planning.

  • This is the Goldilocks regime “on the edge of chaos,” where neither complete order nor complete noise wins. Friston says he does not see the Moon, weather or evolution thinking about their futures. When extending the argument toward an ultimate scale limit, however, he hedges: “I’m not sure this is correct.”

8. Coarse-graining creates new intelligence without inventing new matter

  • Tim asks whether the apparent unintelligence of Gaia merely reflects observer limitations. Humans abstract by ignoring detail and idealize by distorting it; if intelligence is “doing more with less” and emergence means “more is different,” then a higher-scale system may be genuinely reorganized rather than merely a blurry view of its components.

  • Friston’s mathematical answer is the renormalization group: an operator reduces dimensions and groups fine-scale states into the appropriate higher-scale variables, recursively. Higher levels can display brand-new self-evidencing or intelligent dynamics, yet those variables remain functions of finer-scale activity—real emergence without adding an unexplained substance.

  • Keith argues—and Friston agrees—that large-scale intelligence such as a corporation requires deliberately installed organizational structure. The corporation’s coarse metric is survival—whether it progresses from “Series A to Series whatever” in recognizable form—but increasing size leaves less room for rich recursive organization, making planning progressively harder.

  • Keith extends the limit to intergalactic civilization, where the speed of light eventually blocks centralized coordination across light-years. Friston says the physics arguments suggest a Goldilocks zone, then reframes the disappointment: federated ecosystems can still be beautiful. “Evolution is a beautiful thing,” and “it doesn’t have to be intelligent to be beautiful.”

9. Morphology embodies the model while DNA constrains the policy

  • Tim’s extended-cognition challenge asks where thinking ends: inside the head, in a phone, across a plant’s morphology, or throughout a bidirectional causal process? Friston answers that discussion requires a Markov boundary, but every bounded thing is contextualized by the scale above; even an unintelligent virus requires a host world conducive to its persistence.

  • For plants, morphology is not merely what an internal model represents. Physical structure parameterizes conditional distributions, making “the substrate” part of the generative model itself. Under the good-regulator idea, an organism must embody the causal organization of its environment; a scale-free world should therefore be matched by scale-free hierarchy in the organism.

  • Friston’s best specimen is David Attenborough’s plant footage accelerated by a factor of 10 or 100. Roots and shoots move in particular directions, plants compete with other plants for sunlight and some consume insects; sped into a human temporal register, their behavior looks strikingly animal-like and difficult to dismiss as unintelligent, irrespective of consciousness.

  • Tim initially raises DNA as a possible software-like model, then Keith sharpens the point: DNA specifies a policy—what each cell does given sensory states—not a completed body plan. Friston agrees that inherited code supplies slowly changing structural priors and expected constraints, while each organism must learn particulars such as where to grow; the same code can differentiate into roots, bark and organs.

10. Practical intelligence discovery must follow dynamics through time

  • Tim relays Maxwell Ramstead’s compressed account: bounded open systems that cannot merge instead exchange information and synchronize. Friston endorses it—loosely coupled systems converge on a chaotic synchronization manifold, and free-energy minimization can be described as generalized synchrony without invoking Bayes, predictive processing or self-evidencing.

  • Chris Fields’s quantum version frames the counterpart in terms of entanglement and unitarity. Friston jokes that after “more than five minutes” together, the participants will become completely entangled: perhaps not literally one system, though Fields might say so, but increasingly indistinguishable as their dynamics occupy the same manifold.

  • For a camera-equipped robot, image segmentation is only a starting commitment that the environment contains things. Tim argues that understanding requires history, not a snapshot; Friston agrees because Markov-blanket discovery must use dynamics. With only one universe-state at a time, the repeated realizations supporting a probability distribution arise across past and future time.

  • Autonomous vehicles therefore need to conserve inferred blankets across time and distinguish persistent things from amorphous stuff such as water or fog. Object-centric Newtonian priors of the kind Josh Tenenbaum studies are one option; broader models are possible, but Friston’s boundary is memorable: “I would relax, but not beyond the renormalization group.”