Large Language Models and Emergence: A Complex Systems Perspective (Prof. David C. Krakauer)
Summary
Krakauer’s core test is whether a system can do more with less, not whether enormous data and compute can make it knowledgeable. He distinguishes capability, knowledge and intelligence, while treating “the capacity to acquire capacity” as a debated definition rather than a settled one. His recurring criterion is: “When it comes to intelligence, less is more, not more is more.” For investors, benchmark gains may validate compute demand without proving that today’s architectures have crossed into a more efficient form of reasoning.
A sudden benchmark jump is not emergence unless the system’s internal organization has changed. Krakauer contrasts an illustrative rise in three-digit addition from under 50% at 100 billion parameters to 80% at 175 billion with an HP-35 calculator doing the job in a 1K ROM—“an order of a billion times smaller in memory footprint.” His verdict is deliberately abrasive: “I would simply call that really shit programming.”
The investable architectural breakpoint would be a break in scaling accompanied by a new, more parsimonious internal representation. Genuine emergence screens off microscopic detail, as fluid dynamics replaces molecule tracking with densities and Navier–Stokes equations. The host reports studies Daniel Hendricks mentioned suggesting capabilities are “something like 96% correlated” with scale; Krakauer says that relationship must first break, then researchers must demonstrate the micro-to-macro reorganization behind it.
Evolution explains why storage, learning and intelligence should not be treated as the same asset. Manfred Eigen’s quasispecies theory imposes a speed limit of roughly “one bit per selective death,” or one bit per genome per generation, forcing large organisms to develop brains and other extra-genomic systems for faster information. Culture goes further: as long as discoveries are archived, it has no upper bound in this sense and “breaks evolutionary light speed”—but Krakauer still classifies that archive as knowledge, not intelligence.
Useful priors may improve model efficiency, but Krakauer rejects symmetry as a universal master key. He presents convolutional networks as an example of encoding real covariance in visual scenes, questioning the field’s “allergy” to world-informed priors. Yet life is historically contingent: “Darwin is in some sense anti-Noether”—change time or space and biological outcomes change—so geometric elegance cannot erase evolution’s broken symmetries.
Agency requires more than autonomous action: it adds a future-directed policy to an adaptive internal record. Krakauer’s ladder runs from physical action, to Darwinian adaptation through a schema, to agency that says, “This is what I want to do.” His account of collective intelligence also spans people and artifacts: maps, abacuses and language compress collective discoveries into low-dimensional structures that can reorganize an individual brain.
Krakauer’s biggest fear is that AI adoption degrades the human capacities it replaces. Maps can be internalized, while a guaranteed superior GPS removes the incentive to learn navigation; he expects the drive to increase fidelity and reduce effort to make humans “eventually outsource ourselves.” His condition for superintelligence is uncompromising: it matters only if “it makes me more intelligent,” not “more stupid, more servile, or more dependent.”
Deep dive
1. Intelligence begins where accumulated knowledge stops carrying the task
Krakauer’s preferred signal is “doing more with less.” He discusses—but does not settle on—the proposal that intelligence is the ability to acquire capability, or “the capacity to acquire capacity.” The host’s shorthand that “stupidity is doing less with more” is immediately corrected as Krakauer turns to his account of emergence. He also makes a sharp distinction between knowledge and intelligence.
Eigen’s quasispecies theory supplies the evolutionary constraint. Selection tests competing hypotheses by killing unsuccessful variants, but can preserve only about “one bit per selective death”—one bit per genome per generation—with generation time setting the speed limit.
Large multicellular organisms encounter adaptive information at frequencies higher than the generational frequency. Brains and epigenomes therefore act as extra-genomic inferential systems for acquiring high-frequency information; Krakauer calls the boundary requiring such a mechanism the “error threshold.”
Culture changes the process again: books, libraries and hard drives “refrigerate” prior discoveries, allowing variants to be generated at any rate so long as accumulated information is preserved and successful additions are stored. Culture can therefore “break evolutionary light speed,” though Krakauer insists storage remains knowledge rather than intelligence.
2. Benchmark discontinuities are not evidence of emergence
Krakauer grounds emergence in Phil Anderson’s 1972 “More Is Different.” As systems become larger, initial conditions can determine which broken-symmetry state they occupy, making clever averaging—or coarse-graining—necessary for a parsimonious macroscopic description.
For three-digit addition, Krakauer recalls illustrative figures—not exact ones—of performance below 50% at 100 billion parameters and roughly 80% at 175 billion. An HP-35 calculator achieves the same narrow result with a 1K ROM, making the LLM route look less like intelligence than “really shit programming.”
A real phase transition is characterized by altered internal organization, not merely a discontinuous score. Fluid dynamics is the model case: average densities and Navier–Stokes equations become sufficient, while tracking every molecule adds no useful predictive information.
Scaling laws themselves are “not evidence of emergence.” The host cites studies mentioned by Daniel Hendricks suggesting capabilities are “something like 96% correlated” with scale; Krakauer’s test is first a broken scaling law, then an internal explanation mapping microscopic dynamics onto the new macroscopic observable.
3. Better representations need world structure without pretending history disappears
The host proposes that current models learn fractured, entangled representations, whereas emergence would produce factored, unified abstractions that “carve the world up by the joints.” Krakauer agrees that nervous systems and bodies respect genuine physical constraints.
Convolutional networks show the value of a stronger prior: visual scenes contain covariance, and the architecture captures it naturally. Krakauer finds the neural-network community’s resistance to priors confused, since the field’s own foundations were inspired by nervous systems: “Why not use a bit of world inspiration too?”
He nevertheless resists a broadly Platonic symmetry program while admitting unfamiliarity with the specific geometric-deep-learning claim. Noether links physical symmetries to conserved quantities; “Darwin is in some sense anti-Noether,” because contingent history makes outcomes differ across time and space. Evolution can look convergent at coarse resolution; zooming in reveals its unique fossil record.
4. Emergence, causality and agency require distinct internal descriptions
Physical emergence is usually studied in systems of many identical components receiving one global signal, such as temperature. Biology and machine learning instead involve non-identical components that may be locally parameterized: this is “knowledge in,” unlike “knowledge out,” where changing one variable yields an unexpectedly new state. The challenge is claiming emergence after violating that classical setup.
Emergence can be restated as “a more parsimonious causal mechanism.” Coarse-grained observables may be genuinely causal in Pearl’s interventional sense, a complementary conception to fundamental Newtonian causality rather than the same one.
Krakauer’s hierarchy begins with physical action—a ball rolling downhill—then adaptation, where an internal schema maps situations to responses, and finally agency. A policy adds directedness: rather than merely reacting, the system says, “I would like to do this into the future.”
5. Minds extend through communication, artifacts and collectives
Endogenous coarse-graining is hard to establish in an isolated brain, because linguistic thought still recruits millions of neurons. Communication makes it visible: explaining integration by parts or a Fourier transform transmits a low-dimensional symbolic scheme that another person can use to “program” their neurons.
Krakauer calls collectively constructed external tools “exbodied,” distinguishing them from embodiment. Maps, abacuses, chessboards and Rubik’s Cubes are examples of artifacts whose material structures can carry collective discoveries and provide an external vehicle for computation that the brain would otherwise have to perform.
The “embodiment helix” runs from collective artifact to individual mind and back. Someone can memorize a map, burn it, navigate using the internal representation, explore further and contribute improved information for others to internalize.
Individuality is therefore scale-dependent: a bounded object contains enough information to propagate itself over a chosen time and resolution. A cell may qualify, while Einstein’s work would be incomplete in one mind and probably more fully preserved by a group; evolvability also propagates not just information but operators for generating novelty, from mutation to imagination.
6. AI’s endpoint could be augmentation—or the outsourcing of humanity
Krakauer argues that evolution can be viewed as moving from more “mortal” information processing in organisms toward more persistent cultural, software and hardware stores. Technologies complement human deficits—we calculate poorly, hence calculators and the abacus—whereas walking and singing are abilities humans already perform well and for which technologies are harder to build.
The danger appears when assistance removes any reason to acquire the skill. A map can be internalized; a superior GPS invites total delegation. Krakauer says LLM-assisted emails are already “100% rubbish” and expects the human share to keep shrinking until “eventually everyone will sound the same.”
Krakauer’s physiological analogy is delegating the gym: another person’s exercise cannot preserve your muscles. “The brain is an organ, like a muscle,” so fully outsourced cognition will atrophy; superintelligence is only interesting to him if it makes him more intelligent, not more stupid, servile or dependent.