Michael Nielsen – Why aliens will have a different tech stack than us
Michael Nielsen – Why aliens will have a different tech stack than us
Summary
- The premise Dwarkesh came in with — that AI will race ahead in science the way it did in coding because both have tight verification loops — is what the episode dismantles: “there’s an infinite number of theories that are compatible with any given experiment.” His best exhibit is Prout’s 1815 hypothesis that all atomic nuclei have whole-number weights, where chlorine measured 35.5 and then, more precisely, 35.46 — moving further from the correct fraction, not closer. It took 85 years to learn that isotopes can’t be chemically distinguished, meaning the verification loop was “actively hostile against the correct theory” for the better part of a century.
- Nielsen’s reframe of the flagship AI-for-science result is the most tradeable line in the conversation: “AlphaFold really isn’t about AI.” The win rests on the Protein Data Bank — X-ray diffraction, NMR, cryo-EM, several billion dollars spent to obtain roughly 180,000 structures — after which “we fitted a nice model at the end of it, which was a tiny fraction of the entire investment.” The dominant input is instrumented data acquisition, not model capability.
- Bottlenecks don’t disappear, they relocate — and Nielsen argues that’s near-tautological: “you’re going to get bottlenecked at the places where your existing method doesn’t apply. Definitionally, there’s no crank you can turn.” His live example is his programmer friends, whose prototypes went from three weeks to three hours and who are now stuck on having interesting design ideas — “there’s not really a verification loop for knowing that a design idea is very interesting.”
- On whether an over-100-million-parameter model can be an explanation, Nielsen offers three readings and takes the third seriously: they may be “a new type of object” you can merge and distill, the way Mathematica turned a 100-page equation from a dead end into something workable. Dwarkesh’s counter is sharp — a deep-learning model trained on sky data from 1500 would just discover more epicycles, parameters X to Y encoding the next one — “we don’t have the verbs yet.”
- Diminishing returns may be contingent, not intrinsic. Nielsen’s dessert-buffet analogy concedes the static case (best desserts go first) but insists someone keeps restocking the table — computer science arrived as a side effect of esoteric questions in the philosophy of mathematics and “the diminishing returns argument just didn’t apply there.” Against Bloom’s finding that Moore’s law’s 40% annual density gain needed 9% more semiconductor scientists per year, his reply is that those metrics are deliberately narrow and that one possible unlock is institutional: capital allocation, training, “basic security for researchers, so they’re not worried about the Inquisition.”
- The anti-determinist call: “Most parts of the tech tree are never going to be explored” — too many deep ideas, too few explorers, so choices about direction actually matter. He points at DDT, chlorofluorocarbons and the Non-Proliferation Treaty as early evidence of institutions that preemptively decide “we’re not going to go down that path,” which is the opposite of technology-as-inevitable.
- Dwarkesh lands an observation Nielsen says he hadn’t considered: if tech trees are path-dependent, there are “humongous gains to trade” between civilizations far into the future — “it makes friendliness much more rewarding.” Nielsen’s hedge: ideas diffuse cheaply, capacity doesn’t, and comparative advantage is “a special limited model” — “Chimpanzees can do interesting things, but we don’t trade with them.”
- The portfolio lesson for anyone funding research: the same move wins and loses unpredictably. Uranus’s wobble produced Neptune in 1846; Mercury’s 43 arcseconds per century produced Vulcan, which doesn’t exist and needed general relativity instead. Nielsen’s rule — “A priori, you can’t tell which of these is the thing to do, and you actually need to do both” — plus Dwarkesh’s added base rate that 99.9% of anomalies are mundane, like the Pioneer spacecraft’s asymmetric thermal radiation.
Deep dive
1. Michelson-Morley did not kill the ether — and falsification never worked the way the textbook says
- Nielsen’s correction to the YouTube version: Michelson and Morley were “an experiment to test different theories of the ether against one another,” specifically hunting an ether wind that would speed light travelling with it and slow light travelling against it. They found no wind. That “ruled out some theories of the ether, but not all” — and the leading physicists of the day read it as “this gives us a lot of information about what the ether must be, but it doesn’t tell us that there is no ether.”
- The load-bearing detail is how long the losers held on. Michelson ran the first version in 1881, redid it in 1887 after Rayleigh flagged problems, and was still running ether experiments in the 1920s, believing in it until his death around 1929, with a public statement to that effect a year or two before he died.
- Miller kept going too, climbing Mount Wilson on the theory that at altitude the ether wind wouldn’t be dragged along by the Earth — and announced he’d measured the effect. Einstein’s response is the episode’s first great line: “Subtle is the Lord, but malicious He is not.”
- Nielsen’s methodological point, hedged exactly as he hedged it: this “certainly doesn’t show that ideas about falsification are wrong or falsified, but it does show that the most naive ideas… Things are often much more complicated than you think.” Dwarkesh adds that even the definite article is a tell — “even just the word ’the’ there is a misnomer.”
2. The community favored relativity before experiments preferred it
- Lorentz got the mathematics of frame conversion right before Einstein but interpreted it through the ether: length contraction and time dilation as “the effect of moving through the ether,” a pressure “warping clocks” and warping measures of length. His “local time” was a quantity he wasn’t trying to give physical meaning; Poincaré, Nielsen thinks, got closer to realising it is the time registered by clocks.
- The muon evidence came another forty-odd years later, and Nielsen tells it as a story rather than a datum: cosmic rays hit the atmosphere, shower muons, and the muons decay far too slowly to survive the trip under classical assumptions. The measured decay rates Nielsen dates to around 1940, possibly published in 1941, match special relativity exactly. Had Lorentz still been alive, “it seems quite likely that he would have tried to save his theory by patching it up yet again.”
- Dwarkesh restates his claim more carefully — the community “adopted what we in retrospect consider the more correct interpretation before it was actually experimentally shown to be preferred” — and Nielsen interrupts the word “process”: it “carries connotations of something set in advance,” and there is no centralized authority or method. Lorentz, Poincaré and Michelson, all outstanding scientists, “never reconciled themselves.”
- The through-line Nielsen leaves standing: “Great scientists can remain wrong for a very long time after the scientific community has broadly changed its opinion.” Progress happens anyway, without an articulable procedure for it.
3. Expertise as a prison — Poincaré had the premises and still missed
- The Poincaré case is the one Nielsen calls amazing: he seems to have had the principle of relativity and the constancy of the speed of light in all inertial frames — “basically the ideas that Einstein uses to deduce special relativity” — while still treating length contraction as a dynamical effect, particles pushed together by some external force, rather than pure kinematics.
- In a paper Nielsen dates to 1909, the dynamical picture survives, and the honest non-answer is the point: “Why is he clinging onto this idea? I don’t know. I’ve obviously never met the man.” The diagnosis offered in the studio — “It’s almost like he knew too much. He had almost too grand a vision in mind” — with Einstein winning by subtraction.
- Nielsen’s hypothesis, hedged and flagged as contested: in the 1890s a teenage Einstein believed in the ether too, but “he’s not quite as attached as these older people were. Maybe they were a little bit prisoners of their own expertise. That’s my guess. Some historians of science would certainly disagree.” Nielsen notes Einstein later played the same role on quantum mechanics and cosmology.
4. Copernicus was neither more accurate nor simpler — unification is what actually persuades
- Dwarkesh’s setup is the sharpest test of scientific taste in the episode: the Ptolemaic model was more accurate, having had centuries of epicycles added, and — less appreciated — Copernicus needed more epicycles, not fewer, because he insisted the Earth move in a perfect circle in equal time. “So how could you have known ex ante that Copernicus was correct and Ptolemy was not?”
- Nielsen’s partial answer, offered as compelling to him rather than as the historical record: Newton’s gravitation explained Kepler’s planetary motions, parabolas of thrown objects on Earth, and the tides from lunar and solar pull on water. “You have what seem like three very different disconnected phenomena all being explained by this one set of ideas. That starts to feel very compelling.”
- The Keynes essay on Newton supplies the frame for what kind of mind does this — “Newton was not the first of the age of reason. He was the last of the magicians, the last great mind which looked out on the visible and intellectual world with the same eyes as those who began to build our intellectual inheritance rather less than ten thousand years ago” — plus the line that his esoteric work showed “extreme method in his madness,” “just as sane as the Principia if their whole matter and purpose were not magical.”
- Dwarkesh converts this into the actual AI question: were the same parsimony-and-aesthetics heuristics that served Newton and Einstein transferable across time and discipline? “Even if we can’t build a verification loop for science, maybe if the taste tests point in the same direction, you can at least encode that bias into the AIs.”
5. Bottlenecks migrate, definitionally, to wherever your method stops working
- Nielsen’s answer to the encode-the-taste hope is structural: “If you’re attempting to reduce science to a process, you’re attempting to reduce it to something where there is just a method which you can apply, and you turn the crank and out pops insight. You can do a certain amount of that, but you’re going to get bottlenecked at the places where your existing method doesn’t apply. Definitionally, there’s no crank you can turn.”
- Because people study what worked, they stop getting stuck where their predecessors got stuck — “they keep getting bottlenecked in different places” — and the difficulty of the idea sets both the size of the blockage and the size of the payoff: “the greater the bottleneck, but then also the greater the triumph.”
- Dwarkesh’s puzzle for the theory: Principia in 1687, Origin of Species in 1859, and Huxley’s reaction to Darwin was “How extremely stupid to not have thought of this,” while nobody reads the Principia wishing they’d beaten Newton to it. Nielsen relocates Darwin’s genius from the idea — animal breeders knew artificial selection — to “understanding just how central it was to biology” and doing the grinding work of connecting it to geology and everything else.
- The contrast is exactly the verification-loop contrast: Newton could check the Moon’s orbital period and the tides, while for Darwin “there’s no individual piece that is overwhelmingly powerful,” and he had no mechanism, no genes. Dwarkesh’s timing evidence — Lyell’s deep time in the 1830s, paleontology, biogeography from the age of colonization, and Wallace’s near-simultaneous manuscript (“Darwin’s like, ‘Fuck.’ I don’t think that’s an exact quote, but it’s pretty much correct”) — suggests the building blocks were necessary. Nielsen agrees deep time was load-bearing: on a 6,000-year Ussher timescale “you would need to see evolution occurring at a massive rate during human lifetimes, and we’re just not seeing that.”
6. AlphaFold is a data-acquisition story wearing an AI costume
- Nielsen’s flat statement, offered as the thing that should give pause about closing the RL loop on discovery: “the big signature success so far, which is certainly AlphaFold. AlphaFold really isn’t about AI. A massive fraction of the success there is the Protein Data Bank. It’s X-ray diffraction, NMR, cryo-EM, and the several billion dollars that were spent obtaining those 180,000-odd protein structures.”
- The sequencing matters for anyone underwriting AI-for-science: “we spent many decades obtaining protein structure just by going out and looking very hard at the world experimentally, and then we fitted a nice model at the end of it, which was a tiny fraction of the entire investment.” The AI part he calls “very impressive and quite remarkable, but it is only a small part of the total story.”
- He does not use this to dismiss the result — structural biologists “seem to think that AlphaFold was an enormous advance. It was a shock.” His question is narrower and better posed: “Are we primarily bottlenecked on one type of thing, or are we bottlenecked on multiple types of things?”
7. Are trained models explanations? Three answers, and “we don’t have the verbs yet”
- Dwarkesh frames the pivot: general relativity nets out to equations and predicts things it was never built for, like Mercury’s precession, whereas AlphaFold “is encoding these different relationships between things we can’t even interpret over 100 million parameters. Are those really the same thing?” Nielsen calls it “maybe a really pivotal question.”
- Answer one is the conservative one — you want few free parameters and simple models that explain a lot, so AlphaFold is “nice and maybe helpful as a model, but it’s not a scientific explanation.” Answer two is that it contains lots of little explanations inside it, recoverable by “an archeology of AlphaFold” through interpretability; the precedent he cites is Magnus Carlsen apparently changing his game after public forensics on AlphaZero, though “I don’t think there’s any public confirmation of this.”
- Answer three is the one he finds most interesting: models are “a new type of object” that should be taken seriously as explanations, because now “we can merge them, we can distill them.” The anticipation is Mathematica — a 100-page equation in 1920 meant you gave up, and today “that’s an object now, a thing that you can work with,” sometimes yielding simple answers at the end.
- Dwarkesh’s pushback is the strongest moment in the exchange: in 1500, train a deep-learning model on sky observations and interpret it, and “you’d just be able to keep building on Ptolemy’s model… Parameters X to Y encode this epicycle.” Distillation reduces error locally, but the Copernican move is globally better while locally worse — “with raw gradient descent, I don’t really feel like it would do that.” His summary: “We don’t have the verbs yet.” Nielsen’s counter is the forcing function — Einstein saw immediately that Newtonian action-at-a-distance plus special relativity permits faster-than-light signalling, “and it’s not a big leap to realize we have a big problem here.”
8. Verification loops can run hostile for 85 years — so keep every research program alive
- Nielsen’s paired example is the cleanest argument for portfolio breadth: Uranus off its predicted spot produced Neptune, “wonderful, massive success for Newtonian gravity”; Mercury off its predicted spot — Dwarkesh supplies 43 arcseconds per century — produced Vulcan, which isn’t there. “You’ve pursued very similar ideas, and it’s been very successful in one case, and it’s been completely and utterly unsuccessful in the other case. A priori, you can’t tell which of these is the thing to do, and you actually need to do both.”
- Dwarkesh spells out how a committed Newtonian saves the theory anyway: cosmic dust occluding the planet, a planet too small to see, build a bigger telescope, a magnetic field spoiling the measurement. Nielsen’s 1990s version is the Pioneer spacecraft being slightly off course — you can get excited that general relativity is wrong, but the accepted answer is asymmetric thermal radiation producing a tiny acceleration toward the sun.
- The base rate is the punchline, and it cuts against how these stories get told: “99.9% of the time, it just turns out to be some effect like this thermal acceleration… Unfortunately, there’s a lot of selection bias going into those stories.” Dwarkesh’s addition: “there’s no ex ante heuristic which tells you which case you’re in.”
- Prout is the extreme case, from Lakatos. A chemist hypothesises in 1815 that all atomic nuclei have whole-number weights; chlorine comes out at 35.5; the half-integer patch dies when better measurement gives 35.46, further from the correct fraction, not closer. Isotopes — physically but not chemically distinguishable — arrive 85 years later. “You just need this remnant to be defending… There’s no ex ante reason it’s the preferred theory.”
9. The tech tree is enormous, diminishing returns may be contingent, and most branches will never be walked
- Nielsen’s core intuition pump is computer science having its theory of everything in the 1930s and then spending “ninety-odd years exploring the consequences”: public-key cryptography was “incredibly deep, very non-obvious” and “lay hidden already in the 1930s,” and inside that again sit the ideas behind cryptocurrency and collectively maintained ledgers. Knuth’s anecdote seals it — told to come back “when there’s a thousand deep theorems,” he noted decades later that “there clearly are a thousand deep theorems now.”
- Phases of matter make the same case physically: three, or sometimes four or five in school, then superconductors, superfluids, Bose-Einstein condensates, quantum Hall and fractional quantum Hall systems — “we’re going to be able to start to design them in some sense.” His self-assessment of the explorers: “We’re basically slightly jumped-up chimpanzees, so we’re slow and it’s taking us time.”
- Against Dwarkesh’s Bloom-style challenge — Moore’s law delivering 40% annual transistor density growth while requiring 9% more semiconductor scientists per year, replicated industry after industry — Nielsen attacks the framing first (“all of their examples are narrow… GPUs don’t show up there”) but concedes the phenomenon may be real, then points to external conditions as one possible cause: pre-1700 progress was slow and irregular partly because you need institutions around the ideas — training, allocation of capital, “even just basic security for researchers, so they’re not worried about the Inquisition.” The dessert buffet gets restocked; that’s why new fields let a twenty-one-year-old make breakthroughs.
- The conclusion he holds most firmly, and the one with the most consequence: “Most parts of the tech tree are never going to be explored. There are just too many deep ideas waiting to be discovered, and not only we, but nobody ever, is going to discover most of them.” Hence his dislike of technological determinism, and his read of DDT, chlorofluorocarbons and the Non-Proliferation Treaty as institutions that shape which branch we take. Dwarkesh’s extension — that path-dependent tech trees imply far-future gains from trade between civilizations, so “it makes friendliness much more rewarding” — draws Nielsen’s “I hadn’t thought about that at all,” then his limits: ideas spread cheaply, manufacturing capacity doesn’t, and “Chimpanzees can do interesting things, but we don’t trade with them.” Dwarkesh’s own caveat is that comparative advantage doesn’t guarantee terms above subsistence: we don’t keep horses on the roads, and sustaining one in San Francisco isn’t worth $100,000 a year.
10. Fields open when the conditions arrive, not when the geniuses do
- Asked why quantum computing wasn’t a thing in the 1950s — von Neumann pioneered computation and wrote the important book on quantum mechanics — Nielsen gives a deliberately banal answer: two unrelated things matured around 1980. Computation became salient because “you could go and buy an Apple II. You could buy a Commodore 64,” and the Paul trap gave us the ability to trap single ions and manipulate single quantum states for the first time.
- The image he leaves it with: Feynman got one of the first PCs around 1980 or 1981 and “was apparently so excited with this device, he actually tripped and hurt himself quite badly carrying his brand-new computing device.” Having someone that talented in quantum mechanics also that excited about the machines is “a very historically contingent coincidence… What similar story could you have told 10 years earlier? The conditions don’t exist for it.”
- His own entry is the first-person version of the market for follow-ups. “I’m 11 in 1985. I’m not thinking about this. I’m playing soccer.” Then a 1992 quantum mechanics class from Gerard Milburn, an ask for papers, and “a giant stack” including Deutsch and Feynman, at a time when essentially nobody in the world was working on it: “in some sense, I’m benefiting from the taste of this other person.” What hooked him was Deutsch’s conjecture that a quantum Turing machine could efficiently simulate any physical system — “I think in that paper, he more or less claims that he’s proved it. I’m not sure everybody would agree with that.”
- On quantum’s ceiling he refuses Dwarkesh’s tight bounds: “We’ve only been thinking about it for 40 or so years… and we haven’t thought that hard about it as a civilization,” and crucially “we’ve been doing it without the benefit of having the devices. That’s a pretty big bottleneck to have.” His speculative flourish — labelled as such — is that there may be a brief AGI-on-classical-computers era followed by AQGI on quantum machines, “probably capable of a strictly larger class of potentially interesting computations.”
11. The attribution economy is invented, and the real bottleneck is how you learn
- Open science’s win, Nielsen argues, is that Dwarkesh didn’t have to define it — twenty years ago he would have. The deeper point is that credit is constructed: three centuries ago Galileo and Kepler sometimes published results as scrambled anagrams, to be unscrambled later as a priority claim — “not an ideal foundation for a discovery system” — and it took over a century to reach papers, attribution and a reputation economy. Now code, data and in-progress ideas can be shared with no credit attached.
- His favourite demonstration that this is pure social construction: biologists told him “Biology is so much more competitive than physics that we need to protect our priority, so we can’t possibly upload to the preprint archive,” while physicists told him “Physics is so much more competitive than biology that we need to establish our priority by uploading as rapidly as possible to the preprint archive.” “Any attempt to change that economy results in a different system by which we construct knowledge.”
- On output volume, Dwarkesh cites Dean Keith Simonton’s equal odds rule — any given paper or book has roughly the same chance of mattering, so productivity is publication count, “Shakespeare was just publishing a lot,” with Gödel as the counterexample who “published almost nothing.” Nielsen splits work into routine (get good at it, outsource it, don’t prolong it) and high-variance (be willing to lose the time), and admits the ledger: “Sometimes I feel I’m too slow.” The book he wants written is “a very large number of biographies of people who are fantastically talented who just missed” — IMO gold medalists who tried to become mathematicians and failed — because “in many cases that’s actually more informative than anything else.”
- The closing therapy session is the most practically useful stretch. Nielsen’s demandingness test: when students hadn’t solved the problem, ask “If a million dollars had been at stake, would you have put the same effort in? And the answer is no, invariably. They’ve tried, but they haven’t really tried.” Dwarkesh’s diagnosis of his own work is that AI episodes have a forcing function — implement the transformer and you’ve clamped something down — while history episodes don’t, and that being stuck is the mechanism: essays written in a couple of days taught him nothing, ones that took three months he still remembers fifteen years later. Nielsen’s warning about LLMs closes it, via Alan Kay on Linux — “It’s just a great big ball of mud” — because “for a certain type of mind, there is a seductiveness in just learning systems and confusing that with understanding.”