Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
Summary
- Naveen Rao publicly unveiled what he called the “first physical dynamical computer ever built”: Unconventional AI taped out the chip June 1, five months after the company “started in earnest in January,” and it is already generating images in the lab. The headline number: ~500 nanojoules per image versus millijoules on a GPU — “many orders of magnitude more efficient than a standard computer, and it’s because it just doesn’t move information around.”
- Rao’s prior track record frames the bet: he founded Nervana Systems in 2014, sold it to Intel (saying he sold it “way too early”), ran Intel’s AI group, then built GPU-scaling infrastructure with his team before joining forces with Databricks in 2023.
- The company’s goal is 1,000× power efficiency versus existing hardware, and Rao has pulled the timeline in from five years to three and a half — “we’ve actually solved very deep scientific problems more quickly because of AI, interestingly enough.” The stated endgame is bolder still: “the overarching goal of this company is to beat biology.”
- The energy math driving the thesis: Google alone crosses 3.2 quadrillion tokens per month; at 10 joules/token that’s 12 GW — against ~40 GW of total US data-center power and under 100 GW worldwide. Rao’s estimate: “we’re going to run out of energy pretty fast, in about 3 years or so,” and since ~50% of the cost of serving a token is energy, the business case is “we’re going to monetize that 1,000× better than existing hardware.”
- Biology is the existence proof: a human brain runs on 20 watts, a monkey brain on one watt — the same as your phone — and a squirrel’s on 8 milliwatts. The mechanism gap is data movement: the human cortex moves ~16 billion bits/second while a GPU moves ~30 trillion bits in and out of memory, “10–100× more than that” inside the chip.
- The architecture moves beyond conventional von Neumann designs — “compute and memory in one thing, we don’t have a memory interface” — branded “4D computing”: three physical dimensions via die stacking plus time as the fourth. A supporting breakthrough on sparsity turned n-squared connection scaling into a “holy grail” result: throwing away connections made systems more efficient, more scalable, and more trainable at once.
- Under Chamath’s questioning, Rao said a full product is “within two years”: a rack-scale data-center system, “tokens in, tokens out through a network cable, but the inner guts are completely different.” Existing models will work — porting happens “at the model layer,” not the operations layer — but matmul as such doesn’t exist in the hardware, and the software stack is Python libraries, not CUDA.
- The investor kicker is Jevons’s paradox: Rao says halving the price can lead to more than 2× consumption, while making something 1/1,000th the price can lead to consuming more than 1/1,000th as much. He thinks a 1,000× disruption of a roughly $1 trillion 2030 AI market — “maybe it’s bigger than that” — will create “the largest market that humanity’s ever seen,” with the topology shifting from gigawatt data centers to “many small data centers all over the place” and eventually “billions of robots.”
Deep dive
1. AI hits an energy wall in ~3 years — and energy is now the scarce asset
- Rao’s track record frames the bet: he founded Nervana Systems in 2014, sold it to Intel (saying he sold it “way too early”), and ran Intel’s AI group; after leaving in 2020, he and his team built GPU-scaling infrastructure and joined forces with Databricks in 2023.
- Rao’s back-of-envelope using Google’s own public figures: 3.2 quadrillion tokens/month × 10 joules/token (the lower end for models) = 12 GW for one company’s AI services, versus ~40 GW of US data-center power and under 100 GW globally. If models grow and demand grows, “we’re going to run out of energy pretty fast, in about 3 years or so, is my estimate.”
- The industry’s constraint has migrated: floor space → networking → GPUs → energy. “First you think about energy. I get the energy contract and then I have to figure out how to fill it.” With ~50% of per-token serving cost being energy, Unconventional’s pitch is blunt: monetize every watt “1,000× better than existing hardware.”
2. Biology proves efficient intelligence is physically possible
- The reference points as told: human brain 20W; monkey brain 1W — “the cellphone in your pocket runs on about one watt”; a squirrel’s brain, nailing branch-to-branch jumps “a thousand times out of a thousand,” runs on 8 milliwatts. “You could run over 100 squirrel brains on your phone.”
- The inefficiency is data movement: the human cortex moves ~16 billion bits/second; a GPU moves nearly 30 trillion bits in and out of memory — and 10–100× more inside the chip. “Most of the energy in a computing system goes into moving information around.”
- Rao’s historical frame: computers were built to be faster than the alternative — ENIAC’s alternative was humans doing artillery-trajectory calculations — rather than with energy efficiency as the central metric. As transistor shrinking stops delivering the old efficiency gains, he argues the paradigm must be rethought.
3. Cut out the middleman: compute directly in physics
- The design philosophy is stripping lossy abstractions between neural networks and silicon: “there’s no linear algebra [in your brain]… it’s actually the physics of the neurons that gives rise to intelligence. We want to mimic some of that with a semiconductor.” His metronome example carries the idea: metronomes on a rolling plank synchronize purely through physical coupling — a dynamical system that can produce emergent computation, like bird flocks and ant colonies.
- The proof-of-concept was UNO, an open-source image-generation model built on oscillators; a sparsity result then improved the scaling story — throwing away some n-squared connections not only saved energy but made systems more trainable: “one of these rare things where you get something that’s more efficient, more scalable, and actually gives you more performance.”
4. The reveal: a working chip in five months, ~500 nanojoules per image
- Talking about it publicly for the first time, Rao showed what he called “the first physical dynamical computer ever built”: the company started in earnest in January without a team, taped out June 1, and had the chip back with results. Generated images cost ~500 nanojoules each versus millijoules on a GPU — “proof positive that it works” — and the approach can support sequence and language modeling.
- The architecture claim: unlike CPU/GPU/compute-in-memory, all von Neumann architectures, “each individual computing element is a memory” — “4D computing,” three physical dimensions from die stacking plus time in the dynamics.
- The roadmap: today’s hardware sits ~10 billion times from the thermodynamic limit of intelligence-per-watt; mammalian brains are within one or two orders of it. In 3.5 years Rao thinks they can hit “the limits of 2D lithography,” and the goal is to beat biology, enabling small distributed data centers and billions of robots. Under Jevons’s paradox, sharply cheaper compute could expand demand rather than simply reduce consumption.
5. Chamath’s Q&A: path to product and the ecosystem problem
- On timing and form factor: “within two years” of a full product — a VM in a managed environment, effectively a whole-rack data-center product: “tokens in, tokens out through a network cable, but the inner guts are completely different from an existing computer.”
- Chamath’s pushback on switching costs — the ecosystem is “mechanically reductive” around KV caches and existing abstractions — drew Rao’s answer: “let’s make it really, really compelling to move.” Porting happens at the model layer, not the operations layer; existing models will work, though “there’s a fair bit of compute required to make that transition happen.” Matmul doesn’t exist natively — each timestep is analyzable as current-state matrix × transition matrix.
- The team-building candor: dynamical-systems theorists and chip designers “don’t talk to each other,” and bridging that span “is actually one of the most challenging things about this company.” The CUDA-like layer, though Rao stressed it is not CUDA, is a set of Python libraries for expressing “time-varying elements that have stochastic behavior.”