Ep. 004: The Impact of AI Datacenters on Consumer Power Costs
Ep. 004: The Impact of AI Datacenters on Consumer Power Costs
Summary
- Jeremie Eliahou Ontiveros’s core call: the 15–20% electricity-bill spikes in PJM territory are mostly a market-design failure, not raw AI demand. Data centers contribute “to some extent,” but the real culprit is a capacity auction run once a year on regulator forecasts hitting the first US load growth in 20 years — “this market was not designed for an environment where load is growing.” ERCOT’s real-time pricing lets participants see demand coming and place generation ahead; PJM’s signals arrive late, so “everything is just late and you just create a mechanism that is constrained.”
- PJM’s supply side quietly shrank
40 GW (-20%) in four years, and 41% of the decline — 14 GW — came from methodology changes, not retired plants. The thermal accreditation reform followed bad 2021–22 winters that revealed plants fail in clusters when it’s cold, so the cut from 170 GW offered to 156 GW is “legit” — but it means capacity numbers are moving opposite to the political push for more generation, with new gas only +3 GW and renewables +1.5 GW over the same window. - The ratepayer risk mechanism to watch: regulated transmission earns a required return on equity, so an overbuilt, under-utilized asset gets passed straight to consumers — “this is the worst-case scenario outcome.” The fix is custom tariffs; the specimen deal is Oracle’s ~1 GW IT (1.4 GW gross) Stargate site in Michigan: a $4B utility commitment alongside a battery-purchase commitment and a 15–20-year power agreement with a minimum demand charge, meaning Oracle pays even if the load never materializes.
- The desk believes Anthropic is crossing OpenAI on total ARR as of the February exit — roughly doubling ARR (+$10B) in two months — and that reported Claude Code revenue of $2.5B run-rate is “way too low.” Doug says Anthropic counts Claude Max subscription usage inside Claude Code as Claude Code, while API usage—including Claude Code API mode—is counted as API; SemiAnalysis’s own spend runs 90–95% through the API and isn’t attributed. Jordan reads the acceleration as true business adoption, not hype.
- The Department of War spat is a distribution gift, not clearly a revenue driver — and the speakers disagree on how much it matters. Doug notes Claude hit #1 on app downloads, above OpenAI, right on the DoW news (“that’s a pretty hardcore mog… a zeitgeist moment”); Jeremie pushes back that downloads may be random app trials and that the revenue impact was “most likely unrelated to that specific issue.” The narrative inversion stands either way: Anthropic came from the business world and is now beating OpenAI on consumer.
- Jeremie’s tradeable thesis: the “models are commodities, value accrues to the application layer” view is “just completely wrong.” Anthropic has models nobody else gets, owns the coding interface, and can build applications using them better than anybody else can, effectively for free — “the model companies will become the everything companies.” Doug’s counter: as with Moore’s Law CPUs, “until the generalized models don’t get better every three to six months, there’s no point of doing any specialization.”
- China is an underappreciated demand vector leaking into global GPU numbers. Doug’s speculative take is that Qwen leadership was pushed out over Chinese New Year user-acquisition misses after perhaps $1B of spending; ByteDance “won, for sure,” with CDance running a months-long paid queue (Jordan cites 14K of 30K). Doug says Chinese labs VPN through Japan/Korea — claims that region is 15% of Anthropic revenue are “Chinese demand” — and “if Eastern companies had access to Western compute… I think they would be ahead.”
- Pre-GTC positioning note from Doug: the SRAM-will-kill-HBM scare ahead of the LPU session is “the dumbest narrative I’ve heard recently.” Everyone is “super long HBM” and spooked into selling it — investor positioning, not a grounded technology conclusion.
Deep dive
1. Pre-GTC investor positioning: the SRAM scare is not analysis
- Doug expects the LPU session — presented by the former founder-CEO of Roth, Jonathan Roth — to be the sold-out, standing-room event at GTC, precisely because positioning is stretched: “everyone is super long HBM right now, they are all scared that SRAM is gonna take over HBM… that’s the dumbest narrative I’ve heard recently.”
- Housekeeping worth knowing: SemiAnalysis hosts a hackathon the Sunday before GTC with compute grants and Claude Code tokens; Doug skips GTC for OFC — “we have to have some presence at the other place.”
2. PJM’s bill spike is 20 years of no load growth meeting a once-a-year auction
- Jeremie’s framing of the free-tier article (“Are AI Data Centers Increasing Consumers’ Electricity Prices?” — his first Claude Code project, using free public data): everyone says energy is the constraint and it’s AI’s fault, and “to some extent it is,” but people over-index on the flashy short-term data point — 15–20% bill increases in PJM, including Virginia, New Jersey and Ohio, versus two years ago. The bigger driver is market design meeting “the first time in 20 years we see load growth.”
- The mechanism: PJM runs a capacity auction roughly once a year, two years ahead, with demand set by its own forecast and supply by its own accreditation judgments. Jeremie’s contrast — “you basically go back to the debate, communism versus capitalism” — is that ERCOT lets participants price supply and demand in real time and place generation ahead, while in PJM “everything is just late and you just create a mechanism that is constrained and just basically becomes bottlenecked.”
- Jordan’s read of the key chart: PJM keeps revising its load forecast because it isn’t tracking data centers properly — at least not to the fidelity possible with SemiAnalysis’s data center model — and every forecast miss in an auction system “costs everybody a whole bunch of money.”
3. The supply base shrank 40 GW — and 14 GW of that was a measurement change
- PJM’s offered supply fell ~40 GW, roughly -20%, in four years; the bridge shows a substantial chunk is methodology. The thermal accreditation reform alone accounts for 41% of the decline — 14 GW — after bad winters in 2021 and 2022 revealed clustered failures: plants were accredited on individual historical reliability, not “the clustering impact of, like, when a summer is bad, many power plants stop working at the same time.” Jeremie calls the reform “legit” — you never actually had 170 GW, so marking down to 156 is honest — but a real-time market would have priced the risk faster.
- Jordan’s irony: at the exact moment everyone wants more generation online, accreditation moved the numbers 14 GW the other way, while new gas added just +3 GW and renewables +1.5 GW.
- On coal, Doug asks whether retirements get delayed. Jeremie: at the federal level, stopping coal-plant retirement plans has already happened (“Fair Coal Order 206, if I remember correctly, or whatever”), and economically 40–50-year-old plants pencil if prices rise. That’s the core tension of the whole piece: consumers expect low prices, but “to incentivize new generation to come online, you need prices to actually be high” — which is exactly why the discussion is moving behind the meter.
4. Custom tariffs are how hyperscalers keep their costs off your bill — Oracle’s Michigan deal is the specimen
- The worst-case transmission math: transmission is regulated with a required return on equity, so if a gigantic asset is under-utilized, “those costs actually end up being passed on to the consumer” — the builder must turn a profit regardless and takes no risk. Custom tariffs are the counter: hyperscalers directly fund grid upgrades and generation via long-dated agreements so their impact “doesn’t flow through everyone else.”
- The specimen: Oracle’s ~1 GW IT / 1.4 GW gross Stargate site for OpenAI in Michigan — a $4B utility commitment alongside a battery-purchase commitment, plus a 15- or 20-year power-price agreement with a minimum demand charge: “if the load somehow doesn’t materialize, they’re gonna have to pay a charge regardless.”
- Timeline reality-check from Jordan and Jeremie: Jeremie thinks site development began about two years ago with a third-party developer he identifies as “correlated digital,” if he remembers correctly. He thinks Oracle signed around October; the data center is expected to be operational in 2027, with the full site hoped for by mid-2028 — call it roughly four to five years start to finish. A caveat from Jeremie on why not everyone does this: “you have to take speculative risk… you have to make a forward bet” before demand is proven, and sometimes “it’s just easier to just build your own power plant.”
5. Anthropic crosses OpenAI on ARR — and its reported Claude Code number understates reality
- The crossover call: Anthropic is at the run rate OpenAI exited last year with, in March — “it’s like a two or three-month lead” — with the crossover pegged around the February ARR exit. Jordan says Anthropic has doubled ARR in two months, roughly $10B added. Doug’s rueful aside: “there’s a reason why I didn’t make this bet in ‘26” — he and Jordan had a beer bet on Anthropic winning dating to the summer.
- The attribution catch, from the team’s own usage: Doug says Anthropic defines “Claude Code” as Claude Max subscription usage inside Claude Code; API usage — including Claude Code in API mode — is counted as API rather than Claude Code. Jeremie adds that he thinks fast mode is on subscription, while one-million-token context is only on API. SemiAnalysis’s spend is “90, 95% going to API,” so the stated $2.5B Claude Code run rate is “actually way too low.” Doug’s theory on why: it’s in Anthropic’s interest to show “generalized enterprise demand without having a ginormous vector” — one hyperviral thing people are “yoloing $8 billion on.”
- The demand isn’t one pipe anyway, per Jordan: Cursor claims a huge acceleration, Windsurf/Vercel/Replit use plenty of API tokens, and Claude for Finance, Legal, and Security Software “tanked a bunch of SaaS stocks one day after the other.”
6. Department of War vs Anthropic: eyeballs, not clearly revenue — a live disagreement
- Doug’s case that the fight helped: Claude hit #1 on app downloads, above OpenAI, timed to the DoW news — “that’s a pretty hardcore mog… it’s a zeitgeist moment.” Jordan’s version: it’s eyeballs — “more people know what Claude and Anthropic means now because they are following politics, not technology.”
- Jeremie’s pushback is worth keeping: “Do you guys actually think that it had any impact on their revenue?” Downloads could be random app trials; he thinks the revenue is probably driven by Claude Code and is “most likely unrelated to that specific issue.” Jordan separately says the two-month ARR acceleration is true business adoption rather than temporary hype; he points to companies “removing Salesforce and building their own.” Doug concedes that downloads are “a cherry on top,” not what’s moving revenue.
- Jordan frames the narrative inversion: the old thesis was consumer-household-name OpenAI penetrating enterprise; “now we’re actually seeing the opposite way — no one knew Anthropic, they come from the business world, and now they’re actually beating OpenAI on consumer.”
7. Everything companies vs. the Moore’s Law of models
- Jeremie’s thesis: the commoditized-models/application-layer view is “just completely wrong” — Anthropic has models we don’t have, controls the coding interface, and can build applications “better than anybody else can, effectively for free”; he names Salesforce, Harvey and CrowdStrike as potential targets after GitHub Copilot. “The model companies will become the everything companies.”
- Doug’s counter, via semiconductor history: specialization loses while the general thing compounds. CPUs won because they got 50% better every year, “so it never made any sense for any specialization ever to happen” — only when Moore’s Law topped out did specialists win. “Until the generalized models don’t get better every three to six months, there’s no point of doing any specialization.”
- Why coding specifically: Jordan traces the flywheel from chat thumbs-up/down data to RL, which needs far richer signal per trace — “the quality of data in coding is much higher, and so therefore the flywheel is in coding.” Doug goes bigger: “software historically is a layer on top of code for humans to interact with the machine,” and now the agent sits underneath software — “humans are just getting closer to the computer.” Analysts can ask the computer to perform analyses without learning R, statistics or specialized software.
- Tooling churn as a tell: some software accelerates on this wave (Tailscale — used by CoreWeave for cluster authentication — QMD and Rippling) while others get “totally left in the dust”; Doug is even graduating from Obsidian (“why fuck around with a markdown when I can just do it in the terminal?”), and has gone VPS-first with a DigitalOcean droplet while also using remote SSH to a Mac mini.
8. China: a Qwen leadership shakeup, a months-long video queue, and demand leaking worldwide
- Doug’s “pretty cold take” on the Qwen shakeup: leadership was pushed out because it didn’t acquire enough users during Chinese New Year, after spending maybe a billion dollars — KPIs are now all user acquisition, and “ByteDance won, for sure.” Its video model, referred to in the exchange as CDance, is the one everyone at SemiAnalysis actually wants: the paid queue is reportedly 14K of 30K, and even paying users can’t skip the line — “old school internet.”
- The galaxy-brain macro point: accelerating Chinese consumer AI adoption “might be part of the demand that we’re not recognizing well in GPUs” — China domestically can’t support the compute, so it “ends up in Singapore or Europe or everywhere else.” Supporting data from Jeremie: something like 15% of Cursor users were from China last year, and people he spoke with said Western AI agents generally aren’t blocked there because the local ones suck; he explicitly says he does not know whether Claude Code itself is blocked.
- Doug’s reported allegation: Chinese labs VPN into Japan and Korea and shift sleep schedules “to replicate an average Korean vibe coder” — so the earlier claim that Japan/Korea represented 15% of Anthropic revenue may actually reflect Chinese demand. The distillation joke concerns a GLM model whose version is disputed in the exchange: its “distillation of Claude is better than Claude can do itself.” Jordan’s defense — GLM does “a lot of pretty good engineering” — and Doug’s concession-turned-escalation: with equal compute, “I think they would be ahead.” Both await DeepSeek V4.