GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink
GPT 5.5 vs Claude 4.7: OpenAI's Comeback From the Brink
Summary
- OpenAI is back “in the conversation” with GPT-5.5 after a stretch Max Kan calls “really dire.” Anthropic—riding Opus 4.5’s step change in coding and agentic ability—passed OpenAI on a like-for-like revenue basis in early-to-mid April (a leaked ~$19B ARR versus ~$24B at the start of the year), and GPT-5.4 “was honestly just an embarrassment” whose model card didn’t even compare against Opus. 5.5 is back on the frontier but is not “definitively better” than Opus 4.6/4.7 “despite what the Twitter propaganda machine was trying to push.”
- Dylan Patel’s release calendar: “Everyone’s releasing in two weeks”—Google and OpenAI definitely. The “Spud” (5.5) shipped with pre-training unfinished, so the next drop is finished pre-training plus more RL; Google’s is expected to be “mostly just” a multimodal swap.
- Fast-mode premiums are decaying while the price stays 6x: Opus 4.6 fast has slipped from 2.5x to under 2x speedup (90→70 tok/s against an unchanged 35–40 base). Yet this is the first time SemiAnalysis engineers chose fast over higher-quality tokens—Jordan Nanos’s read is that “4.7 is just not meaningfully better quality than 4.6 for people today.”
- Token pricing is starting to price out even heavy professional users. Dylan says the desk is “on the cusp” of being priced out; Mythos was quoted ambiguously at $25/$150 or $25/$125 versus Opus’s $5/$25, while Doug characterized it as roughly 5x, with fast mode 6x on top. Doug’s $800 boil-the-ocean scraping run versus a $55–100 data-enrichment API is the cautionary tale—and cost growth comes from new tasks (Jevons), not repricing old ones.
- Doug’s structural bear case for frontier pricing: Opus 4.5 may have crossed the threshold where day-to-day tasks are one-shottable without supervision. GPT-5.5-level intelligence in a 100–200B-parameter form factor in “probably less than a year” might mean the majority of people do not need frontier-level intelligence.
- Benchmarks have degraded to “a vibe check to make sure that the model’s not total trash.” Humanity’s Last Exam is esoteric multiple choice; SWE-bench scrapes GitHub issues with implementation-scoped unit tests. Meanwhile, 4.7’s new tokenizer can cost 35% more for identical output, and the desk’s truther take—“Opus 4.7 is actually Sonnet”—captures how small the model smells.
- The China open-source gap is widening again, per Dylan’s flat “Yes”—and it’s a compute-constraint story. Chinese frontier weights conveniently fit an 8x H200 pod’s memory domain, Ascend kernels only partially serve DeepSeek V4, and Max expects Meta—behind today but signing monster compute deals—to “pull away from all the Chinese guys” by H2’26 or H1’27.
- The form-factor fight is live: Dylan calls the CLI “a foregone relic” and says OpenAI’s app holds the true agent-orchestration vision; Max’s rebuttal: “it’s CLI all the way down. Pure maxi vision.” Tri Dao’s cracked-kernel workflow, via Dylan’s group chat: have Codex write it, then Opus fix the slop—“you can’t go the other way around”—though “everyone else at the firm prefers the other way around.”
Deep dive
1. GPT-5.5 pulls OpenAI back from the brink — Anthropic had passed them on revenue
- Max’s TLDR of his debut SemiAnalysis article: The Information leaked ~$19B ARR for Anthropic versus ~$24B for OpenAI at the start of the year, and even setting aside net-versus-gross hyperscaler accounting, Anthropic “pretty clearly surpassed them on a like-for-like basis in early to mid-April.” The driver: Opus 4.5 was “a real step change” in coding and agentic ability—from late November through early April, “everyone was basically just spamming Opus 4.5, 4.6 for all their workloads.”
- The GPT-5.4 verdict, unsoftened: “honestly just an embarrassment”—its release card compared only against past OpenAI models, not Opus. “That kind of tells you all you need to know.” With 5.5, Opus is back in the card and OpenAI is “back on the frontier”—not definitively better than 4.6/4.7 “despite what the Twitter propaganda machine was trying to push on release date,” but “definitely in the conversation.”
- Dylan’s forward calendar: “New release in two weeks… Google, OpenAI, maybe Anthropic.” The Spud “kind of like didn’t finish the pre-training and just released it”—so expect finished pre-training, more RL, then the drop; Google’s will be “mostly just gonna be like multimodal swap.”
2. Fast mode: the speed premium is decaying, and buying speed over quality is new behavior
- Doug can’t stay on Codex—“the usage limits raw dog me every day, bro”—and the desk agrees OpenAI’s fast mode is “pretty fake”: reduced reasoning depth, but not noticeably faster. Jordan’s taxonomy: priority mode is a roughly 2x premium for an SLA, not faster interactivity, and 5.3 Codex Spark is “definitely fake. That’s just a different model.”
- Jordan’s data on Opus: 4.6 fast launched around 90 tok/s per user versus 35–40 base; it now runs around 70 against an unchanged base—“not even two times faster for six times the price.”
- The behavioral tell: this is the first time any SemiAnalysis engineer traded fast over higher-quality tokens. Jordan’s explanation: “4.7 is just not meaningfully better quality than 4.6 for people today.”
3. Tokenomics: getting priced out, and whether frontier intelligence is even needed
- Doug’s structural bear case, framed via the OpenAI-Cerebras deal: if Opus 4.5 passed a key capability threshold such that many day-to-day tasks are one-shottable without supervision, then even if Cerebras can never run anything larger than roughly 100–200B parameters, it might have GPT-5.5-level intelligence “in probably less than a year.” The majority of people might no longer need frontier-level intelligence.
- Dylan says SemiAnalysis is “on the cusp” of being priced out of Mythos fast. Dylan floated Mythos at $25/$150 or $25/$125 versus Opus’s $5/$25; Doug called it roughly 5x, with fast mode 6x on top. Doug says current spend is justifiable, “but if you were to even double it… maybe we gotta turn off fast mode, guys”—“margins matter.”
- Doug’s key distinction on cost: “I don’t expect new models to be more expensive to do the same task… the problem of cost is you’re gonna do new tasks.” He also suggests Mythos fast may be cheaper than 4.6 fast for many tasks if it is more token-efficient. Max: “They call that Jevons, dude”—and the sad future is having to ask, “Is this task really worth Mythos fast token pricing?”
- The specimens: Max burned ~$400 of Opus 4.6 fast setting up a benchmark on a DigitalOcean droplet; Doug burned ~$800 in tokens scraping 10,000 employee profiles that a data-enrichment API would deliver for $55–100—“I just tried to boil the ocean using AI.” The “Rick and Morty” coda: “What’s your purpose?” “Pass me the butter.”
4. 4.7 versus 4.6 is a wash—and benchmarks can’t adjudicate it
- Doug’s vibes: instruction-following is “objectively worse,” missing CLAUDE.md instructions and skills, and the model keeps signing off—“We’ve done a lot for today, go enjoy your weekend.” He was using it on a Monday and said, “Get back to work.” His theory: inference optimization plus 10–100x more users degrades the experience—“the 4.6 golden age when it wasn’t quantized… pre-nerf. Those were the days.” His fix: “I need a NIMBY frontier model.”
- Doug on why benchmarks can’t settle it: frontier proximity is necessary, but “being number one on the benchmark ranking does not necessarily imply you’re the best model”—they’re now “a vibe check to make sure that the model’s not total trash.” HLE is “the most esoteric multiple choice questions you’ve ever seen”; SWE-bench scrapes GitHub issues that aren’t well-scoped tasks, with unit tests demanding specific 20-word error messages never mentioned in the prompt.
- The 4.7 changes Jordan catalogs: extra-high reasoning between high and max, high-resolution image support, thinking hidden by default, task budgets, and a tokenizer with 35% more vocabulary—potentially 35% higher cost for identical output. Dylan, waking from an on-air nap, counters that a bigger vocabulary should compress output; Jordan concedes conceptually but says in practice the model is less token-efficient—and if 4.7 isn’t clearly better, why the new tokenizer at all?
- Doug: “I don’t think Anthropic really does half-baked models. OpenAI clearly does”—5.3 Codex was RL’d only on code. Truther corner: “Opus 4.7 is actually Sonnet, dude,” with Dylan adding “and Opus 4.7 is Mitas”—“this model smells small.”
5. DeepSeek V4 and China’s compute wall
- The show’s bait question is whether the open-source gap between China and the US is widening again because of compute constraints. Dylan’s complete answer: “Yes.”
- The transcript highlights V4’s 1M-token context, versus Kimi’s 256K—a potential edge for long-horizon agentic work—and the “Reasoning in Visual Space” repo, posted then pulled, as a possible sign of a multimodal version. The same speaker admits that, unlike the original DeepSeek, which they used “all the time for random stuff,” they do not really use this one.
- Doug says there was no ether moment this time: “It comes out, it’s just state-of-the-open-source art,” not clearly better than Kimi K2.6. The real news is inference optimization: Ascend kernels being partially able to run inference on it “would really unlock more compute for China for the first time,” and the weights conveniently fit an 8x H200 pod’s memory domain—nothing bigger is served at the state of the art. “It just clearly feels like they are starting to hit some kind of wall.”
- Jordan calls the DeepSeek engineering release fascinating for its attention variants and KV-cache compression, while Doug challenges the practical value of the longer context: even Opus’s 256K-to-1M context is “dogshit” in quality, and compaction is painful.
- Max on slope: DeepSeek/Kimi are “probably ahead” of Meta, Grok, Cursor and xAI today, but compute is a key input; Meta is signing monster deals and has overcome the fire-then-rehire overhang. He expects Meta “to pull away from all the Chinese guys” in H2’26 or H1’27. Distillation aside, Dylan says Mistral distills from the Chinese labs, not Anthropic.
6. CLI versus app: innovator’s dilemma or pure maxi vision
- Dylan’s hot take: Anthropic perfected the CLI into an innovator’s dilemma—“the CLI is a dead end, a foregone relic of H1’26 and H2’25”—while OpenAI’s app holds “the true vision of the agent orchestration platform,” including voice and multimodality. Max separately notes that the Codex app is adding generative UI.
- Max’s maximalist rebuttal: the operating system doesn’t need to exist—hardware, a terminal, a Claude API, and it builds the OS for you. “It’s CLI all the way down. Pure maxi vision.” He says Claude Code is “clearly just a CLI wrapper,” while Codex CLI is “clearly just an app wrapper.” Jordan, meanwhile, is a VS Code-plugin guy—“don’t accuse me of reading the code.”
- Dylan’s Tri Dao anecdote from his cracked-kernel group chat: Codex is dumb but implements; Claude will waffle on niche microarchitecture details. So “have Codex write it and then have Opus fix it. You can’t go the other way around.” Max: “Everyone else at the firm prefers the other way around, actually.” Dylan: “we’re not writing fucking Tri Dao kernels.”
7. Long context, compaction, and the fake-news frontier
- On compaction, Doug is categorical: “Compaction blows… fuck the compaction”—better to clear and restart. Jordan notes DeepSeek’s 3.2 paper found exactly that: past the context window, fully clearing beat even summarizing on their tested tasks.
- Why hasn’t anyone shipped Llama 4 Scout’s announced 10M context? Dylan: “What fucking data do I have that is useful for next-token generation from 1 million to 10 million context? It’s so little”—the same data drought that makes 250K–1M “trash anyways, even on Opus.”
- SubQ, the day’s viral startup, gets the treatment from Max: “extremely ultra-mega sus.” If a real KV-cache breakthrough existed, “memory stocks should be down a quadrillion percent today,” and “I just don’t think these guys are gonna be the guys to crack the single hardest problem in all of AI.” Dylan’s pattern-match: every Asia trip surfaces a KV-cache-reduction paper no American researcher has heard of; TurboQuant was fake news.
- Still, Jordan’s market read stands: “There’s more capital than there is opportunity.” Dylan thinks SubQ could raise $50M at a $1B valuation.
8. Closing take
- Jordan calls Claude Code “the inflection point in February 2026” and says Doug’s May victory lap is complete. Dylan hopes the next inflection point is more exciting.