Pioneers Insight Method Research Author
GPT-5 Arrives, and We Try the New Alexa+
Back to Episodes

GPT-5 Arrives, and We Try the New Alexa+

Summary

  • GPT-5 looks more like a distribution breakthrough than a dramatic expansion of the AI frontier. Free ChatGPT users gain access to OpenAI’s flagship model and reasoning capability, while a router chooses how much compute each request receives. Kevin Roose’s summary: OpenAI may not have “raised the ceiling” much, but it has “raised the floor.”

  • Aggressive API pricing turns GPT-5 into a direct margin challenge for rival model providers. At $1.25 per 1 million input tokens—the same as Gemini 2.5 Pro—it sharply undercuts Claude Opus 4 at $15. Roose compared the market to the era of “$4” Uber rides: heavily capitalized companies are subsidizing artificially cheap AI tokens to win demand.

  • The first hours of use suggested a faster, more polished ChatGPT rather than the revolutionary system implied by the GPT-5 name. Casey Newton found it completed editing work much faster than o3 at similar quality, but screenshots of prediction markets reportedly swung from OpenAI toward Google for the best model at August’s end. The cutting verdict: “My timelines are now longer.”

  • OpenAI says scaling still works, but the mechanism increasingly includes post-training reasoning rather than simply making pre-training runs larger. Sam Altman said the laws “absolutely still hold” and that OpenAI keeps finding “new dimensions to scale on”; Roose believed him, while noting that a really big training run did not visibly produce superintelligence. The “software-on-demand” demos looked good, yet mostly resembled tasks competing models could already perform.

  • Reliability gains could matter more commercially than benchmark leadership, though the hosts immediately found reasons for caution. OpenAI reported hallucination rates around 1% for some question types, 5,000 hours of red teaming and new “safe completions,” but Newton caught GPT-5 hallucinating within hours. His operating rule remained: “Don’t trust these things for anything mission critical.”

  • Alexa+ demonstrates both the strategic promise and the integration risk of putting generative AI inside a mature consumer product. Its voice is more natural, conversations persist without repeating the wake word, and Roose successfully ordered an Uber; yet alarms, document ingestion, recipe navigation and task routing failed during testing. His verdict was “two steps forward, one step back,” while Newton concluded that even “the finest minds in the world” do not yet know how to build this reliably.

  • Amazon is attacking that integration problem with unusual architectural and organizational scale. Alexa+ uses more than 70 models, over 80% of its main inference traffic runs through Amazon Nova, and thousands of employees support hardware, software and integrations with millions of existing capabilities. Daniel Rausch said advertising is “probably the smallest part” of the business plan; the larger objective is making Alexa+ a Prime benefit that strengthens usage, retention and Amazon’s broader flywheel.

Deep dive

1. GPT-5 arrived as the centerpiece of a crowded release week

  • The surrounding releases established how compressed the frontier has become: Google DeepMind demonstrated the interactive world model Genie 3, Anthropic shipped Opus 4.1, and OpenAI released its first open-source models since GPT-2. One reportedly runs on a MacBook; the larger one requires a dedicated GPU.

  • Newton said early reviews placed those open-source models near proprietary o3-mini and o4-mini performance. Roose framed them partly as an answer to criticism that OpenAI had abandoned its founding spirit, and partly as competition for Chinese open-source systems such as DeepSeek.

  • The hosts disclosed that The New York Times Company is suing OpenAI and Microsoft for alleged copyright violations, while Newton’s boyfriend works at Anthropic. Their initial GPT-5 assessment came from an OpenAI briefing before either had meaningfully tested the released model.

2. OpenAI sold GPT-5 as expert intelligence, but stopped short of AGI

  • Sam Altman called GPT-5 a “significant step along the path to AGI,” while explicitly saying it was not AGI. His concrete dividing line was continuous learning: GPT-5 does not do it, and he believes an artificial general intelligence will.

  • Altman’s capability ladder moved from GPT-3 as a high-school student, through GPT-4 as a college student, to GPT-5 as a PhD-level expert. After returning to GPT-4 during development, he reportedly found the experience “quite miserable” and said he never wanted to go back.

  • Roose kept the claims provisional because the model had not yet rolled out during the briefing. What was already consequential was distribution: free ChatGPT users—including large numbers of students—would receive access to OpenAI’s claimed PhD-level intelligence and reasoning rather than being confined to weaker default models.

3. The router raises the floor while creating a new compute conflict

  • OpenAI is doing away with the model picker, or at least making it less necessary, by using a router that evaluates each request and assigns an appropriate model and compute budget. For many free users, this will be their first routine encounter with a reasoning model.

  • Newton saw clear usability value: people no longer need to know whether a question warrants GPT-4o, o3 or another option. But he also identified a conflict—OpenAI bears the inference cost and therefore has an incentive to route users to “the absolute least compute” it believes a task requires.

  • Roose’s synthesis became the release’s sharpest framing: GPT-5 may not have raised the frontier’s ceiling very far, but it “raised the floor” by moving free users onto stronger systems. That distribution shift could change public perceptions even without a spectacular new capability.

4. “Software-on-demand” looked useful, not unprecedented

  • OpenAI demonstrated GPT-5 building a French-learning tool from a text prompt. It generated flashcards and a small game in which a mouse collected cheese and displayed new vocabulary—work Newton thought might have earned an A in an introductory programming course five years earlier.

  • Roose found the result impressive but noted that current rival models can already build similar applications. His preferred test for a new release is no longer its benchmark score but “what is possible for me now that wasn’t before?” The briefing did not give him a satisfying answer.

  • His fatigue was explicit: after many nearly identical launches promising better coding and “agentic capabilities,” he briefly zoned out because the presentations increasingly sound like marketing. GPT-5 would have to prove itself through actual use and his informal “RooseBench,” not curated demonstrations.

  • Newton’s first test was a Fantastic Four-styled to-do application. Roose’s objection captured the mismatch between infrastructure spending and marginal consumer utility: “I can’t believe we’re building giant gigawatt data centers for your stupid to-do apps.”

5. Scaling may continue, but the definition of scaling is broadening

  • Asked whether the industry was reaching limits, Altman answered categorically: scaling laws “absolutely still hold,” and OpenAI keeps discovering “new dimensions to scale on” and new paradigms for improvement. Newton expected users to challenge that confidence if GPT-5 merely performed old tasks somewhat better.

  • Roose suspected the claim now incorporates post-training reinforcement learning and reasoning environments, not just larger pre-training runs. He has heard from people who believe that approach has considerable room remaining, even if pre-training alone may be nearing diminishing returns.

  • OpenAI did not disclose the training data, GPU count or model size. Roose assumed it had pushed pre-training scale as far as practical, yet the demonstrations did not reveal something “that much smarter” or anything resembling a superintelligence emerging from the box.

6. The GPT-5 label is itself a capability claim

  • Newton questioned what makes a bundle of models and systems worthy of the integer “5.” The large-number releases carry exceptional marketing power because GPT-2 to GPT-3 and GPT-3 to GPT-4 created expectations of similarly visible leaps.

  • Roose said labs routinely rename disappointing training runs. OpenAI’s GPT-4.5, he believed, was at one point intended to become GPT-5 but did not perform as hoped; calling this system GPT-5 therefore signals that OpenAI wants it judged as the next major capability step.

  • That signal heightened the mixed day-one reaction. Online screenshots showed prediction markets moving rapidly from OpenAI toward Google as the likely provider of August’s strongest model, suggesting that a large constituency had expected a revolution and instead saw evolution.

7. Speed and price may be GPT-5’s most immediate competitive weapons

  • After several hours, Newton called GPT-5 a meaningful improvement to ChatGPT, especially for free users tackling extended problems. In his own editing workflow, it “blazed through” work faster than o3 while producing output he considered equally good.

  • GPT-5’s API price landed at $1.25 per 1 million input tokens, matching Gemini 2.5 Pro and far below Claude Opus 4 at $15. Newton read both OpenAI’s and Google’s pricing as an effort by well-capitalized labs to apply pressure to competitors.

  • Roose compared the moment to venture-capital-subsidized ride-hailing a decade earlier, when an Uber could cost $4. AI developers are receiving similarly artificial token economics, with the durability of those prices left unresolved.

  • OpenAI’s execution still impressed Roose: despite becoming a large organization, enduring board turmoil and juggling competing teams, it appears to be accelerating toward AGI. Newton’s caveat was talent poaching—recent departures may only reveal their effect on iteration speed over the coming months.

8. Safety improved on paper, but verification remains mandatory

  • OpenAI said GPT-5 hallucinates less, is less deceptive and can use “safe completions”: rather than merely refusing a problematic request, it attempts a safer version. Roose highlighted reported hallucination rates near 1% for some question categories but did not trust the benchmarks without personal testing.

  • Newton had already caught the model hallucinating several times. His practical conclusion remained unchanged: users should double-check facts and avoid relying on any model for mission-critical work, regardless of aggregate benchmark improvements.

  • Nick Turley said OpenAI is consulting physicians and outside experts about intense user relationships and sycophancy, and is “absolutely not optimizing for engagement.” The company wants a useful system that completes a task and “sends you on your way,” though the briefing supplied limited detail.

  • OpenAI reported 5,000 hours of red teaming and said it had shared the models with some external experts for advice. It also rated GPT-5 high for potential creation of novel biological risks and added protections accordingly—a disclosure Newton summarized with understated concern: “That doesn’t seem great.”

9. Alexa+ makes the old assistant conversational—and the old basics fragile

  • Alexa launched in 2014, yet both hosts still use it almost exclusively for three jobs: timers, music and weather. Generative AI promised to extend that narrow utility into open-ended questions, personalized conversations, shopping, smart-home control, restaurant bookings and transportation.

  • Amazon said Alexa+ had reached 1 million users by June 23 and requires newer Echo hardware. The hosts also disclosed that The Times had agreed to a licensing deal giving Amazon access to Times content for AI products including Alexa, while speculating that Anthropic’s Claude may contribute through Amazon’s model ecosystem.

  • The early gains were tangible: a more human-sounding voice with eight options, follow-up conversations without repeating “Alexa,” longer stories, recipe help and multi-step requests. Roose successfully ordered an Uber and asked Alexa to find Wirecutter’s best-rated box grater and add it to his Amazon cart.

  • Roose did not complete an OpenTable booking, but Alexa surfaced relevant options. The experience hinted at a useful ambient agent: state a destination or dinner requirement aloud, let the system gather choices, and confirm the transaction without opening multiple apps.

10. Alexa+ failed where hardware leaves little room for probabilistic behavior

  • Newton’s Echo Show 5 experience began with a $90 screen that displayed art for perhaps “four seconds per minute” before pitching aspirin or paper towels. Amazon then sent him an Echo Show 15, a 15-inch device meant to be wall-mounted; he left it sitting on his desk rather than undertaking the installation. He concluded that the hardware felt like “little windows” for sending money to Amazon.

  • The functional breaking point was a displayed lemon-pasta recipe. Alexa could show the item but could not respond when Newton asked to open “the lemon pasta right there”; intermittent device connectivity may have contributed, so he carefully stopped short of blaming the AI alone.

  • Roose experienced latency, a hallucinated answer about a tennis tournament’s top seed, a research request incorrectly routed into immediate Spotify playback, and a document-ingestion feature that denied receiving the paper he had emailed. Most damagingly, Alexa failed to cancel an alarm—a command he had issued roughly 1,000 times before.

  • Because Alexa+ remains in early access and warns that it may make mistakes, Newton characterized his assessment as a first impression rather than a full review. Roose’s assessment was that Alexa+ feels like putting a “GPT-3.5-class model inside of a smart speaker”: valuable enough to keep developing, but neither state-of-the-art as a language model nor reliable at basic assistant work. “Two steps forward, one step back.”

11. Amazon’s real challenge is translating stochastic language into deterministic action

  • Rausch said Alexa’s AI and model layer was “entirely new,” with some legacy deterministic systems remaining downstream. Natural-language models produce elegant conversation, but APIs use “clunky computer science language”; reliably translating intent into commands across millions of established capabilities is the central engineering problem.

  • Alexa+ uses more than 70 specialized models, central systems that maintain conversational context, and different training corpora for different tasks. Rausch estimated that over 80% of major inference traffic flows through Amazon Nova models, giving Amazon greater control over training, tuning and post-training.

  • The scope includes tens of thousands of integrated services and devices, thousands of employees, and millions of things the original Alexa could do. Rausch said the 2023 “Let’s Chat” concept was too modest; limited instruction-following and reasoning forced Amazon to expand the vision, ground answers in authoritative sources and build deeper personalization.

  • On organization, Rausch declined to comment on a former scientist’s claim of “technical and bureaucratic problems,” but said Alexa is undergoing a startup-culture transformation. He acknowledged that its rate of innovation had slowed, while saying the team is now executing at an “unbelievable pace.”

  • The business model centers on Prime more than ads, which Rausch called “probably the smallest part” of the plan. He argued that Alexa+ unifies Music, Video, Photos and other benefits, strengthening Prime’s flywheel; the hosts’ paper-towel experience showed the risk that cross-selling can instead make the hardware feel invasive.