Pioneers Insight Method Research Author
48. A Conversation with Former OpenAI Scientists: GPT-5 Can Win a Math Olympiad Gold Medal, but That May Be Deceptive
Back to Episodes

48. A Conversation with Former OpenAI Scientists: GPT-5 Can Win a Math Olympiad Gold Medal, but That May Be Deceptive

Summary

  • Former OpenAI scientists Kenneth Stanley and Joel Lehman believe GPT-5 may show that returns from the scaling path are slowing—and that this is precisely the signal that research is becoming interesting again. Joel says bluntly that “the jump from GPT-3 to GPT-4 looks more profound than the jump from GPT-4 to GPT-5,” and even hopes the path is “running out of steam”; Ken’s framework is that “every limitation is an opportunity, because it means someone can come along and disrupt it”—if scaling alone cannot reach AGI, this is the window for new ideas.
  • The two challenge benchmark-driven AI competition head-on: a model can win a math Olympiad gold medal, “but that may be deceptive.” Ken asks: if leaders at every lab say their models are already “PhD level,” “where is the new math? Why hasn’t it invented a lot of new mathematics?” Every new model is setting records across every benchmark, “which is starting to make these benchmarks look gameable”; the field may be following a path in which the objective is fooling it.
  • The most transferable mechanism for investors is this: one big win causes organizations to converge and put all their resources behind what appears to be the winner, while the path to the last innovation “will not work twice.” Ken cites Kodak’s disruption by digital photography; Joel argues that Google identified the foundation-model direction early but missed the GPT revolution and lost AI talent, making it vulnerable to being consumed by its own success, although Ken adds that Google has regained its footing to some extent. Joel’s advice for screening startups is to look for teams built around interesting people and ideas, and assess whether the potential innovation and novelty are enough to justify the risk.
  • OpenAI itself is the best proof of “innovation without a goal”: ChatGPT was an accidental project, and looking back at 2019, Ken says “you couldn’t see what the business plan was,” with success coming from opportunistically following interesting directions. Betting on language models in 2017 did not violate the book’s argument: it was conviction, not a goal—“GPT-1 was garbage, but it was interesting,” and the people willing to keep doubling down at that stage “were very smart.” Ken also stresses that when he joined, OpenAI had not fully converged on language models; he can only partly speculate about whether such a concentrated decision was made in 2017.
  • The two experienced OpenAI’s gear shift from pure exploration to exploit mode firsthand, and Joel is “a little sad” about it: the field used to be “more playful,” whereas now “there are just so many language-model papers.” Ken’s assessment of Sam Altman is limited to 2020-2022: “a good leader—pragmatic, cautious and communicative.” On the later power struggle, he says he was as shocked as everyone else and still does not understand exactly what happened.
  • At Lila Sciences, where both recently joined, they are betting on “scientific superintelligence”: scientific revolutions will not emerge from internet data alone but require interaction with the real world and hypothesis testing. They have launched an open-endedness team there, viewing “science itself as the most open-ended endeavor—a tree of discovery that never ends.” Ken also notes that coding models in 2025 brought enormous acceleration, but a new drag: debugging is harder when you did not write the code yourself.
  • DeepSeek drew enormous attention, but the trajectory of the US-China AI competition remains unpredictable. Ken calls its achievement a shock and says international events sometimes push the two countries apart and sometimes bring them closer. His warning to China points directly to the book’s title: China is good at planning and benefits from it, but “planning may be harmful, especially in disruptive innovation”; he also sees high-pressure exam culture as a shared US-China problem. On predictions of “AI hell mode” beginning in 2027 and the middle class disappearing, Ken says forecasts made during periods of upheaval may be unreliable—risk and enormous upside coexist, and social systems may not be prepared for AGI.

Deep dive

1. The Picbreeder Paradox: The Best Way to Find Something Great Is Not to Look for It

  • Ken explains the book’s origin: Picbreeder was an online experiment in which users “bred images,” originally designed to study open-ended systems—“human civilization itself is an open-ended system.” The experiment revealed a disruptive phenomenon: if you start out looking for a bird, you will fail. “The images that lead to a butterfly don’t look anything like a butterfly. You have to not think about butterflies in order to select the images that lead to one.” This ran against everything he had learned in engineering; “set a goal and move toward it” simply did not work.
  • The system succeeded not by reaching a target but by accumulating diverse stepping stones. Joel turned that insight into the novelty search algorithm: if you do not tell a robot where the exit to a maze is, it can actually navigate the maze more effectively.
  • The decision to write the book came after a talk Joel gave at the Rhode Island School of Design. Art students became “so emotionally moved they were almost crying,” telling him, “This is the first time I’ve been able to prove to my parents and teachers why I chose to live this way—because I’m following the interesting path.” Ken heard about it and said, “If people are going to cry over this, it must be important.”
  • ChatGPT fits the same narrative. Ken believes it qualifies as a form of greatness in terms of its global impact, although “greatness” has no single definition; even something that profoundly affects only 2 people might be called great.

2. Inside OpenAI, 2020-2022: The Gear Shift from Pure Exploration to Exploit Mode

  • Ken’s first-hand assessment of Sam Altman is explicitly limited to the period before the later power struggle: “He was a good leader—pragmatic, cautious and communicative. He weighed issues very carefully, and at the time I couldn’t see much to criticize.” On the subsequent power struggle, Ken says he was as shocked as everyone else and still does not know what happened in detail; Joel describes it as a series of extraordinary events that were difficult to make sense of.
  • As the product’s impact became visible and continued to expand, Joel says OpenAI’s atmosphere gradually shifted from pure exploration toward a more deliberate exploit mode. The competitive dynamic followed: larger datasets, more GPUs and other teams joining the race. Ken says, “When you strike gold, you naturally focus,” but the most dramatic change may have come after they left, when ChatGPT truly changed the company.
  • OpenAI initially looked more like a research lab, but its applications and commercial functions expanded rapidly, bringing a more pronounced commercial sensibility inside the company. Asked whether the company in 2023 had to choose between being a research institution and a technology giant, Ken can only speculate: it may have wanted both, but a huge commercial opportunity inevitably affects company culture.
  • On how to identify an innovative startup, Joel says the outcome is very difficult to predict in advance, but interesting companies are usually built around interesting people and interesting ideas. For employees, the question is whether the work matches their own interests; for investors, it is whether the innovation, novelty and potential upside after success are enough to make the risk worthwhile.
  • Did they regret OpenAI’s transformation? The two disagree. Ken does not want to say he regrets it, arguing that the entire AI field changed: a relatively small discipline with enormous social impact suddenly became “the most important thing in the world,” and that transition was bound to be complicated. Joel is “a little sad,” missing the more diverse and “playful” research environment of the past. He acknowledges that scaling is important science, but simply making things bigger and bigger is less intellectually compelling to him than foundational research.

3. GPT-5: Slowing Down Is Not Bad News—It Is When Research Gets Interesting Again

  • Ken’s core view is that the path of the past few years has been scaling plus data. “If that path alone doesn’t get us to the holy grail—what some people call AGI—that wouldn’t be surprising. Intelligence is not only about this one thing.” He treats limitations as opportunities: “If things start slowing down, there’s room for more interesting ideas to come in and do something new.” The past few years have been exciting from a product perspective but increasingly unexciting as research, because everyone has been following the same established route.
  • Joel puts the comparison more directly: “The jump from GPT-3 to GPT-4 was more profound than the jump from GPT-4 to GPT-5. I don’t think that should be controversial.” He retains some uncertainty—“we don’t know whether this paradigm is at a plateau or still has juice”—but his personal position is clear: “I hope it’s running out of steam.” He suspects the way these models are trained differs from how humans understand things, and does not believe the transformer will be the endpoint of the story.

4. The Deceptive Goal: It Won the Olympiad Gold—Where Is the New Math?

  • Ken maps the book’s idea of “goal deception” directly onto today’s AI race: everything is organized around math and programming benchmarks. “We may be seeing scores rise continuously without any real improvement in the intelligence we understand. When we become obsessed with these metrics, we lose the big picture—and losing the big picture is itself deceptive.”
  • The sharpest point is this: leaders at every lab like to say their models are “PhD level,” and the models can even win gold medals at math Olympiads—but “where is the new math? Why hasn’t it invented a lot of new mathematics?” One possibility is that getting better and better on these tests will not lead to new mathematics because the path itself is deceptive. “Every time a new model comes out, it leads on every benchmark,” which is starting to make the benchmarks look gameable; people may be competing for the benchmarks rather than for intelligence.
  • Joel points to history for support: neural networks went through a winter and a revival after debates such as those surrounding Perceptrons 20 or 30 years ago, and AI research has repeatedly risen and fallen. “If AGI is just around the corner, I don’t know whether we’re ready, but we may be charging toward it without looking back.” He agrees that the risk of deception is substantial.

5. Why Big Companies Kill Innovation: The Self-Reinforcing OKR Loop and the Cultural Risks of Innovation Labs

  • Ken’s diagnosis of the mature-company trap begins with goal management. Google was one of the major promoters of OKRs and KPIs: “Goals brought success, so impose more goals.” That creates a self-reinforcing loop. The deeper problem is that “the path that led you to innovation will not work twice”; the heuristics behind the last success may not apply to the next innovation, and every innovator can fall into this trap.
  • On Google, Joel says it identified the foundation-model direction relatively early but missed the GPT revolution and lost a large amount of AI talent. As a bigger company, it may have been more cautious about deploying the technology and more vulnerable to being consumed by its own success and bureaucracy. Ken adds that Google has regained its footing to some extent, now competing directly with OpenAI and Anthropic, while DeepMind has also delivered important results outside language models, including AlphaFold.
  • Why are internal innovation labs so difficult to make work? Researchers are not fooled by slogans. When they see colleagues promoted for helping the company’s bottom line, they respond to the signal—“because I want this year’s bonus too.” People outside the lab become jealous and angry: “Why do those people get paid to play around like children in a playground?” Ken’s paradoxical conclusion is that “doing useless things is exactly what is best for innovation.” Protecting a real innovation lab requires courage from leadership.
  • On who pays for open-ended exploration, Ken says university basic research is usually publicly funded, with long-term returns that may exceed the investment. Corporate environments face profit timelines and survival pressures, making completely unconstrained exploration harder to sustain. But as long as a company is healthy, success should not be an obstacle to open-ended exploration; even small companies can innovate in open-ended ways. A society under too much pressure to take risks will not innovate. Innovation begins when people have enough slack to experiment and fail.
  • Ken stresses that this is not an argument against all planning. In disruptive innovation and similar areas, excessive goal orientation can be harmful. For an innovation lab to fulfill its purpose, leadership must have the courage to protect exploration that appears useless for the time being.

6. Convergence Is Not Betraying Openness: The Difference Between Conviction and Goals, and Lila Sciences’ New Bet

  • The host frames the question by saying that OpenAI was running out of money in 2017 and therefore concentrated its resources on large language models. Ken does not fully confirm that account: when they joined, the organization was still working on more than language models. Someone may indeed have decided that language models were the path forward, but he can only partly speculate.
  • That does not conflict with the book’s argument. Ken says the book never rejects conviction: if you strongly believe a path is interesting, you should pursue it, but not merely because it appears to lead to some distant goal. Even if someone decided to invest in language models in 2017, they could not have anticipated ChatGPT 6 or 7 years later. GPT-1 “was garbage”; it had no real product value, but it was interesting. In hindsight, the people willing to keep doubling down at that stage made an extremely smart decision.
  • Joel offers a neat formulation: you can “be divergent by converging.” Early OpenAI converged intensely on the hypothesis that scaling neural networks would produce interesting results, but that high conviction could itself become a stepping stone along an exploratory path.
  • Asked which company most resembles OpenAI in its early days, the two point, with an admitted bias, to their new employer, Lila Sciences. Its goal is “scientific superintelligence,” built around 2 core assumptions: scientific revolutions will not emerge from internet data alone but require interaction with the real world, hypothesis testing and the acquisition of more information; and science itself is the most open-ended endeavor, a tree of discovery that never ends. Ken says launching an open-endedness team there reminds him of his time at OpenAI. Joel concedes that the next big event is always invisible before it becomes visible: “If I knew what it was, I would tell you.”
  • On AI’s development in 2025, Ken still sees progress moving quickly, particularly in the practical impact of coding models. They bring enormous acceleration but also a new form of deceleration: because the code was not written by you, debugging a problem may first require understanding what the model did, making the process more time-consuming than before.

7. DeepSeek, China’s Planning Culture and the “Hell Mode” Forecast

  • Ken says DeepSeek’s impact is difficult to ignore: “The dramatic shock to the stock market” was itself a clear signal. He calls it an impressive achievement and another important stage in the AI race. The host says the episode pushed Silicon Valley to pay closer attention to Chinese AI companies and changed many competitive dynamics; Ken responds that US-China technology competition is already tense but highly unpredictable, with some events pushing the countries apart and others unexpectedly bringing them closer.
  • On China’s planning culture, Ken begins with self-deprecation: “If China really likes planning, and our book is called Why Greatness Cannot Be Planned,” the answer may seem obvious. But he immediately qualifies the point: he is not against all planning; in disruptive innovation and similar areas, planning can be harmful. On high-pressure exams, he notes that both the US and China have exam cultures: test scores may measure the ability to focus on tests rather than the potential required for real-world innovation. He hopes both countries can reduce the heavy, tedious testing burden placed on children.
  • On the forecast that a 15-year “hell mode” begins in 2027, the middle class disappears, and society is left with 0.1% of rich people and 99.9% of the poorest, Ken says predictions made during periods of upheaval may not be reliable. Employment, social structure, the economy, physical safety and well-being all face risks, but scientific progress could also bring enormous benefits; the destruction of one industry often accompanies the creation of another. In theory, if AGI performs much of people’s work, they could rest, but society’s systems may not be prepared for that outcome.
  • A personal postscript closes the conversation. The host says Ken appeared frustrated and confused when he left OpenAI, and Ken confirms it: he had only just begun to grasp AI’s world-scale impact. After thinking about it for several years, he at least understands more of what is happening. An unexpected benefit of writing the book was meeting people from many different backgrounds; one grandmother even called his office to ask whether he could meet her grandson and tell him not to be so goal-oriented. Ken’s own stepping stones include encountering computers through BASIC programs in the magazine 3-2-1 Contact as a child, and later deciding to recruit Joel—a choice that led to many consequences he could never have predicted, including this interview.