[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Summary
- Ashvin Nair’s core thesis is that today’s RL is a “very peaky” instrument: it can “kill the training distribution completely” yet generalize weakly outside it. The near-term value pool therefore belongs to products that pull a worker’s real context into distribution—codebases, terminals, documents, Slack, accumulated experiments—and co-design the model around that workflow; raw model capacity may not be the bottleneck.
- Olympiad gold no longer maps cleanly to economic automation: Ashvin once thought IMO gold meant “we could all just go on vacation,” yet “life is still the same.” He and the hosts explain the mismatch through benchmark selection and community-level overfitting: a model can jump from ordinary developer competence to elite competition performance while still missing the context required for everyday jobs.
- Robotics remains earlier and less monetizable than software agents despite impressive laundry-folding demos. The host compared robotics to the “GPT-1 to GPT-2 era”; Ashvin emphasized hints of generalization but said robotics still feels more like an investment in a team than a proven technology. He expects LLM agents to become a trillion-dollar market before AI robotics is maybe even a $10 billion market, because physical systems must still prove usefulness, reliability, maintenance economics, and generalization.
- Reasoning RL arrived as a smooth internal scaling curve, not one miraculous release, once pretrained models became capable enough. A small 2023 prototype produced unusually accurate reasoning traces and surprisingly strong math scores on a small model—performance that otherwise would have required much more pretraining; by early 2024, Ashvin says the recipe made IMO and IOI wins predictable, while public “leaps” concealed stacked experiments and steady month-to-month gains.
- At OpenAI, the internal-to-public lead has shrunk from roughly six months to perhaps one or two, while labs converge on “similarish” RL recipes. Ashvin saw no major internal lesson from DeepSeek—OpenAI already had a better model and smarter models remained valuable—but noted that even Anthropic’s Opus 4.5 showed an ARC-AGI-2 plot resembling OpenAI’s.
- Cursor’s wedge is tight product-model co-design, embodied by online Tab policy updates about every two hours and a 20–25-person ML group behind Composer. Composer is “smart enough” to use but fast enough to avoid context switching; the larger ambition is to automate the whole software-engineering loop—write code, inspect Datadog, form hypotheses, rerun, and learn.
- The biggest prospective discontinuity is continual learning with human-like data efficiency, not another static benchmark win. Models repeat bugs even within one context, whereas a person can watch someone touch a hot stove once and learn. Ashvin calls the idea deeply interesting; the host suspects it could be paradigm-shifting within a year, but Ashvin explicitly says he has “no idea” what that shift might be.
- Governance remains unresolved whether AGI arrives in two years or ten. During OpenAI’s “blip,” Ashvin signed the employee letter but was willing to set equity aside for a real governance debate; he wonders whether broad public-company ownership could be more democratic than seven nonprofit directors, while conceding that capitalism already mishandles social media and unhealthy food.
Deep dive
1. Software agents should monetize long before robotics
Ashvin’s path from Berkeley robotics to OpenAI and Cursor felt less discontinuous than it sounds: both domains demand inspecting large amounts of messy data and persisting when systems refuse to work. Robotics builds “very gritty people” because, unlike simulation, they have no choice but to confront the real world.
Ashvin had not seen the Sunday robots himself, though he found the reported demos “kind of cool”; the host had seen Physical Intelligence robots live, folding laundry in an ordinary living room. Ashvin tied the potential inflection point to the kind of generalization associated with GPT-2, while saying the details matter and robotics still feels more like a bet on a team than a proven technology.
His market call is stark: “LLM agents are going to be like a trillion-dollar market before robotics is maybe even like a $10 billion market.” Agents already create value; robots still need useful tasks, reliable hardware, repairs, and viable unit economics.
2. Olympiad gold revealed how little benchmark supremacy guarantees
Ashvin once treated IMO gold as an end state: “I would have just assumed that we could all just go on vacation—AI is solved.” The result arrived, yet ordinary life barely changed. Chess and Go produced the same surprise, but the mismatch remains jarring each time.
The hosts’ explanation was that people move the AGI goalposts; Ashvin partly defended that move because the community collectively optimizes whatever benchmark it chooses. The host juxtaposed this with language models that appear at roughly junior-to-senior developer level on normal work yet can win elite programming competitions; Ashvin called that mismatch suspicious.
His own 2017–22 RL field supplied the warning. Academic work appeared to advance through off-policy learning, value functions, and new mathematical machinery, while an RL winter saw entire startups founded on the premise later give up. In retrospect, researchers had created knobs and implicitly tuned them to shared benchmarks at community scale.
Academia compounded the problem by rewarding mathematically interesting novelty over “simple ideas that work.” Methods that actually work tend to be simple, have fewer knobs, and rely more on compute, but also offer less publishable “secret sauce”—a poor incentive structure for discovering what generalizes.
3. Reasoning RL worked when pretraining crossed the right threshold
Ashvin credits a long OpenAI lineage, including Ilya Sutskever and Jakub Pachocki, with having “AGI in their bones.” Dota already contained the template: copying the internet would eventually plateau, while RL could generate better intelligence. RLHF was a limited side branch because human feedback could not absorb comparable compute.
The turning point came around 2023. Running RL on even a small model produced reasoning traces that were unusually accurate and surprisingly strong math scores—performance that otherwise would have required much more pretraining. Like GPT-2 before GPT-3, the prototype demanded first-principles conviction before its raw performance made the opportunity obvious.
Progress then felt continuous inside OpenAI: experiments produced incremental gains, inconclusive ideas were stacked, and resources followed the new line. By early 2024, Ashvin thought it predictable that the recipe could “really smash” IMO or IOI, even though external observers experienced releases as abrupt breakthroughs.
The host cited roughly 300 people around reasoning versus a dozen faces in the original o1 video. Ashvin resisted that accounting: the early contributor set was probably 50–200, and as o3 became a product, safety, evaluation, and other functions expanded participation until he had “lost track of the numbers.”
4. Job automation requires putting the entire workflow in distribution
Scaling is not over, but Ashvin thinks its shape has changed. Current RL generalizes somewhat and in interesting ways, yet remains “very peaky”: it can dominate its training distribution with modest effort without transferring broadly enough to automate work.
The requirement is therefore to bring economically useful tasks into the training distribution. GDPval has roughly the right form: it covers 128 tasks across white-collar jobs accounting for more than 5% of GDP and tries to stay close to source documents rather than sanitized model inputs. Ashvin had not examined its traces closely enough to know what an accountant’s job and required context actually entail, underscoring that the product must expose the real workflow, not just an eval.
Ashvin’s OpenAI research job illustrates the missing context. He wrote relatively little code while spending a year running sweeps, studying hyperparameter interactions, and accumulating knowledge across graphs that mostly “were just sitting in my head.” A coding model could write the scripts yet remain unable to reproduce the job without that accumulated context.
The host read separate GPT-5 and GPT-5 Codex lines as evidence that “one model fits all” was dying. Ashvin’s pushback: “OpenAI has a tendency to shift the org chart.” Specialization may reflect which organization owns the data, not model capacity; with all relevant data, joint training might still produce useful cross-domain generalization.
5. Frontier competition compresses both forecasts and release windows
At the Curve conference, forecasters expected only 10–20% on Epoch AI’s FrontierMath and Humanity’s Last Exam around 2027, while Ashvin had already seen internal models exceeding those estimates. Some of the same people contemplated Dyson spheres by 2035: potentially “too pessimistic in the short term, too optimistic in the long term.”
He nevertheless respects that community for registering predictions rather than claiming afterward, “I saw this the whole time.” Its older capability forecasts were directionally stronger than the prevailing view that AI was a sham, and Ashvin put human-level intelligence somewhere in the “2030-ish” range.
At OpenAI, the host said internal models had once been about six months ahead of public releases; Ashvin estimated the current lead time at roughly one to two months. The host invoked Nano Banana Pro as an example of how quickly a lead can matter, while Ashvin emphasized that the release window is now “tiny.”
DeepSeek surprised Ashvin more through its market impact than its technical message: he said it showed NVIDIA chips were more useful than previously thought, yet NVIDIA’s stock fell. OpenAI already had a better model, and labs soon converged on similar RL forms; Anthropic’s Opus 4.5 even displayed an ARC-AGI-2 plot resembling OpenAI’s.
6. OpenAI’s crisis left the central governance question unanswered
Ashvin learned of Sam Altman’s firing during Thanksgiving with two OpenAI friends and initially took it as a joke. After “a crazy weekend of just ups and downs,” he joined roughly 95% of people in signing the letter—but his reasoning was more conflicted than simple loyalty.
He considers governance important whether AGI is two years away, ten years away, or further. During the crisis he was prepared to “forget about the equity” and debate the structure seriously; he also wondered whether a Microsoft-style board, whose stakeholders might include the public through pensions, could be more democratic than seven people running a nonprofit.
When the host later asked whether OpenAI’s nonprofit had a better “secret shadow board” structure for determining AGI, Ashvin had no answer. His honest conclusion was that society has “not solved governance at all.” If capitalistic incentives already fail to produce healthy outcomes around food and social media, there is little basis for confidence in AGI governance.
7. Cursor is betting that product-model proximity beats laboratory scale
The host’s pushback was direct: OpenAI has effectively unlimited resources, abundant data, and Codex, so why leave? Ashvin’s answer was organizational. Cursor offered a small, focused environment where product and ML teams sit together and can deliberately pull the product’s test distribution into RL training.
Online Tab is the clearest specimen: Cursor can update its policy about every two hours. Ashvin disputed that this was merely easier because autocomplete uses a smaller model; the enabling factor was a compact organization able to connect user behavior, product decisions, and training rapidly.
Cursor’s ML group is only 20–25 people, and Ashvin was pleasantly surprised by Composer’s quality. Its appeal is not just intelligence but latency: it is “smart enough” that users want it, yet fast enough for them to stay in the loop. Slower smart models induce context switching that “kind of gives you ADHD.”
The destination is broader than answering prompts. Cursor wants the model to perform software engineering as a process: write code, inspect Datadog, diagnose behavior, form a hypothesis, rerun the system, and iterate. Internal tooling lets researchers SSH into user environments and stay close to the data—the durable advantage Ashvin sees in keeping model and product together.
8. Continual learning may be the next genuine paradigm shift
The host challenged naïve online learning: ingesting every user action could pull a model toward mediocre behavior. Ashvin’s hot-stove analogy reframed the problem. Humans do not merely filter out a bad example; they presumably have a value function that makes them avoid repeating it after one observation.
Models remain “a few orders of magnitude” behind that data efficiency. They may introduce the same code bug repeatedly—even within one context—whereas a person should make the mistake once and retain the lesson across contexts. Ashvin floated “in-context learning with infinite memory or something,” where an experience in context should become part of the weights so the model does not repeat it.
He sees little immediate capacity problem: deployment may add thousands or perhaps millions of tokens to a model originally trained on trillions, “a drop in the water bucket.” The deeper uncertainty is whether weights behave primarily like a hard drive storing facts or a CPU whose reusable circuits compute them.
The host suspects continual learning could be “paradigm-shifting in the next year or something”; Ashvin agreed there was something deeply interesting there but said he had no idea what the shift might be. Cursor is hiring especially for code data and rewards, and favors two-day work trials over trivia, though “why is off-policy RL unstable?” remains his revealing interview question.