Pioneers Insight Method Research Author
Back to Pioneers
Radically Better Reasoning
Innovators 1 Curated Dialogues

Radically Better Reasoning

Key Views & Dialogues

Radically Better Reasoning: Elicit’s Andreas Stuhlmüller & Jungwon Byun on World Models for Research

  • 🗓️ Date2026-06-17 | 🎙️ Show:The Cognitive Revolution

Elicit’s differentiation is “trust at scale”: a domain-specific language makes research workflows, coverage, and citations auditable. Formal work with seven of the top 20 life-sciences companies shows traction where evidence faces scientific, regulatory, or payer scrutiny. An inspectable external world model is next, but unstable probabilities and 80% automated-review accuracy keep evaluation and reliability unresolved.

View Dialogue Notes & Key Takeaways
  • Elicit’s differentiation is “trust at scale”: frontier intelligence orchestrated through workflows that execute exactly as specified. When Claude and ChatGPT were asked to analyze roughly 100 toxicology papers, they produced reports before admitting, “I did not analyze 100 papers”; Elicit’s domain-specific language instead guarantees that document No. 5 and No. 9,999 receive the same declared process. The bet is that high-stakes customers will pay for systematicity, citations, and auditability—not merely fluent answers.

  • Commercial traction is strongest where evidence must survive scientific, regulatory, or payer scrutiny. Elicit now works formally with seven of the top 20 life-sciences companies, spanning target and gene ranking, toxicology, experiment-related research, drug-launch strategy, and cost-effectiveness arguments after validated Phase 2 and Phase 3 trials. Its broader funnel runs from individual academics seeking technical, cited synthesis to teams screening thousands of papers and extracting data from figures.

  • The next product frontier is an external world model that turns sprawling evidence into an inspectable, continually updated decision system. After a cancer search still left Stuhlmüller with about 5,000 relevant papers, the problem was no longer retrieval but coherent reasoning across predictions, interventions, and counterfactuals. Elicit is exploring heterogeneous representations—causal graphs, spreadsheets, technology trees, SQL tables, and text—that preserve the flexibility of language while making continual learning “available to humans as a representation we can inspect and understand.”

  • Evaluation, not generation, is becoming the binding constraint because current models remain remarkably easy to push around. A model might forecast a 30% clinical-trial failure probability, then reverse itself when prompted with either a base rate or molecule-specific weakness; unlike an expert, it often lacks a stable underlying world model. Byun expects generation to become “more or less a solved problem,” shifting human work toward defining good performance, documenting failure modes, building verifiers, and deciding when an output is usable.

  • Hidden chain of thought does not eliminate process oversight: tool traces and “certificates of reasoning” can expose whether the necessary work occurred. A certificate might show sensitivity to changed inputs, what was examined, and which literature supports each claim; tool calls can reveal that a model never read the methodology section supporting its conclusion. The durable opportunity lies in independent consistency checks and checkable certificates that let people assess reasoning without replaying every step.

  • Elicit’s internal automation shows both the leverage and the reliability ceiling of agentic software. Its system, “The Line,” carries simple requests from a Slack reaction through specification, implementation, recorded testing, review, and deployment, merging roughly 30 to 50 issues per week fully automatically. Yet 80% accuracy in deciding what is safe for automated review is nowhere near sufficient when the other 20% might break production; Stuhlmüller suggested an “ultra-reliable mode,” not simply more inference.

  • Compute spending will rise selectively, with orchestration and model routing mattering more than defaulting to the largest model. Stuhlmüller personally spends about $2,000 per week on API tokens and might double or triple that, but avoids fast mode when marginal returns do not justify it; Elicit routes work among specialist models and uses cross-checking where it adds value. Even closely converged models retain “micro-jagged” differences: Claude Opus 4.5 beat Gemini 3 Pro on extraction accuracy, while Gemini reportedly led by at least 5 percentage points on direct evidentiary support.

  • A long-term design question is whether legible, discrete reasoning remains valuable as models integrate more modalities. Byun argued that discretization supplies “a little bit of error correction” at every step, which may keep tools, programs, and explicit representations relevant even if end-to-end multimodal models grow stronger. Stuhlmüller thinks AI could either worsen epistemics or enter a truth-seeking “basin of attraction”; the decisive variable is whether institutions explicitly optimize for truth before the largest decisions arrive.

  • 🔗 Original source & video: Radically Better Reasoning: Elicit’s Andreas Stuhlmüller & Jungwon Byun on World Models for Research

Listen to full conversation →