Pioneers Insight Method Research Author
Notion's Token Town: MCP vs CLIs and the Software Factory Future
Back to Episodes

Notion's Token Town: MCP vs CLIs and the Software Factory Future

Summary

  • Notion’s Custom Agents launch was its strongest yet for free trials and conversion, but it followed four or five rebuilds dating to late 2022. swyx notes that making it free for three months helped; Simon Last says earlier models lacked tool concepts, intelligence, and context length, leaving only “glimmers” of usefulness. The capability became viable around “Sonnet 3.6 or 3.7” early last year, while reliable background execution and enterprise permissions required additional product work.
  • Notion’s value is its role as an enterprise system of record and its collaboration expertise, not ownership of the underlying model or agent harness. Sara Ma compares its position to Datadog’s on top of AWS: foundational infrastructure is necessary, but understanding how customers collaborate is the value layer. Notion expects “a majority of our traffic” eventually to come from agents, giving its accumulated documents, meetings, tasks, and permissions increasing strategic importance.
  • The company has organized itself to rebuild continuously as model capabilities move, with product teams formed after shipping rather than before. Sara’s rule is to avoid “swimming upstream,” then determine where the river is flowing; Simon deliberately rethinks the stack roughly every six months. The “Simon vortex,” loose reporting boundaries, “demos over memos,” and a culture comfortable deleting its own code turn curiosity into an operating advantage.
  • Evals have become core infrastructure for product quality and model-provider feedback. Launch report cards target 80%-90% on defined journeys, while “Notion’s Last Exam” is intentionally held near a 30% pass rate to expose headroom. Notion sees quality differences between nominally identical models served through different vendors and has influenced prerelease snapshots using enterprise-work feedback.
  • Simon’s “coding agents are the kernel of AGI” thesis points toward a software factory in which humans supervise the outer system rather than type every line. The factory needs human-readable specifications, strong self-verification, and workflows that turn bugs into reviewed and merged fixes with minimal intervention. Sara describes the near-term human shift as an “identity crisis”: coding matters less than delegation and context switching, although Simon argues the resulting control plane remains deeply technical.
  • CLIs and MCP serve different agent architectures rather than representing a winner-take-all protocol contest. Boris Power describes CLIs as offering progressive disclosure and the bootstrapping power to debug or create their own tools; Simon Last calls MCP the “dumb simple thing that works” for narrow, lightweight, tightly permissioned agents. Notion will keep supporting MCP, but Sara argues that repeatedly spending language-model tokens on deterministic operations is wasteful when code can execute them once.
  • Usage credits let Notion meter models, GPU-served fine-tunes, web search, sandboxes, caching, and serving tiers without exposing every underlying cost. Charging by perceived business value proved too complicated, while agentic autofill—especially “Opus on every single database cell”—could cost billions of dollars. Auto is currently designed to select the right model and reduce user stress, not maximize margin; open-source options such as MiniMax help fill the missing middle of the intelligence-price-latency triangle.
  • Meeting Notes and composable Custom Agents create the clearest data flywheel: capture more work, make the system more useful, then automate the processes around it. One internal operator reduced more than 70 daily notifications from 30-plus agents to roughly five through a manager agent, while ordinary pages and databases provide memory and coordination. Sara’s strategic boundary is crisp: “Our job isn’t to build the best wearable to capture Meeting Notes. Our job is to build the best place where Meeting Notes live.”

Deep dive

1. Custom Agents needed several swings before the market caught up

  • Sara Ma called Custom Agents Notion’s most successful launch for free trials and conversion, while swyx supplied the useful caveat: “Making it free for three months helps.” Because teams were already two or three milestones ahead, launch day felt like delayed satisfaction rather than completion.

  • Simon said this was probably the fourth or fifth rebuild. The first effort began after access to GPT-4 in late 2022, when the team called the concept an “assistant”: give it every Notion capability, let it run in the background, and have it perform work autonomously.

  • Before native function calling, Notion worked with Anthropic, OpenAI, and Fireworks on its own multi-turn tool framework and fine-tuning. The models were “just too dumb,” context windows were too short, and promising demonstrations never became reliably delightful.

  • Simon places the model-level unlock around “Sonnet 3.6 or 3.7” early last year. Custom Agents then took longer than the earlier agent because unattended execution demanded greater reliability and a comprehensible permissions interface across partially overlapping Slack groups and document audiences.

2. Notion runs an AGI portfolio without swimming upstream

  • Simon described a portfolio approach balancing maintenance, capabilities that work now, and “a few projects that are a little bit crazy.” The company wants to be “AGI-pilled” without sacrificing useful shipments while it builds toward where models are going.

  • Sara’s two-part discipline is first recognizing when a team is “swimming upstream” against model limits rather than suffering from bad context or infrastructure, then asking which direction the river is flowing. The trick is to begin building for that direction without persisting too long at an impossible implementation.

  • Asked what might look obvious in 18 months, Simon answered that “coding agents are the kernel of AGI. Everything is a coding agent.” Because an agent can bootstrap, debug, and maintain its own capabilities, Notion is exploring a “software factory” where multiple agents develop, review, merge, and operate a service together.

3. The value is collaboration expertise, not model ownership

  • Sara’s analogy is Datadog and AWS: Datadog requires cloud infrastructure even though AWS offers CloudWatch, but its expertise lies in how customers want observability. Likewise, Notion’s expertise is “understanding how people wanna collaborate,” regardless of which models provide the underlying capability.

  • Simon distinguished Notion from narrow vertical SaaS. Its job is to listen across a broad customer base, decompose varied requests into reusable primitives, and preserve a system that remains coherent and pleasant rather than accumulating disconnected vertical features.

  • Sara warned that focusing on “cool tools” produces the team’s lowest velocity. Each Friday, the group examines the P99 most token-intensive Custom Agent transcript and cuts failed tasks against concrete journeys such as email triage; a sandbox or computer tool earns priority when it solves PDF export, not because the tool itself sounds exciting.

4. The “Simon vortex” institutionalizes rebuilding

  • Sara does not see her job as supplying the ideas or being the deepest technical expert. Leadership establishes the objective and prioritization mechanism, then lets prototypes from people close to user problems redirect the roadmap: “Proof is in the pudding.”

  • Rebuilding the harness three or four times required people comfortable deleting their own work, without treating design documents as promotion artifacts. Sara credits Simon Last and Notion co-founder Ivan for a low-ego culture where saying “I wrote that code” does not become an organizational veto.

  • The “Simon vortex” is a skunkworks-like rotation of trusted senior engineers around rapidly changing prototypes. Reporting and working relationships remain loose, and Notion historically forms organizational structures “after we ship things, not before.”

  • Company hackathons teach the broader workforce—one recent exercise asked everyone to build an agentic tool loop—but Simon’s warning is categorical: if hackathons are the only route to invention, “you’re toast.” Image generation shipped because Jimmy, an engineer outside the AI team, pursued it with Gemini access, token tracking, and eval support until it became a full project.

5. Demos replace mocks, while the platform absorbs their blast radius

  • Sara’s core AI capabilities and infrastructure organization has about 50 people, with another 30-40 packaging the technology into chat, Custom Agents, and Meeting Notes. Every product team also owns the agent-facing version of its service, from competing CRDT edits to SQL queries.

  • This ownership reflects the forecast that most product traffic will eventually come from agents rather than humans. The editor, database, and other product teams therefore build simultaneously for both constituencies instead of routing every agent feature through one central AI group.

  • Notion’s Design Playground gives designers reusable components and a working agent, so they deliver URLs rather than mocks. For engineers, Simon says the prototype bar is essentially “a feature flag that actually works,” amplified by company-wide dogfooding on a development instance full of experimental flags.

  • “Demos over memos” forces stronger product conviction because almost anything can now be demonstrated. Sara’s test is whether the work builds “one tower” rather than “a really flat hill”; behind it, the agent-platform-velocity organization supplies eval tooling, compliance, vendor work, and operational hardening so prototype owners can continue maintaining what they ship.

6. Evals are a product system, not a single quality score

  • Sara rejects “evals” as a synonym for one quality number. CI contains unit-like regression tests with stochastic tolerances; product report cards demand roughly 80%-90% across launch-critical journeys; frontier or headroom evals are deliberately constructed to pass only about 30% of the time.

  • The 30% suite is “Notion’s Last Exam,” created after older evals saturated and could say little beyond “it wasn’t worse.” Notion staffs it with a data scientist, a model behavior engineer, and a full-time eval engineer, both to anticipate the river’s direction and to give Anthropic and OpenAI useful frontier feedback.

  • Notion observes different quality from ostensibly identical models served first-party or through Bedrock, Azure, and other vendors, plus slower service during working hours. Sara says labs have also sent multiple prerelease snapshots and, in some cases, changed the version ultimately shipped after Notion identified enterprise-work regressions overlooked by coding-heavy benchmarks.

  • Model behavior engineers evolved from “data specialists” manually judging Google Sheets into a distinct path combining data science, test PM, prompting, linguistics, and taste. Coding agents now help them download datasets, run evals, diagnose failures, and implement fixes, but Sara insists supervision need not come from software engineers.

7. Software engineers move upward into a technical control plane

  • Sara says every Notion engineer experienced an identity crisis similar to a new manager’s: “their ability to write code is less important than their ability to delegate and context switch.” Simon frames the same shift as a continuum from manually typed code, through autocomplete, to agents that debug, verify, merge, and deploy longer tasks.

  • Simon rejects the idea that this merely turns engineers into people managers. Humans are fuzzy; agents can be modeled as a rigorous system of PRs, blocked states, approvals, memory, and recovery. Designing that outer system remains “a hard engineering problem” and a deeply technical one.

  • His software-factory requirements begin with a human-readable specification layer—Markdown files or a database of Notion pages—followed by strong self-verification and testing. The process layer must define how a reported bug reaches a sub-agent, becomes a PR, gets reviewed, and merges while preserving required invariants with minimal human intervention.

8. Custom Agents compose through ordinary records, not exotic orchestration

  • Alessio’s Kernel Labs demo turned incoming coworking applications into an enriched Notion database: the agent checked email, added rows, searched the web, and extracted move-in timing. Setup took roughly 15 minutes, and the information remained where he would already have kept it.

  • Sara’s strongest internal specimen is bug triage: a Slack-resident agent uses a routing constitution, creates an item in the appropriate task database, and posts back to the channel. Her formulation is precise: “It’s not replacing people, it’s replacing processes.”

  • Simon described two composition modes. Loosely coupled agents can coordinate by watching and writing databases; a forthcoming setting lets one agent invoke another directly. Alessio immediately raised recursion and infinite-loop risk—“Everything’s gonna be paperclips”—and Simon and Sara acknowledged that some limit exists without supplying the number.

  • One go-to-market operator had more than 30 agents producing over 70 blocked-task notifications daily. A manager agent reading their issue database cut the human-facing load to about five and could help diagnose failures. Notion similarly avoids a dedicated memory primitive: memory is simply a page or database that both humans and agents can edit.

9. MCP and CLIs win at different layers of the stack

  • Boris Power’s CLI case starts with terminal-native leverage: pagination, files, help commands, and progressive disclosure hide irrelevant capabilities until needed. More importantly, the environment is bootstrappable—a browserless agent reportedly wrote a roughly 100-line Chromium wrapper for itself and could repair the tool if it failed.

  • Boris’s Chrome DevTools MCP counterexample exposes the trade-off: when its transport breaks, the agent loses the browser and cannot repair the external server. Yet Simon still calls MCP the “dumb simple thing that works” for narrow, lightweight agents whose permissions should stop at explicitly exposed tool calls.

  • CLIs create harder token and credential questions because an agent with runtime access might reach or exfiltrate an API token. Sara adds an economic argument: using language to repeatedly execute deterministic third-party actions wastes tokens, especially outside cache windows, whereas generated code calling a CLI can impose a one-time reasoning cost.

  • Notion therefore mixes layers. Linear and GitHub integrations may use MCP, while Slack, mail, calendar, and search receive higher-touch native tooling and triggers; MCP has no trigger protocol. Internal abstractions normalize tools, agents, completions, tasks, and chat archetypes, with MCP treated as one integration type rather than the entire architecture.

10. Many rebuilds converged on model-native abstractions and 100-plus tools

  • The first late-2022 architecture was itself a coding agent: represent every action as JavaScript and expose JavaScript APIs. Models then were not good enough at code, so Notion moved toward tool calling before standard tool calling existed.

  • The replacement used an XML representation designed to map losslessly onto Notion blocks. It fit Notion’s internals but not the model’s learned environment, prompting the larger lesson: “Give the models what they want.” Notion-flavored Markdown kept plain Markdown at the core and accepted that conversion need not be lossless.

  • Database access followed the same path. A complex JSON query format mapped neatly to internal structures but burdened the model, so Notion exposed SQLite-style queries instead. That choice benefited from an existing system in which Notion databases were already queried across clusters of SQLite databases.

  • The broader arc stripped away one-shot prompts and few-shot examples in favor of goal-driven tool definitions and feedback loops. Ownership could then move from five or six prompt gatekeepers to individual product teams; progressive disclosure now protects quality and token usage as the latest agent exceeds 100 tools and even saying hello would otherwise consume thousands of tokens.

11. “Teach to the top of the class” shaped the agent UX

  • Notion does not treat its system prompt or tool list as secret sauce; operators can ask the agent what tools exist. Simon’s principle is to “teach to the top of the class,” preserving enough depth and interpretability for power users to understand how the agent works and prompt it precisely.

  • The team sharpened the trade-off: making setup maximally easy can abstract away interpretability and “nerf” the agent. A decisive product moment came when the team agreed Custom Agents were not for everyone, which clarified the intended operator and accelerated work.

  • The agent can configure itself because it receives setup and debugging tools plus a development guide explaining good instructions and end-to-end testing. When it fails, the user can ask why and request an instruction update; fully automatic self-healing remains roadmap work.

  • Permissions constrain that bootstrapping. Background agents begin with no access and cannot silently edit their own permissions; pressing Fix enters a synchronous admin mode where proposed changes are visible and confirmed. The chat-first “Flippy” redesign made setup and use the same conversation, delaying launch about a month but replacing settings as the primary experience.

12. Credits price compute, while Auto selects model choice

  • Credits sit above raw tokens because Notion’s costs also include GPU-served fine-tunes, differently priced web search, potential sandboxes, serving tiers, cache rates, and asynchronous processing. Credit packs also fit enterprise procurement and volume discounts better than exposing each infrastructure unit separately.

  • Notion initially considered charging per agent run or task value, but complexity repeatedly mapped back to token throughput. Usage pricing also prevents catastrophic subsidy: Madhu Muthukumar says running Opus agentically across every database-autofill cell could cost “billions of dollars.”

  • Auto is intended to choose the best model for the task, not the cheapest model for Notion; Madhu says it is not currently used as a margin maker. Because asynchronous users care less about speed, Notion adds cost cues and may nudge someone away from using Opus to triage every email.

  • Madhu sees an unfilled middle in the intelligence-price-latency triangle: models cluster at a few capability, speed, and cost points, while smaller options have not always become proportionally cheaper. Notion offers MiniMax and collaborates with open-source labs to expand choice; Simon adds that the ideal agent may “automate itself out of a job” by replacing repeated reasoning with code.

13. Training loses to the outer loop—except where retrieval truly changes

  • Madhu rejected training a Notion foundation model as a necessary core competency. Simon is more interested in enterprise-specific fine-tuning that knows a company’s context and people, while large customers also ask about bring-your-own-model arrangements; public prompts and tool definitions make those models easier to connect.

  • Simon admits he “burned a lot of time trying to train models.” Notion changes tools daily, making a tool-specialized model stale before the investment pays back; his current diagnosis is that “99% of the time it’s a bug in one of the tools,” so velocity in harnesses, tools, verification, and debugging beats reflexive retraining.

  • His work pattern nevertheless came full circle: training once required starting overnight experiments, and now he starts coding agents before bed, aiming for jobs that will still be running in the morning. One thread ran almost continuously for 17 days and compacted about 100 times because of a harness bug.

  • Retrieval is the major exception because most search traffic on AI-enabled plans now comes from agents. Their queries favor top-K coverage over human click-through position, different snippets, and parallel exhaustive search; swyx described an eight-query fan-out intended to maximize query diversity. The “agentic find” team now treats ranking, query generation, indexing, and retrieval as one journey, with less emphasis on choosing vector embeddings.

14. Meeting Notes turns conversation into compounding enterprise context

  • Sara calls Meeting Notes one of Notion’s strongest growth levers for adoption, virality, and retention. Her personal example captures its system-of-record value: for a self-review, she bases the review on conversations with her manager because work never mentioned in those one-on-ones was probably not material to the review.

  • Internally, a Custom Agent assembles a standup pre-read from Slack and GitHub, creates the meeting note, and asks attendees to read it. After a hands-off-keyboard discussion, another calendar-triggered agent files tasks and sends the follow-up Slack messages decided in the meeting.

  • Transcripts created an explosion of long-form content that forced improvements in search, context management, and compaction. Agentic summaries now attempt to resolve and @mention the correct person—for example, the most probable Simon—using attendance data, generated profiles, and people-similarity machinery, though Sara and Simon acknowledge it can still be wrong.

  • Sara reframes the product as data capture: transcription is the primitive, Meeting Notes packages an agent on top, and future agents might update the relevant task database during the conversation. Wearable partnerships could feed more context into Notion, but the boundary remains collaboration: build “the best place where Meeting Notes live,” not necessarily the best capture hardware.