Calm AI for Crazy Days: Inside Granola's Design Philosophy, with co-founder Sam Stephenson
Summary
Granola’s breakout is powered by a simple sharing loop: users share polished notes, recipients ask how they appeared so quickly, and some become users themselves. Nathan recalls a Ramp report saying Granola ranked second only to Anthropic for January customer additions, while Sam Stephenson says growth is “almost entirely through word of mouth.” Recipes generate social attention, but shared notes—sometimes appearing in Slack 30 seconds after a meeting—are a recurring discovery mechanism.
Granola’s broader appeal comes from designing for the most frazzled user, not the most technically capable one. Inspired by a kitchen-utensil company Stephenson believed was OXO—while acknowledging he might be wrong—Granola targets people with calendars that are “a constant block of stuff,” barely able to use the bathroom between meetings. The intended result is a calm, easy-to-use product for everyone.
Stephenson’s deliberately “surprisingly unambitious” roadmap is a strategic response to AI’s brittleness in nuanced knowledge work. Agents can produce plausible screenshots yet fail when they misread relationships, priorities, or tone even slightly; Granola instead tackles bounded jobs it can perform reliably, beginning with notes and postponed menial work. Its design budgets roughly 2% of a user’s attention during meetings but perhaps 80% in deliberate, full-screen multi-meeting chat.
Enterprise context is a major opportunity and a difficult governance problem. Company information remains fragmented behind admins, security reviews, and permissions, while a single personal remark can “pollute a transcript” and make broad sharing unacceptable. Granola defaults every note to private and uses an LLM to suggest a folder; Stephenson says even a hypothetical 99.98% routing accuracy might not be sufficient when one sensitive disclosure can make sharing unacceptable.
The current seat-based product hides inference complexity, but heavier agentic work may force usage-based pricing. Granola initially declared that “there is no budget”; at one point transcription consumed roughly half its burn, and extrapolated growth looked “terrifyingly expensive.” Those costs fell with scale and commoditization, but sophisticated multi-meeting users still generate bills Granola swallows, making a Cursor-like usage model plausible once the product directly performs more work.
AI coding has sharply shortened Granola’s idea-to-evaluation cycle without eliminating design judgment or Figma. The roughly 60-person company includes 25–30 engineers, three product people, three designers plus Stephenson, and a couple of design engineers; designers now prototype inside the live app because actual meetings expose qualities mocks cannot. Figma has shifted from “the start, middle and end” to a specialist ideation tool, while Friday demos, intensive dogfooding, and close user observation outweigh sheer feedback volume despite the roughly 10,000-person beta program.
The existential bear case is that powerful model providers absorb specialized applications, while the bull case is that focused products remain a little better at painful workflows. Nathan connects this to Andrew Critch’s “big tech singularity”; Stephenson calls that the main risk but argues that specialized tools may retain merit. Granola must stand on model builders’ shoulders while preserving a superior meeting experience. His end state is less screen servitude and more human presence—people “firing on all cylinders” because the computer handles capture and clerical movement of information.
Deep dive
1. One sharing loop explains Granola’s breakout
Nathan Labenz opens with a strong commercial marker: he recalls a Ramp report saying Granola ranked second only to Anthropic in new customers added during January. Stephenson was “taken aback” to see the company in that peer group, even though its internal growth trajectory was already clear.
Stephenson’s attribution is emphatic: growth comes “almost entirely through word of mouth,” either from explicit recommendations or shared notes. Granola assumed that a sufficiently good product with embedded virality would eventually compound; unlocking those loops took time, but now “the thing’s really starting to snowball” with strong month-over-month growth.
Labenz’s growth lesson—worth keeping—is that products usually need one mechanism that truly works, not a collection of clever hooks. Granola’s clearest example is concrete output: a colleague sees polished notes appear almost immediately and cannot reconcile their quality with the time elapsed.
2. Granola designs for the frazzled edge case
Granola began with a broader ambition than meeting notes: invent interfaces that let ordinary workers access AI without making computers an object of study. Notes were the “foot in the door”—a habitual, bounded workflow from which the product could earn permission to help with more over time.
The target archetype emerged from open-ended research: someone in back-to-back meetings, repeatedly context-switching and “barely” finding time to use the bathroom. That includes salespeople, account managers, recruiters, investors, founders, and client-service workers, but Stephenson treats the archetype as an extreme version of conditions nearly every knowledge worker encounters.
Stephenson credits an analogy from co-founder Chris: a kitchen-utensil brand—he thought OXO, while explicitly allowing he might be wrong—designed for people with impaired grip, one arm, or another disability. Solving for the extreme produced friendlier tools for everyone; Granola similarly seeks an easy, calm experience by serving the most chaotic calendar first.
3. Software should assume users are reactive, not calmly rational
Stephenson argues that builders routinely imagine users arriving calm, attentive, and ready to master complicated sequences. Actual work is more reactive: somebody opens an inbox, finds “three things” burning, sees another meeting starting in 20 minutes, and remains behind and overwhelmed for the rest of the day.
His design implication is blunt: much workplace behavior occurs in “System 1,” not the rational, methodical mode software assumes. Sophistication that looks defensible in a focused usability session can collapse when the product receives only a fragment of attention amid real work.
Labenz spots the testing problem: asking somebody to demonstrate software creates artificial focus, partly because the participant does not want to look foolish. He jokes about staging a fire alarm and spilling coffee to reproduce reality, underscoring how easily conventional research validates a user state that rarely exists.
4. Concrete artifacts beat users’ imagined accounts of themselves
Stephenson grounds research in evidence users cannot abstract away. While exploring folders, Granola asked people to share their actual home screen, then opened meetings one by one: what happened, who attended, who should see the notes, and how would the user mentally organize them?
Calendars work as the same forcing function. Once an interview moves into what somebody “theoretically or generally” does, Stephenson says, “you can’t trust any of that”; the researcher is hearing the person’s imagined identity rather than behavior demonstrated by meetings, notes, and schedules.
For evaluating built features, Granola leans even harder on dogfooding. Stephenson can quietly watch a teammate use Granola during a real sales call, obtaining a less filtered signal than a retrospective description from somebody giving the product full attention.
5. “Surprisingly unambitious” automation is the credible near-term wedge
Labenz presses on an apparent tension: should AI merely accommodate overloaded, System 1 work, or free people into strategic, System 2 thought? Stephenson’s answer is “both,” but at different speeds and on different surfaces.
General work assistants still miss social nuance: relationship history, comparative priority, appropriate tone, and how one request ranks against everything else. A generated email or to-do list can look impressive in a screenshot, yet if it is “almost there” rather than right, Stephenson says it becomes unusable; better models will help, but the transition “might take a surprisingly long time.”
Granola’s response is to be “surprisingly unambitious” about promised tasks. It may not resolve a user’s thorniest burning issue soon, but it can clear boring, low-priority work the user keeps postponing; reliability on that bounded layer creates real assistance before a complete executive agent is possible.
The System 2 surface already looks different: full-screen chat across a user’s or company’s meetings supports long-form work such as writing posts, drafting job descriptions, or analyzing a business area. That mode can demand sustained thought instead of disappearing into the background of a call.
6. Deep context is constrained by institutions as much as models
Labenz describes exporting five years of email, Slack, direct messages, transcribed calls, and diarized podcasts into one local database. With that context, an agent can inspect relationship history and drafting patterns well enough to produce an introduction requiring only a small edit—evidence that the “promised land” is visible, though not yet reached.
Stephenson agrees that a machine is only as good as its context, but separates personal experiments from company deployment. Organizational knowledge is fragmented behind permissions, security approval, and tool boundaries; Labenz could unify his own data largely because he administered every system, an option ordinary employees usually lack.
Personalization is also deeper than many products assume. A useful agent needs knowledge of the user, colleagues, projects, priorities, and changing circumstances; highly proactive early adopters can teach general-purpose agents today, but Granola’s job is to make that learning “happen in the background” for average users.
7. Shared organizational memory creates a permissions paradox
Granola defaults every note to private until the user deliberately moves it into a shared space. Stephenson’s reason is conversational contamination: “You only have to say one personal or dodgy thing” to pollute an otherwise useful transcript and make organization-wide sharing inappropriate.
The upside of collective context is nevertheless immense. Teams can access and combine information across meetings, but the product must place each transcript where it belongs without leaking sensitive material; Stephenson calls this a design problem Granola “hasn’t nailed” despite ongoing improvement.
Granola currently uses an LLM to inspect a call, available folders, and the characteristics of notes already stored in them, then recommend a destination. It does not file automatically because accuracy is not high enough, and Stephenson questions whether even a hypothetical 99.98% success rate would satisfy users when a single failure could be unacceptable.
Labenz’s pushback is that models already identify potentially regrettable podcast remarks reasonably well. Stephenson accepts the direction and says Granola is tinkering, but internal conversations require background knowledge the model may lack; the remaining issue is less average accuracy than the asymmetric cost of one bad share.
8. Fixed-seat simplicity currently outranks perfect inference margins
Granola’s early instruction to itself was that “there is no budget”: with few users, optimizing cost would distract from the harder task of making something people loved. The product was expensive at launch, chiefly because real-time transcription APIs carried substantial cost.
Today, core note-taking has predictable economics because human meeting time has physical limits and usage averages across a month. Sophisticated users who chat across many transcripts can still create large bills that Granola “swallow[s] at the moment,” preserving the simple experience of paying one seat price and receiving one comprehensible product.
The frightening phase was about a year before the conversation, or roughly six months after launch, when extrapolating user growth and then-current costs produced a “terrifyingly expensive” curve; at one point, about half the company’s burn went to transcription. Scale and commoditization brought that under control.
Stephenson expects the pendulum might swing back as Granola performs more LLM-driven work rather than merely assisting. A Cursor-like usage model could eventually make sense, but only when direct work changes the product’s economic shape; for a meeting-notes app, transparent seat pricing still matches customer expectations.
9. Real-time transcription trades accuracy for trust and immediacy
Granola captures microphone and system audio, then streams both to cloud transcription APIs. Real-time text reassures users that recording is working, allows notes to arrive immediately after a meeting, and lets Granola discard audio rather than retain a more sensitive artifact.
Stephenson openly accepts the compromise: real-time transcription is worse than uploading a complete recording and waiting, particularly for quality and speaker separation. Yet the whole product benefits from incremental gains; improving transcription by 10% creates multiple downstream effects, while better speaker separation could make everything else “10 times better.”
Deepgram and AssemblyAI are Granola’s two principal providers, with desktop traffic split because their quality is comparable and redundancy matters for core infrastructure. The architecture permits rapid model swaps, leaving Granola free to chase better speed, quality, or speaker separation and potentially move on-device at some point.
Labenz notes that API reliability alone justifies a fallback. Stephenson agrees that users notice outages immediately and painfully; transcription is too foundational to let one provider’s downtime take the entire product with it.
10. OS-level capture maximizes coverage while shifting consent outward
Granola chose a desktop architecture so it would work wherever users were having conversations, without their having to ask whether Granola was suitable for a particular meeting. Unlike a bot participant, it listens at the operating-system audio layer—and sacrifices the discovery benefit of being a “huge glowing orb with a logo” visible to everyone.
The earliest version briefly stored audio in an S3 bucket in case it later became valuable for training. After about a week, the team stopped: retaining voices felt “too creepy,” made the product heavier and more serious, and suggested a tool being used against people rather than for them.
Granola retains transcripts because notes serve humans while transcripts give LLMs the detail needed for subsequent work. Stephenson does not seek a word-for-word record that would establish in court exactly who said what; raw audio is not necessary for that narrower product purpose.
The company treats Granola like voice memos or a notepad: users must follow applicable laws and disclose recording appropriately. Granola offers calendar notices and is developing more transparency because open use benefits consent and growth, but quiet capture remains technically possible.
11. Work meetings offer a social contract that ambient wearables lack
Stephenson says privacy and consent questions have declined since Granola began two years earlier. He believes workplace transcription will become increasingly normal because its benefits are substantial and organizations will develop ways to mitigate the downsides.
His boundary is situational, not absolute. Meetings are set pieces where participants agree to work together, creating a moment for an explicit recording contract; “always-on wearable thingamajigs” face a much harder problem because personal life lacks that contained social negotiation.
Labenz remains less certain that transcripts are dramatically safer than audio, especially when answers can be grounded to exact conversational moments. Stephenson’s defense rests less on perfect anonymity than minimization: Granola stores only what its notes and LLM workflows require, avoiding the richer and more emotionally charged original recording.
12. Forgetfulness may be a safety feature, not a defect
Asked how autonomous agents should access a person’s history, Stephenson gives an unusually honest answer: “No, I don’t know.” He reaches instead for a human analogy—an executive assistant needs sufficient, current context, but should not attend every private conversation between founders; partial access has costs, yet protected spaces preserve candid relationships.
Stephenson argues that AI design does not think enough about forgetfulness. Human memories lose fidelity quickly, and an agent whose recollections become fuzzy could diffuse dangerous details from transcripts or email: in many cases, forgetting is “a feature not a bug.”
Labenz extends the idea to agents that literally shrink as model weights are pruned, becoming efficient and increasingly incapable outside a defined niche. His broader critique is that the industry is conducting a depth-first search on one “everything AI” architecture instead of exploring diverse trade-offs before driving the first successful approach toward AGI.
13. Granola allocates interface complexity according to attention
Stephenson is “almost paranoid” about adding anything to the meeting notepad. The core sequence—popup, start, write alongside the conversation, receive generated notes—must clear a very high feature bar because every added control makes the workspace less calm.
That restraint creates deliberate discoverability costs. Templates can structure notes almost arbitrarily, yet many users never find them; Granola keeps the feature tucked away because most people should not need to think about it, while specialists can optimize once motivated.
The governing ratio is striking: during a meeting, Granola expects roughly 2% of attention while the other person receives 98%. In multi-meeting chat, it may receive 80%, permitting denser controls and faster experimentation; Stephenson concedes that the company’s calmness bias sometimes delays genuinely useful ideas.
14. Recipes convert conversational residue into repeatable work
Recipes emerged because power users were discovering valuable prompts across many meetings but lacked an easy way to repeat them. A recipe packages that prompt behind one click, turning accumulated conversation into a reusable workflow and showing mainstream users what deeper context can do.
Granola’s customer-experience team holds an unstructured discussion of bugs and recurring confusion, then runs a recipe that inspects the transcript alongside existing documentation and proposes targeted edits. Hiring works similarly: teammates talk through a role, supply previous job descriptions, and receive a strong first draft in language grounded in the company.
The most emotionally powerful category is personal coaching. In research, users ran a “Coach Me” recipe written by famous CEO coach Matt Mochary; nobody quite cried, Stephenson recalls, but reactions were profound because Granola identified personal patterns across conversations and suggested how the user could improve.
Other examples reinforce reflection rather than clerical speed: Labenz’s “Blind Spot Finder” surveys his AI coverage, while Dan Shipper’s implicit-company-culture recipe contrasts a company’s behavior with its stated values. Granola is 100% in person and uses its mobile app to record discussions on a phone, then uses AI to amalgamate debates into shared product principles.
15. Live prototypes accelerate judgment, but specialized products still face existential risk
Granola is roughly 60 people: 25–30 engineers, three product people, three designers plus Stephenson, a couple of design engineers, and the remaining business functions. Designers are expected at least to be curious about tools such as Claude Code; they prototype inside the real app because authentic content and peripheral attention determine whether the meeting interface works.
Moving chat from a sidebar to a floating, globally accessible control illustrates the gain. Mockups had prolonged debate, but once Stephenson built the live version, one day of internal use made its superiority obvious; AI coding compresses idea-to-evaluation even when production edge cases still require conventional design work.
Figma remains useful for comparing ten variants side by side, mapping copy-heavy signup flows, and seeing states together, but it is no longer “the start, middle and end.” Full-app schematics are gone; Figma is an ad hoc specialist tool, and Stephenson says he would be nervous in the company’s position as coding tools approach ideation from the other direction.
Granola’s culture assumes “nine times out of the 10” an idea will fail with users. Engineers talk directly to customers, Friday demos welcome roadmap work and left-field experiments, and even founders publicly retire bad concepts; despite roughly 10,000 beta users, core decisions remain qualitative, based on dogfooding, close observation, and “gut feel.”
Nathan raises Andrew Critch’s “big tech singularity”: powerful AI companies could turn their systems toward any niche, copy specialized products, and potentially subsidize them. Stephenson calls that the main risk, but says technology companies repeatedly overestimate one interface swallowing everything; specialized tools may retain merit. His test is that if each new Claude or GPT does not improve Granola, “over time we’re cooked”; survival requires riding the rising tide while staying better at painful meetings.
His positive endpoint is not more software consumption. Knowledge workers are currently “a slave to your computer,” moving a mouse and shuffling information; Granola’s most valued outcome is that users feel present and can have their brains “firing on all cylinders.” Labenz supplies the practical metric: if AI does not help him get outside more, is it really serving him?