Pioneers Insight Method Research Author
How Decagon Runs 90% of Its Agents on Open-Source Models
Back to Episodes

How Decagon Runs 90% of Its Agents on Open-Source Models

Summary

  • Decagon now runs 90% of its workflow on open-source models because production voice agents reward task-specific accuracy and low latency, not general intelligence. Fine-tuned smaller models can outperform frontier systems on narrow jobs such as topic identification or bad-actor detection while also being cheaper and faster: “We end up getting all three things.” The remaining 10% supports new, experimental, or unusually open-ended work.

  • The durable model split is frontier for discovery, open source for scaled production. Frontier APIs remain the easiest way to launch an uncertain use case, while a stable workflow creates strong incentives to fine-tune and control smaller models; Decagon Autopilot still uses frontier intelligence to review as many as a million conversations, detect trends, generate variants, and test improvements. Enterprise migration will be slow because proprietary data, custom evals, security, and model-risk governance matter more than model availability.

  • Decagon treats its research organization as a continuously operating “model factory,” not a one-time infrastructure project. New capabilities create new tasks to automate, while stronger base models make older fine-tunes obsolete; the company therefore trains and retires models continually. It builds tightly coupled system-level evaluation internally, buys commodity labeling and dataset-diversity tools, and optimizes for the customer’s unit of value—a conversation—even as model calls and tokens per conversation rise.

  • The application-layer thesis is that models do not encode changing business processes, integrations, controls, or systems of record. Decagon fine-tunes for reusable customer-service behavior, but keeps each enterprise’s procedures in context so they can change without retraining; the surrounding product must handle testing, QA, compliance, collaboration, and cases such as rebooking three travelers after a canceled flight. Even with AGI, the founders argue, agents will still need software “to store work and pull information from and reason about things.”

  • Forward deployment is valuable only when customer pain compounds into reusable product. Early AI workflows require people to “lay out the track as they see which way the train is going,” but known workflows should be productized for the next 10 customers; otherwise, the company becomes “a glorified consulting shop.” Ashwin Sreenivas preserves Palantir’s sharper formulation: forward-deployed engineers “eat pain and excrete product.”

  • Duo demonstrates Decagon’s compounding product loop: an expensive human deployment task becomes a feature, then Duo Autopilot automates that feature’s ongoing improvement. The larger, slower agent can turn transcripts and documentation into agent operating procedures, integrations, tests, simulations, conversation monitoring, and drafted fixes—work the core conversational agent cannot do. Decagon’s near-term moat, Jesse Zhang argues, is the enterprise infrastructure around that intelligence, though he concedes that if agents eventually generate all such infrastructure on demand, “I don’t know, and we’ll figure out in three years.”

  • Decagon’s commercial wedge is a “glass box” customers can operate themselves, widening from support into an AI front door for every customer interaction. One customer reportedly left Sierra after producing three journeys in roughly a year, then built seven on Decagon within a month; founder involvement remains heavy, with Jesse Zhang estimating that sales consumes about 80% of his time. Lower service costs can also unlock latent demand rather than translate mechanically into layoffs: Ashwin Sreenivas argues that “AI will kill jobs but not careers.”

Deep dive

1. Production scale pushed Decagon from frontier APIs to 90% open source

  • Jesse Zhang’s starting point was pragmatic: when Decagon only needed to prove that an agent could deliver value, OpenAI and Anthropic were the obvious choices. The calculus changed with larger enterprises, millions of end customers, and voice, where an otherwise good answer arriving too slowly is still a poor product.

  • A conversational agent actually performs many bounded jobs: identify the customer’s topic, detect a possible bad actor, or perform other narrow checks. None requires a model that is simultaneously excellent at mathematics, coding, and every other general-purpose task; each requires exceptional performance on one defined behavior.

  • Jesse Zhang rejected the standard “smart and expensive versus dumb and cheap” framing. A smaller model is less general, not necessarily worse: after task-specific fine-tuning, Decagon sees it outperform “the large, smart, state-of-the-art models” while lowering both latency and cost. “We end up getting all three things.”

  • The resulting production mix is 90% open source and 10% closed-source frontier models. Zhang framed model selection across cost, intelligence, and latency: Decagon initially accepted less general intelligence to improve voice latency, then found that specialization could raise accuracy on the actual task as well.

2. Frontier models remain the discovery engine for open-ended work

  • Decagon still reaches for frontier intelligence when the job is broad and exploratory. Decagon Autopilot, for example, may review a million completed conversations, discover patterns, generate variants of the core agent, and determine which variants perform better—far less bounded work than executing an established rebooking or healthcare procedure.

  • The founders’ production rule is lifecycle-based: use convenient frontier APIs while a use case is new, then consider open source once its shape is stable and running at scale. At that point, Zhang argued, open source can become “strictly better” because the company gains control, speed, and cost advantages without sacrificing task performance.

  • That explains the apparent paradox that open-source inference share might decline while enthusiasm rises: enterprises are launching many new experiments, and frontier APIs are the fastest starting point. Successful experiments will migrate later, but proprietary data, custom benchmarks, security reviews, model-risk governance, and limited organizational capacity make that transition slower than the discourse suggests.

3. Decagon Labs operates as a continuously rebuilding model factory

  • Decagon does not fine-tune one permanent model stack. As base capabilities advance, the team finds newly repeatable tasks worth specializing, trains net-new models, and deprecates fine-tunes whose purpose has been absorbed by stronger open-source bases. Jesse Zhang called Decagon Labs “a model factory of sorts.”

  • Evaluation is inseparable from the application. Public benchmarks cannot show whether one specialized model, working alongside several others, delivered the final customer outcome; Decagon therefore measures the system end to end rather than celebrating a favorable loss curve or an isolated task score.

  • That coupling determines the build-versus-buy line. The company builds training and evaluation infrastructure when it embodies Decagon’s architecture and customer outcomes, but buys broadly reusable capabilities such as data labeling and measurement of dataset diversity. The governing question is simply how quickly the best model can reach production.

  • Token cost is secondary at Decagon’s current growth stage. Customers buy a successful conversation, not Decagon’s internal token count, and tokens per conversation have actually increased as the company adds checks, parallel work, and model calls to improve quality. Zhang’s priority is growth; deeper cost optimization becomes central “if we’ve won the market.”

4. Application companies own the business logic that base models do not

  • Zhang distinguished two forms of customization. Fine-tuning teaches reusable customer-service behaviors across Decagon’s customers; customer-specific procedures stay in context because retraining and reversing a model every time an enterprise changes policy would be impractical.

  • The application layer captures what happens when, for example, a canceled flight forces the rebooking of three people together. It must connect systems, express permissions and procedures, run experiments, inspect conversations, support QA, and give compliance teams visibility. Those requirements are business software, not latent knowledge inside a foundation model.

  • The labs and applications are nevertheless converging. Frontier labs add general applications to help enterprises realize value, while Decagon builds specialized models to improve performance, latency, and cost. Sreenivas suggested that application companies may ultimately become “labs for specific verticals,” with their models serving as the primary product.

  • His rebuttal to the “last startups” thesis was deliberately simple: humans are already broadly intelligent, yet they still need databases and CRMs. AGI agents will likewise need places to store work and retrieve facts; some software designed only around human interaction may face pressure, but “I don’t think software as a whole in any meaningful way is going away.”

5. Forward deployment must discover a product, not subsidize consulting

  • Sreenivas sees forward-deployed engineers as newly necessary because AI workflows have not yet been mapped. They embed while both vendor and customer learn what the product should do, “laying out the track as they see which way the train is going.” Once the path is understood, the work should become software.

  • His long-term test is unforgiving: if a workflow can be productized, productize it; if it cannot, “you’re just building a glorified consulting shop.” From his Palantir experience, he retained the more vivid internal maxim that forward-deployed engineers “eat pain and excrete product.”

  • Decagon’s engineers therefore contribute missing capabilities to the core product so the next 10 customers receive them automatically. Its agent PMs similarly translate enterprise deployment failures into product or repeatable process improvements, including how a large organization changes operations while handing meaningful work to agents.

  • Zhang cautioned that few startups can copy Palantir’s economics of landing enormous contracts and then spending heavily. A team promising to implement any AI use case may win near-term revenue, but without reusable output it has built a modern services firm. Decagon calls itself product-led, with sales supplying the evidence for what product to build.

6. Duo turns the labor of building an agent into another agent

  • Decagon’s first conversational agent required people to write agent operating procedures, build tools and API integrations, create simulations, and manually review live conversations. The fact that the delivered product was an agent did not remove the substantial human work required to construct, test, and operate it.

  • Duo is a second, “much bigger, much slower” agent that performs that surrounding work. A user can provide documentation and transcripts, ask it to infer the procedures, and receive AOPs, tools, tests, and simulations; after launch, it can inspect a thousand conversations, identify a weak topic, and draft improvements.

  • Zhang described the capability as “very magical” because it was impossible when Decagon began. Better reasoning models—built for broader use cases—became sufficiently general to perform Decagon-specific procedure writing, integration, testing, and monitoring despite not being trained explicitly for that workflow.

  • Every layer came from productizing deployment pain. Procedures first moved from code into plain-text AOPs; Duo then automated their creation; Duo Autopilot addressed the remaining labor of reviewing production traffic and iterating. The objective is progressively reducing how much customer-facing engineering each deployment consumes.

7. The near-term moat is making powerful agents governable inside enterprises

  • On the hosts’ AGI challenge, Jesse Zhang first rejected the premise that careers disappear: most modern jobs are already constructed layers of abstraction rather than food production or infrastructure. AGI may change their content, but people will still perform valuable work for other people.

  • Zhang located Decagon’s nearer moat in the software required to deploy even a hypothetically perfect model. Enterprises need enforceable boundaries on what it may do, collaboration among hundreds of domain experts, regulatory testing, controlled rollout, monitoring across millions of conversations, insight extraction, and connectivity to legacy systems.

  • His long-range answer remained explicitly uncertain. That infrastructure should matter for the next few years; if agents eventually generate it on demand and it becomes commoditized, “I don’t know, and we’ll figure out in three years from now.” The concession preserves the distinction between a current deployment advantage and a guaranteed decade-long moat.

8. A glass-box product is Decagon’s answer to services-heavy competition

  • The hosts framed the market as increasingly centered on Decagon and Sierra, while Zhang emphasized respect for Sierra and other capable platforms. His concrete contrast came from a recent customer that switched after finding Sierra’s forward-deployed model a “black box”: building journeys or understanding conversations repeatedly required going back through its engineers.

  • That customer had reportedly launched about three journeys over a year on the prior system. With Decagon’s “glass box” approach—where technical and nontechnical employees can inspect and change the agent—it launched seven new journeys in roughly one month.

  • Decagon also productizes the path to production, not merely the agent. For a regulated enterprise, it maps the likely model-risk review, testing program, staged rollout, issue-detection process, remediation, and controls that prevent recurrence. Large buyers must believe both “this works” and “we can actually get this live.”

  • Zhang estimated that roughly 80% of his time goes to sales. Founder involvement accelerates calls, but its higher leverage is redesigning product, process, and organizational sequencing—often starting with one or two high-value use cases—while salespeople build champions and navigate the account.

9. Speed comes from coupling enterprise sales directly to product judgment

  • Zhang said neither founder arrived with enterprise-sales experience; the hot market reduced the need to sell AI as a category, leaving Decagon to prove that its approach was right. Success required empathy for what buyers value, what they fear, and how decisions actually travel through a large organization.

  • The founders credit an unusually strong early commercial team, including cold applicants who already understood the market and people with nontraditional sales backgrounds. Scaling that group rapidly created continuing work in enablement and structure, but the company-wide premise remained that demand enters through sales and propagates into product.

  • Founder participation also keeps the feedback loop short enough for a market changing weekly or daily. Customers see new model capabilities elsewhere and immediately ask why Decagon cannot do the same; staying close reveals what the company must add before a conventional roadmap or reporting chain catches up.

  • That is also why a precise 12-month roadmap is unrealistic. The company holds longer-term themes, but its current view is that if a feature is already clearly valuable, it should be built immediately; customer evidence and changing model capabilities continually reorder the nearer-term work.

10. Customer support is expanding into an AI front door for the business

  • Support was Decagon’s first market because it combined urgent customer pain with the limits of early models. As instruction-following improved, the same system could accept broader guidance, tolerate conversations that “bob and weave,” and fill reasonable gaps instead of remaining inside a tightly scripted path.

  • One customer noticed that a support agent already knew its products, capabilities, brand voice, and customers, then extended it into inbound sales. The agent could answer questions, conduct discovery, and route sufficiently valuable opportunities to the appropriate enterprise representative.

  • Another customer used Decagon proactively for operational workflows when account problems appeared. The through-line was that Decagon never fundamentally built “an agent that does customer support well,” but “an agent that follows business process well”; support, sales qualification, and proactive operations are variants of that same primitive.

  • The longer-term “concierge” vision is therefore an AI agent as “the front door of your business or your brand.” Every reactive or proactive customer interaction could pass through it, with the product roadmap following observed customer demand rather than Decagon inventing an exhaustive set of interactions in isolation.

11. Company-building, not raw model capability, is now the binding constraint

  • Asked for the biggest bottleneck, Jesse Zhang answered “Hiring.” Agents can outsource specific execution steps, but he does not yet trust them with taste: deciding what to build, what to omit, or whether something is finished. Voice-to-voice systems and smarter small models still offer technical upside, but they are not the primary business constraint.

  • Zhang’s counterexample to the one-person-unicorn thesis is that AI coding companies—the most sophisticated users of coding agents—are hiring aggressively. If everyone can complete a roadmap in one-third the time, competitors do not stop; they attempt three times as much. Decagon’s hiring plan has therefore not materially contracted.

  • The founders reject “grind” as the objective. Engineers join sales calls, salespeople debug product, and agent PMs span both ends; the office supports that team sport and fast communication. Scaling the culture across offices remains “not a solved problem,” so new hires spend weeks in San Francisco and experienced employees seed new locations.

  • International demand is arriving unusually early because executive pressure is global and multilingual adaptation is easier with AI. Decagon still limits investment to markets with real customer pull because data residency and local competitors remain consequential; horizontally, Zhang expects consolidation because scale and product depth should outweigh narrowly vertical features.

12. Systems of record persist while AI lowers the price of personalized service

  • Sreenivas framed the concierge as democratized attention: a customer spending $100,000 may receive deeply personalized service, while one spending $10 cannot support the same labor. If AI can provide it for “10 cents,” the experience becomes economical—but the resulting preferences and history still need somewhere durable to live.

  • Sreenivas said CRMs “could do quite well.” A human concierge already records information in one, and an AI concierge will do the same; graphical interfaces may matter less when agents operate them directly, but the source of truth may receive more activity, not less. Decagon has “zero desire” to build a CRM.

  • Business context is also what limits personal AI. Sreenivas built an agent that continually compiles hiring needs, candidates, deals, and company challenges, allowing it to spot duplicated skill gaps or warn that a deal is repeating an earlier failure because key facts were not validated soon enough.

  • Better models have become useful sparring partners rather than automatic agreement machines. Zhang connected that to the advantage of having a cofounder: talking accelerates conclusions, while a context-rich model can now say, “No, that’s a bad idea. Don’t do that.” Sarah Wang said she used an aggressively contrarian Claude prompt, though her wife found it simply mean.

13. Cheaper support can expand demand even when particular jobs disappear

  • Sreenivas argued that many companies face more latent demand for customer help than they can economically supply. If support costs fall 30%, management may reinvest the savings in faster activation, better retention, and more availability instead of automatically shrinking the team by a corresponding percentage.

  • One early customer handled roughly 50,000 inquiries per month before Decagon. After discovering how many customers still needed help, it made support visible on more pages, placed it near likely friction points, and offered immediate assistance even to free users—the host’s candidate for a real-world Jevons-paradox example.

  • Zhang kept the employment claim hedged by customer circumstance. Some companies materially reduce or eliminate BPO usage; rapidly growing businesses may instead keep operations headcount flat, while others redirect people into higher-value or revenue-generating work. Cost reduction is the first easy use case, but conversational agents can increasingly support revenue creation too.

  • Sreenivas’s final distinction was categorical about tasks but not people: repetitive clicking and scripted phone work “should be done by AI,” while the possible work of improving customer outcomes is nearly unlimited. “AI will kill jobs but not careers.”