Pioneers Insight Method Research Author
Kavak's Playbook for Rebuilding a Company Around AI
Back to Episodes

Kavak's Playbook for Rebuilding a Company Around AI

Summary

  • Kavak has rebuilt itself so that between 100 and 200,000 agents are instantiated daily—one per customer, each with its own virtual machine, years of interaction memory, and a long-term goal of maximizing lifetime value. Head of AI Alejandro Maza says 96% of all interactions and 95% of transactions are now fully agent-handled, and frames the design question as “how would we build Kavak in 2035 with GPT-10-level intelligence.” The relational shift matters commercially: with 10M customers in the database, “just activating 1% of this customer base” could be worth hundreds of millions of dollars.
  • The agents aren’t support bots—they’re salespeople, and they now convert 2.1× better than Kavak’s human team while tripling NPS and customer satisfaction scores. Maza’s mechanism: a single “mega-expert” replaces 15 human specialists across financing, insurance, and trade-ins, is “infinitely patient,” and when one agent errs, the other agents learn from that mistake by the next day. In lending, Kavak usually approves car loans in under three minutes versus two months or more in Mexico and some emerging markets, with personalization of rate, risk, and loan amount.
  • Evals are the constraint on speed, not caution: Kavak spends roughly equal engineer time, tokens, and money on evals as on the agents themselves. “How fast can we go? It depends on the quality of our evals”—the car-brakes analogy—and the first checks are business outcomes: conversion, customer value and satisfaction, and willingness to re-engage.
  • When Opus 4.5 came out, Maza destroyed two years of working multi-agent architecture because “this isn’t the right paradigm anymore.” Thousands—possibly tens of thousands—of agents were running the business in December; he concluded that graphs and multi-agent latticework could constrain newer intelligence and rebuilt around a virtual-machine harness with memory, evals, a CLI, and long-term goals. His advice: “don’t build agentic workflows.”
  • An AI CEO has run the city of Cuernavaca for six weeks—it missed its goal of doubling profits but delivered 1.5×, or 50% more profits—micromanaging every number and messaging physical workers their daily plans. Maza says the remaining human-intensive work is mainly in the physical world: Kavak’s roughly 800 mechanics in Mexico get “El Mike,” a Ratatouille-style sidekick that helps with inspections, while warranties fell around 26%.
  • For enterprise buyers, Maza offers a token-quality framework investors can use to grade AI spend: tier 3 tokens go to agents with measurable per-token ROI, tier 2 are indirectly measurable, and tier 1 is raw Claude Code, ChatGPT, Cowork, or similar usage—“what happens with those? I have no idea.” Transformation must be top-down: “an army doesn’t really work if everyone comes up with ideas on the strategy”—hackathons and bottom-up use-case sponsorship “doesn’t work.”
  • The macro thesis is Schumpeterian: incumbents adopting AI superficially get 6–10% gains, while deep organizational redesign is needed to pursue 10× improvements. His electricity analogy: Ford’s production-line technologies existed by 1879 and 1881, but simply swapping a coal engine for an electric one produced about a 6% improvement; rebuilding the factory around electricity produced roughly 3× productivity. “It’s the innovator’s dilemma at an industrial scale,” and his founder advice is to “just map a trend that’s linear” in AI capability and build for that.

Deep dive

1. One agent per customer, up to 200,000 spawned a day — the bet that defines the company

  • Maza’s background frames the conviction: he founded Opi Analytics in 2013, pre-transformers—“we were 10 years ahead of our time”—serving Fortune 500 companies in risk algorithms, logistics, forecasting, and marketing, before joining Kavak’s vertically integrated used-car marketplace. Kavak also built a fintech, logistics operation, and Carfax, because the infrastructure needed to serve customers did not exist in LATAM.
  • The architecture: when a customer arrives, “an agent will get spawned specifically for this customer with its own virtual machine,” remembering years of interactions—web visits, a call from two years ago—and setting a long-term goal to maximize lifetime value. Between 100 and 200,000 agents instantiate daily, working “sometimes for three minutes, sometimes for eight hours, sometimes for three days,” then setting an alarm clock and going back to sleep.
  • Three decisions drove the transformation: redesign the company around agents by rebuilding most APIs rather than just handing staff ChatGPT; bet on superhuman agents by putting them in front of hard problems and generating data and feedback loops; and shift from transactional metrics—cars bought, cars sold, brake pads ordered—to relational ones. Kavak has 10M customers in its database, agents assigned to most of them, and estimates that activating 1% could be worth hundreds of millions of dollars.

2. Evals as brakes, agents as sellers — and regulated lending

  • Maza’s operating rule is to spend about equal engineer time, tokens, and money on evals as on agents. “I like to move extremely fast, but in order to move fast, you need to have brakes.” The first checks are business outcomes: did the customer convert, receive value, remain happy, and re-engage later?
  • Kavak “never built customer support or customer service agents. We built sales agents.” Selling a car in Latin America means choosing among 20,000 SKUs, arranging financing, insurance, and coverage, and quoting a trade-in—historically requiring 15 experts across 15 teams. The agent is a mega-expert across those functions: NPS and customer satisfaction tripled, conversion began 50% above the human team and now runs 2.1× higher, because agents “never get tired” and mistakes propagate to the other agents by the next day.
  • The hosts raised whether regulated financial services could be handled end to end by AI. Maza answered that the experience was “so much better,” describing car-loan approvals in under three minutes versus two months or more in Mexico and some emerging markets. Kavak personalizes the interest rate, risk level, and maximum loan amount while optimizing for the portfolio as a whole. If a borrower can no longer pay, vertical integration lets Kavak take back the car and offer a cheaper one, reducing the monthly payment.

3. The AI CEO experiment and the limits of the physical world

  • The hardest capability question—could AI do the CEO’s job?—was tested live. Kavak carved out Cuernavaca and installed an agent as CEO six weeks ago, with a first-month goal of doubling profits. Maza said it missed that target but reached 1.5×, or 50% more profits, by examining every number and customer, forecasting, and micromanaging execution. It messaged physical workers their daily plans and requested voice notes on their progress; inventory, financing penetration, and customer satisfaction all improved.
  • Maza sees remaining human-intensive work as mainly physical-world work involving dexterity and senses. Kavak’s roughly 800 mechanics in Mexico get “El Mike,” a sidekick he likens to the mouse in Ratatouille, which explains how to inspect cars, offers tips, and shows mechanics how to perform the work. Inspections and repairs became faster and cheaper, delivered-car quality improved, and warranties fell around 26%.

4. Jedi Academy, humans working for agents, and advice for incumbents

  • Everyone retrains through Kavak’s internal Jedi Academy, which Maza designed and continually upgrades because “you can’t send these people to Stanford to learn this.” From the CEO and AI engineers to finance staff and mechanics, employees attend; after six weeks, they launch state-of-the-art AI agents into production. Maza’s message is explicit: train for the new reality, or “maybe leave Kavak if this is not for you.” He says the program strengthened the culture and gave people practical skills for collaborating with agents.
  • The organization is now made up largely of flat, senior, empowered teams doing one of three things: “building the agents, working for the agents, or being in the physical world in front of the customer.” Sometimes agents are the bosses of humans; sometimes humans design the agents. In production, a stuck agent calls an “I need help” API answered by a human, rather than sending the case to Tier 2 and forgetting it. That closes the loop and generates data for improving the agent.
  • Maza’s counsel to executives is twofold: transformation must be top-down, with a clear picture of what the company should look like in three or five years—“an army doesn’t really work if everyone comes up with ideas on the strategy”—and token quality matters more than adoption. Tier 3 tokens go to agents with measurable per-token ROI; tier 2 has value that can be measured indirectly, such as developer work reaching the codebase and production; tier 1 is usage of Claude Code, ChatGPT, Cowork, or similar tools where “I have no idea” what the value is.

5. Destroying the working system, the self-improving organization, and Schumpeter’s warning

  • The riskiest call of the episode: thousands—possibly tens of thousands—of multi-agent systems were working at scale and running the business in December. Then Opus 4.5 came out, and Maza concluded that the graph, multi-agent latticework, and harness would constrain the newer intelligence. He destroyed two years of work that had brought Kavak to profitability and rebuilt around a virtual machine with an agent, memory, evals, and a CLI, designed to leverage recursive self-improvement and newer, more intelligent models. His advice is categorical: “don’t build agentic workflows.”
  • The RSI reframe is that “economic value in humanity for the past 4,000 years has been delivered by organizations, not individuals.” The loop companies should pursue is a self-improving organization that can use better models and intelligence to compound not only its intelligence, but also the economic value it generates.
  • Why he thinks startups, not incumbents, will capture this: Schumpeter’s creative destruction, illustrated with electricity. The technologies for Ford’s production line existed by 1879 and 1881, but merely replacing a coal engine in a four-floor factory of shafts and belts produced about a 6% efficiency gain. Rebuilding the factory on a flat surface around small dynamos produced roughly 3× productivity. Today’s superficial adopters may get 6–10%; deep redesign is needed to pursue 10×. “It’s the innovator’s dilemma at an industrial scale.”
  • Closing advice for founders: this is “the most exciting time in human history,” with access to powerful tools and intelligence for almost nothing or about $20 a month. No exponential extrapolation is required: “just map a trend that’s linear” in AI capability and build around it.