Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)
Summary
- Brex’s AI thesis is a three-part operating model: accelerate every corporate function, automate regulated financial operations, and sell agents that become part of customers’ own AI strategies. Reggio calls the internal platform “the thing that ties it all together,” creating a potential loop between lower service costs and product differentiation. The ambitions are specific: “10x” internal workflows and reach an 80% automated acceptance rate for startup and commercial applicants, with decisions inside 60 seconds and no humans involved.
- The clearest near-term economic unlock is making previously unprofitable commercial accounts economical to acquire and serve. Human-heavy onboarding made law firms, dental practices, and similar slower-growth businesses “ROI negative,” while Brex’s earlier high-volume small-business push became “almost existential.” The current lower bound is still selective: roughly $1 million in annual revenue or at least $10,000 in monthly card transactions.
- Brex concluded that useful agentic finance requires a network of specialists, not one assistant overloaded with tools. An employee-facing assistant delegates to travel, reimbursement, expense, and policy agents through multi-turn conversations; MCP-style tools connect those agents to conventional systems. Reggio’s framing is an “org chart,” with an executive assistant DMing specialists, rather than reducing collaboration to a single tool call or deterministic DAG.
- The AI product group is being run like a roughly 10-person startup with distribution to about 40,000 customers. Three-person pods pair a product- or customer-focused teammate, a staff-level Brex veteran who knows “where the skeletons are,” and a young AI-native builder unconstrained by established solution patterns. That structure lets Brex test what “a company that was founded today to disrupt Brex” would build without reorganizing its roughly 300-person engineering department.
- Brex refuses to pick a permanent winner among foundation models and coding tools, turning employee choice into both adaptation capacity and procurement leverage. Through ConductorOne, employees can provision ChatGPT, Claude, or Gemini and developers can choose among Cursor, Windsurf, and Claude Code. Usage becomes a market signal at renewal time: “our employees are voting with their feet. They’re voting with their dollars.”
- Reggio rejects the simplistic inference that more AI-generated code immediately means fewer engineers. The host cited Brex’s 5x growth and 99% burn reduction over 18 months, but Reggio emphasized broader execution discipline and said agentic development amplifies “all the good” and “all the bad”—including slop, weak architecture, knowledge drift, and harder incident response. His preferred outcome is still about 300 engineers a year from now, serving a much larger business at perhaps “30, 50, 100% more efficient.”
- Reggio’s account points to a potential advantage in Brex’s operating knowledge, evaluation loops, and cross-agent workflow design—and all three remain unfinished. Operations errors become regression evals, multi-turn product tests simulate users and apply an LLM judge, and the audit network divides detection, judgment, and employee follow-up among separate agents. Yet the assistant can still promise to “reach out to the finance team” when no such capability exists, while Brex’s product knowledge remains fragmented across internal, customer, sales, and support systems.
Deep dive
1. Founder experience, not backend pedigree, prepared Reggio for CTO
Reggio acknowledged that few leaders with front-end and mobile backgrounds reach CTO, but attributed his progression less to technical specialization than to founding companies twice. The CTO role, in his telling, is “a leadership and general business role as much as it is a technical role.”
He was considering leaving Brex to start another company when Pedro offered him the CTO job roughly two years ago. Brex now leans into that tension with “Quitters Welcome,” celebrating employees who later become founders or department heads rather than pretending retention must be permanent.
The pitch to former or future founders is “instant distribution”: they can build financial AI applications and deploy them across roughly 40,000 customers, from Fortune 100 companies to tens of thousands of startups. The organizational challenge is preserving enough startup texture that those builders do not feel swallowed by a corporate environment.
2. Brex built a startup-sized AI team beside its product organization
Brex has about 300 engineers and roughly 350 people across engineering, product, and design. Most engineers sit in 30-to-40-person, full-stack domains covering cards, banking, expense management, travel, and accounting, alongside shared infrastructure and security functions.
The exception is a centralized LLM group of about 10 people, up from four or five only months earlier. Its founding question was explicit: “What would a company that was founded today to disrupt Brex look like?” Reggio then used the answer to shape an internal challenger.
Its typical three-person pod combines customer or product intuition, a staff engineer who understands the existing codebase, and a younger AI-native engineer. Reggio’s provocative observation: “Too much experience or too much knowledge of how to solve a problem can actually be an impediment” to seeing an AI-first solution.
Centralization has not produced the resentment Reggio expected. Brex already optimizes engineering culture around measurable business impact; the card group, for example, drives about 60% of direct revenue. AI-tool adoption is also company-wide—one of Brex’s largest Cursor users is an engineering manager.
3. A January 2023 gateway became the base layer for two generations of agents
Reggio’s original AI labs team built an internal LLM gateway around January 2023 for deploying, versioning, and evaluating prompts; controlling data egress and model routing; and monitoring observability and cost. “Simple is elegant” remains his architectural preference.
That platform still powers precise operational applications, including research agents that help automate underwriting and KYC. Much of its interface lives in Retool, where operators can manage prompts, tools, and related workflows without waiting for engineers to mediate every refinement.
The newer customer-facing agent layer uses TypeScript, deliberately separated from Brex’s Kotlin- and Elixir-based backend through public interfaces. Its storage mix includes pgvector and Pinecone, while approximately half the current applications use Mastra and half use Brex’s evolving internal multi-agent framework.
Mastra won because its ergonomics resembled Brex’s existing framework, particularly around tracing and observability. Reggio expects continued stack churn: because agentic coding has reduced “the half-life of code,” teams can test technologies and migrate far more cheaply than before.
4. One omnipotent assistant failed where specialist conversations worked
Brex serves finance professionals and ordinary employees issued a company card. For the latter, Reggio’s desired experience is disappearance: “The best UI UX for Brex is just the card,” with SMS and an AI assistant eliminating expense documentation, policy questions, and travel administration.
The model is Reggio’s own executive assistant, who can infer business purpose from his calendar, email, and travel context. Brex wants a software equivalent for every employee, connected to the same kinds of contextual sources.
A single agent with many tools performed poorly across expenses, travel, reimbursements, procurement, and policy. Dynamically swapping prompt context also underperformed, so Brex split responsibilities into specialist agents behind an orchestrator, allowing each product team to improve its domain without redesigning the total system or making one team own every possible action.
Multi-turn delegation is the key distinction. A policy agent asked about a dinner limit might need to determine whether it is a customer event, team event, or travel meal; it tells the assistant what clarification to obtain, then resumes after the user replies. Reggio therefore treats MCP and tools as interfaces to “conventional imperative systems, not the AI space.”
5. Three AI pillars turn automation into a product feedback loop
Brex’s corporate pillar asks how purchased AI tools can “10x” workflows across every function. Its operational pillar targets the cost of running a regulated financial institution—fraud, underwriting, KYC, disputes, and support—while the product pillar builds features customers can cite as “part of our corporate AI strategy.”
Corporate adoption is led largely by IT and the people organization; Reggio concentrates on operational and product AI. The platform is an unofficial fourth pillar, supplying the gateways, tools, models, and interfaces reused across both internal automation and customer-facing agents.
Operations carries the fastest immediate impact because Brex employs hundreds of people in service-heavy workflows. COO Camila and Reggio are reframing those roles from executing SOPs to “build prompts, build evals,” and encode domain knowledge—while insisting automation must not degrade Brex’s high customer-satisfaction levels.
6. Vendor optionality doubles as product discovery and negotiating leverage
Brex’s deliberate policy is not to “pick winners in the horse race” among foundation models, chat products, or coding agents. Employees request approved ChatGPT, Claude, or Gemini access through Slack and ConductorOne; developers similarly assemble their preferred coding stack from options such as Cursor, Windsurf, and Claude Code.
Enterprise agreements preserve privacy and non-training guarantees, but Brex avoids mandatory wall-to-wall deployment. At renewal, actual adoption shows whether a formerly hot product has lost relevance, giving procurement a factual basis for reducing or reallocating seats.
The limiting factor has shifted from initial adoption to workflow inertia. Reggio observed that developers may decline to test a better but slower Codex because they have spent nine months mastering Claude Code: “I’m an iPhone person and I’m just going to stay with an iPhone.”
7. AI coding’s second-order costs are now more important than adoption
Reggio does not index on headline claims such as “80% of our code is written by AI,” because co-author metadata does not yield a defensible measure. Adoption is already broad; the current problems are “a little bit too much slop,” insufficiently rigorous reviews, and long-term maintainability.
Faster independent changes also create knowledge drift. Engineers understand their services less deeply as code evolves over months, surfacing during incident response when on-call staff encounter systems they did not meaningfully author or review.
The hosts argued that human attention cannot simply be replaced by placing an AI reviewer atop AI-generated code. Brex uses conventional linters, repository rule files, and Greptile, whose comments Reggio praised as unusually high-signal even when it leaves 65 on a diff.
After working “effectively 996” for a month inside the AI team, Reggio cycled from “this is going to change everything” to fears that engineers would disappear, then toward uncertainty. College students surprised him by using agents as co-architects for design documents while still writing much of the code themselves: “Everything looks like mentorship and management.”
8. Plain research agents beat a sophisticated credit-learning bet
Brex initially expected reinforcement learning to replicate a human underwriter’s credit-limit decisions and invested with an outside specialist. Reggio’s change of mind is categorical: its performance was “inferior to just building a web research agent.”
The reason is operational structure. Regulated teams already decompose work into granular, repeatable, auditable SOPs, which map cleanly onto prompts, tools, and sometimes even single-turn completions. The difficult work is extracting unwritten institutional knowledge—not inventing a more elaborate learning technique.
Brex prioritizes frequent workflows affecting the broadest customer base. Business-legitimacy research came before card-dispute documentation, where an issuer must assemble a three- or four-page Word document for the card network and acquiring bank; disputes are costly but relatively uncommon.
Automation supported a push into commercial businesses such as law firms and dental practices. Brex’s earlier volume-led SMB expansion left tens of thousands of ROI-negative customers and became “almost existential”; today it remains above true small business, generally requiring $1 million in annual revenue or $10,000-plus in monthly card spend.
9. Institutional knowledge is both the grounding layer and an unfinished liability
A base model’s picture of Brex can lag the actual company by years—describing it only as a startup card or, conversely, as enterprise-only. Agents therefore need curated product and process documentation to understand current capabilities, customer eligibility, and Brex’s ideal customer profile.
That knowledge is currently fragmented across internal operations and go-to-market documents, external customer materials, sales-oriented enablement, and Sierra’s support corpus. Reggio wants these applications to draw from a unified source because maintaining parallel truths is “wasteful” and increases hallucination risk.
Brex nevertheless buys Sierra rather than recreating it. Reggio considers customer support insufficiently differentiated to justify building every layer, while Sierra gives CX operators a low-code, workflow-oriented administration interface plus reporting and telemetry in “the language of customers.”
10. Evals are becoming production infrastructure, but guardrails remain lighter than expected
Operational agents launch with eval sets co-developed by an engineer and subject-matter expert. Existing QA continues after deployment, and almost every discovered mistake becomes a regression test—mirroring the QA process applied to human and LLM decisions.
Multi-agent product evaluation is harder. Brex gives a simulated user-agent an objective, runs a multi-turn conversation, and applies an LLM judge afterward; handwritten preambles can isolate narrower behaviors when a full conversation would resemble an overly broad integration test.
Accuracy failures can block release, while tone and coherence are tracked over time as metrics. A host proposed preserving tests for unsupported behaviors until models or products can satisfy them; Reggio embraced the idea as a way to show the assistant progressing from persistently red evals to eventual capability.
The sharpest failure is invented delegation: an assistant may promise to “reach out to the finance team” despite having no team or tool to contact. Brex has mostly fought that through system prompts. The host was surprised that hard guardrails remain uncommon even in finance; Reggio said the gateway supports circuit breakers but that he did not believe they were currently being used.
11. Fluency, headcount, and agent networks remain open organizational bets
The discussion referenced a four-level AI-fluency framework—user, advocate, builder, and native—and Reggio said operations is ahead of engineering in structured training. Leadership’s message is frank: many responsibilities will disappear, but “we don’t anticipate that meaning that your job has to go away. It’s just that your job has to change.”
Brex reinforces adoption with spot bonuses and an AI showcase at its company all-hands every two weeks, usually featuring operations, finance, or people teams rather than engineers. Its revamped engineering interview requires agentic coding, and every existing engineer and manager retook it without pass/fail records to expose personal skill gaps.
Reggio will not yet translate AI into a layoff formula or a junior-versus-senior prescription. Brex has held engineering near 300 while expanding customers and product lines, and he would prefer the same headcount with far greater output; whether cuts elsewhere reflect AI or ordinary performance management remains unresolved.
His strongest product conviction is the agent network. An audit agent zealously flags patterns such as repeated $74 charges below a $75 receipt threshold; a review agent applies judgment; then the employee assistant gathers context. The topology is “a tree more than it is a graph” until finance agents communicate with employee assistants—evidence, for Reggio, that deterministic DAGs undersell fluid agent-to-agent planning.