Pioneers Insight Method Research Author
Fireworks Co-founder Benny Chen on Open-Source Models and Inference
Back to Episodes

Fireworks Co-founder Benny Chen on Open-Source Models and Inference

Summary

  • Fireworks co-founder Benny Chen says open source is catching up faster than he expected: the platform now handles 40-50T tokens a day, more than the B2B API volumes publicly disclosed by Gemini and OpenAI. But distillation “is not a particularly reliable path” — it depends on a security gap in closed-source APIs, and “this can all be fixed very quickly”; as long as open-source evals are large and clear and data prices fall, open-source success is not contingent on distillation.
  • Closed-source models’ long-term advantage may lie more in user mindshare and go-to-market than in capability leadership alone. “No one gets fired for buying IBM” — enterprise procurement is conservative and mindshare changes slowly, which gives closed source an advantage over the next 2-3 years; most of the eventual run rate may go to open source, while closed source can retain price-insensitive customers. The share paid to cloud providers by SOTA vendors may keep expanding, meaning “ARR can rise while GAAP revenue does not necessarily keep rising.”
  • He explicitly takes the vertical-agent side: building one vertical well and solving the Riemann Hypothesis “may be two completely different kinds of work,” and Frontier Lab may care more about the latter. Under the host’s hypothetical, a closed-source vendor going all-in on legal would “definitely” beat Harvey/EvenUp, “but that would not support a $1T valuation”; many fine-tuned models displaced by new closed-source models last year have held up this year — “far fewer have been washed out,” which is a very positive signal.
  • Most work is being rebuilt as coding: even PowerPoint and Excel are coding problems, and vertical tools may be turned into RL targets. The counterintuitive incremental demand is coming from “taking care of liberal-arts users” — slides, Word, and video generation are all taking off; computer use left him “astonished” 1.5 years ago but now looks like a smaller vertical, because text models are cheap and companies are pushing workflows toward pure text to cut costs.
  • Token consumption alone is “a little deceptive”; revenue is the most truthful signal. Willingness to pay flows to the best models, and Fireworks’ benchmark for custom models is “can it beat Opus on this task”; open source still has low penetration in coding, while Claude Code/Codex ARR is far above open-source counterparts — fundamentally a go-to-market issue more often than a technical one.
  • Eval is the scarcest muscle on the enterprise side: even in 2026, many companies still run POCs by vibe, a major reason POCs fail. “People who can do evals are truly few and far between”; writing good evals and teaching a boss how to choose between open- and closed-source procurement “can save 3-5x your salary”; the difference between old and new SaaS is not that large, with the task simply shifting from writing unit tests to writing evals.
  • On moats and macro, Fireworks is currently taking the “sublessor” route, avoiding hardware and aligning incentives with customers’ inference traffic; the endgame boss is the CSP. Large players are not first calculating ROI; they fear that missing this wave means never getting back to the table, so overspending is not an existential risk. His indicators are bond ratings and whether the market has money to rescue a player at its worst moment; open source catching up “absolutely favors infrastructure providers” — when DeepSeek launched last year and NVIDIA fell, “I didn’t understand it either.”

Deep dive

1. Open Source Is Catching Up Faster Than Expected; Distillation Is Unreliable

  • Benny’s starting point is that the team largely came out of PyTorch/Meta and has always believed valuable software layers will eventually catch up through open source — operating systems and databases are precedents. “We just didn’t expect it to catch up this fast”: K3, DSP4 Flash, and the latest Muse arrived one after another, while Meta has also said it will open-source Spark. “The competition on the open-source side is this intense, and the models’ price-performance is this strong — I was genuinely surprised.”
  • His view on distillation is blunt: it exploits a security gap in closed-source APIs — feeding large-model outputs back into smaller models and extracting thinking tokens. “This can all be fixed very quickly; there is no reason closed-source vendors cannot do it well.” “In the long run, distillation is not a particularly reliable path”; Meta appears to have done relatively little distillation, and Nemotron may not have relied on much either. With large, clear evals and falling data prices, “distillation may not be necessary.”
  • The gap is also narrowing for the non-distillation camp. Meta’s lead is “relatively speaking, not that large anymore”; it has talent and GPUs and can continue to scale. Nemotron, Reflection, and Thinking Machines all have pieces that can catch up, and “the gap will keep getting smaller.”

2. Closed Source’s Long-Term Advantage May Be User Mindshare, Not Just Capability

  • His Apple/Android analogy: he personally finds Google phones “much better to use,” but still buys an iPhone when price is irrelevant. “The lion’s share of profits may still go to whoever remains ahead, which is entirely compatible with open source becoming very large.”
  • Enterprise mindshare changes more slowly. “No one gets fired for buying IBM”: as long as people believe buying Claude or OpenAI means getting the better model, the shift will be slow. “That may be their advantage over the next 2-3 years. Five or 10 years out, I genuinely can’t see it.”
  • ARR can continue to rise while the accounting becomes less certain. Anthropic may be counting close to 100% of some revenue, while the share paid to cloud providers could keep expanding — “ARR can rise, but GAAP revenue does not necessarily keep rising.” SOTA vendors’ bargaining power “will weaken somewhat,” though visibility may improve after they go public.

3. Taking the Vertical Side: Jagged Intelligence and the McDonald’s Analogy

  • Making a vertical work requires extensive testing and a rich language for expressing evals. Harvey is a particularly well-known legal agent; Doximity is North America’s largest ChatGPT-like application used by doctors, focused on Deep Research and ensuring doctors see the correct references; Heidi is the largest medical scribe/search company outside North America. Benny believes a fintech lab trying to handle all these specialized workflows would face extremely high costs and potentially very low ROI.
  • The host posed a pointed hypothetical: could a closed-source model vendor going all-in on legal beat Harvey and EvenUp? Benny’s answer was “definitely,” but “that would not support a $1T valuation” and would not be a smart use of capital for those companies. Vertical tuning produces jagged intelligence: a model can be excellent in one direction, but “can it solve the Riemann Hypothesis? I think that is absolutely impossible.”
  • His restaurant analogy is that a great chef makes one dish exceptionally well before charging a premium; an AGI company’s job is to make McDonald’s work. It can localize the menu for China, but cannot simultaneously perfect the California and Kentucky flavors. “If it changes all of them, I don’t think it is fundamentally different from In-N-Out.” An AGI company that really serves every taste “will definitely make a great deal of money; we just have not observed that yet.”

4. Token Flows: Most Work Is Being Reframed as Coding, While Liberal-Arts Demand Rises Counterintuitively

  • The figure disclosed alongside the financing was 40-50T tokens a day, “larger than the numbers Gemini and OpenAI have disclosed” — meaning the volume of open-source models on the platform already exceeds the B2B API volumes of both companies.
  • The biggest workloads are coding, medical, and co-work, but “most work has been reframed as a coding problem”: identify a vertical SaaS tool, turn it into an RL target, and train the model to use it well. “We genuinely did not realize that PowerPoint and Excel were coding problems.” What is not coding is generally Deep Research; those are the 2 dominant categories.
  • The counterintuitive growth is in serving liberal-arts users. “Most of our time is spent taking care of STEM users,” but demand for slides, Word, and video generation is rising, and “most work may ultimately be about serving people.” OpenAI pushing the math benchmark to very high levels is impressive, but “I am not entirely sure what the significance is.”
  • Computer use has reversed from its initial promise. When it appeared 1.5 years ago, it was “astonishing”: models no longer needed to call an assortment of messy APIs and could simply use a computer to do anything. Now it looks like a relatively small vertical because “most verticals are coming from the text side” — text models are cheap, so companies are moving toward text to cut costs, and Chinese open-source models have reached high ARR through pure text. Vision models remain “very promising” as prices fall and capabilities improve, but the path is still being worked out.

5. Who Is Being Used: Five Chinese Players Rotating Through the Top Spot

  • “It is hard to rank them; whoever just launched sees its traffic jump.” The 5 Chinese players — Kimi, GLM, Qwen, MiniMax, and DeepSeek — rotate through the lead. “For example, if today were August 2026, Kimi would definitely have the most traffic.” Overseas, the leaders may be Nemotron, Llama, and Gemma; it used to be Llama, while Muse may now account for more usage.
  • Public router rankings are noisy because of free traffic. Some companies offer incentives for people to use their models free on OpenRouter, distorting the observed shares.

6. Token Volume Is Deceptive; Revenue Is Honest

  • He volunteered the caveat himself: “Token consumption is not that honest… I think looking only at token consumption is a little deceptive.” Fireworks cut the cost of each token last year from 100 units to 1 unit, but infrastructure grew faster — “it may already be at 70 units.”
  • “Revenue is the most truthful signal.” Willingness to pay moves toward the best model; when buying specialized intelligence, customers care less about price-performance than about the absolute best result. The benchmark is “can this beat Opus on this task,” measured against frontier models and customer revenue.
  • Open-source penetration in coding “has still been relatively low lately.” Twitter is full of jokes about “vibe coding with DeepSeek V4 and never using it up,” but Claude Code/Codex ARR remains far above open-source counterparts. “A lot of this work is fundamentally go-to-market, not technical” — which is also why Fireworks is pushing Fireworks Nexus.
  • TAM needs to be split apart. Token volumes “will definitely keep rising,” while the direction of revenue is uncertain. The General Electric analogy: everyone eventually used electricity, but the most profitable company was not necessarily General Electric.

7. Enterprise Procurement: Trust, Multi-Tenancy, and the Missing Eval Muscle

  • Trust is the core of procurement. Contracts run 1-2 years, and customers “do not particularly care whether we are actually just a middleman”; they care whether the models the platform serves can fulfill their future needs. Closed source has built trust more effectively — it can post on Twitter every day that it hacked another system or solved the Riemann Hypothesis. Open-source model companies remain far weaker at go-to-market, while most mindshare is still with closed source.
  • Deployments are shifting from on-premise to the cloud. The underlying issue is trust plus the economics of multi-tenancy: most enterprises cannot saturate the GPUs they rent, so aggregate demand allows customers to share costs. “Both sides get better ROI.”
  • Unclear evals are a major reason POCs fail. Even in 2026, plenty of customers still operate on vibe — “try it and see if it works” — even when AI accounts for a large share of call-center operating costs. “Many engineers are still worried about losing their jobs. I think people who can do evals are truly few and far between… Write the eval well, and you can save 3-5x your salary.”
  • His deliberately provocative take is that the difference between new and old SaaS companies is not even that large. Previously, a SQL engine had to pass unit tests; now an agent has to pass eval before delivery. Both require institutionalization — the transition is about how many people can adapt, and how quickly.

8. Inference Optimization: Caching, Workload Distribution, and 0.83 Calls per Task

  • Inference optimization “still has a long way to go.” The gap between a model just after launch and its best performance after a year of optimization is substantial; Fireworks’ job is to compress that timeline. Caching is a focus, and “the essence may not be in the serving runtime itself but in the supporting infrastructure” — an extremely labor-intensive task.
  • The advantage over a 1P API comes from workload distribution. Most of what Fireworks serves is customized models, so it sees more varied and messy workload and traffic patterns. It can even tell open-source model vendors about workloads they have not tested.
  • A joint Fireworks-Harvey study cited by the host found that a frontier model is consulted only 0.83 times per task on average, yet the system outperforms using the strongest model alone: Opus serves as executor and GLM as advisor, splitting long-context tasks into short-context ones because LLM performance declines as context length increases. Benny compares it with CPU optimization: treat the underlying model as a text processor, design the harness around its characteristics, and “put the most suitable model in the most suitable position.”

9. Hardware’s Real Question: Who Can Make Models Overfit to Its Own Cards

  • The core issue is the dimensionless ratio between compute and memory; he recommends Dylan Patel’s whiteboard video from 2-3 months ago. The target hardware is effectively decided before model training begins. NVIDIA and AMD have similar ratios, so migration is relatively easy; moving to an ASIC “takes a huge amount of work.”
  • The question therefore flips: if someone has extremely cheap hardware at scale and a good ecosystem, and can offer enough incentives to make model builders overfit to its cards, “whoever can provide that incentive can succeed regardless of how strange the ratio is; whoever cannot will fail even if the ratio is identical to NVIDIA’s.”
  • He used Jensen Huang’s line to make the point: “If you have electricity today and do not buy my cards, you could lose your shirt.” Buying the cards is almost cost-neutral and they pay back quickly; other cards do not have that characteristic.

10. Customized Models: Who Should Customize, Potentially Thousands of Environments, and Revisiting the Open-Source Release Cycle

  • The standard line is that vertical SaaS companies are the right customers for customization. Doximity understands the needs of 100,000 doctors, while Harvey covers many types of law firms, giving them sufficiently broad evals. A single hospital or legal practice is not recommended; it is better off codifying the workflow into a skill or building internal tools.
  • GPU is the main cost, followed by data. The work with customers is to examine reward-hacking behavior, environment stability, and how to scale the environments. The required data volume “may be as few as 1,000 environments”: during RL, the model keeps exploring, and if each step runs 128 rollouts, 1,000 environments can generate 128K data points.
  • Re-training and replacement are mandatory: when a new model launches, Fireworks helps the customer replace the old one. Once the eval is fixed, replacement is fast, and the cadence “depends on the open-source model release cycle.” Many customized models were displaced by new closed-source models last year; far fewer have been this year — “the direction of progress in many verticals may be somewhat different from the direction in which frontier labs are moving.”

11. Moats: The Sublessor Model, Inference Incentive Alignment, and CSPs as the Final Boss

  • Fireworks is not buying GPUs today: “We are more on the sublessor side,” building the software layer on top of compute. Its difference from Nebius and similar companies is a genuine belief in specialized intelligence: “most traffic is not vanilla base models; it is customized models,” and “most profit is on models that perform well.” Customers can charge more, allowing Fireworks to take a larger share.
  • Its difference from AI services or consulting companies is incentive alignment. Consulting firms and pure training-infrastructure companies are aligned with training traffic; Fireworks is aligned with customers’ inference traffic. “They make money through inference, and we can take a share of the money they make.” The real opponent is not a new cloud but the CSPs — “they are the final boss.” Fireworks is currently closest to Azure: on Azure it is 1P, and customers can directly use Azure credits to buy Fireworks services.
  • On the possibility of RSI, he was candid: if AGI automates kernel optimization, inference optimization, training, and post-training, “I do think there would be little core competitive advantage left.” The difficulty is that AGI would still have to talk to customers about their needs. “It may indeed solve the Riemann Hypothesis, but I am not sure it can solve the customer’s problem.” The biggest challenge over the next 1-2 years remains whether the market can provide enough compute at a reasonable price.

12. Capital Markets: The Cost of Missing the Table, Bond Signals, and “Open Source Favors Infrastructure”

  • The capital-allocation logic for large players needs to be reversed. Even if ROI is high, missing this wave “may mean never getting back to the table.” Overspending may make the books look bad for 2 years and take 5-10 years to absorb through depreciation, but “that is not an existential risk.” Google and Meta are not asking first how high the ROI is; they are asking, “What if I do not make this investment and I end up finished?” New clouds “need to be a little more honest.” At least over the next year, he expects compute constraints.
  • The observable signal he uses is the bond market. Ratings are transparent, and project delays pressure ratings. The real test is whether, if a player runs into debt trouble within 3-6 months, cannot borrow, and needs restructuring, there are enough people in the market with the money to rescue it at its worst moment. “As long as it can be rescued, the problem is not that big; if it cannot be rescued, it is genuinely difficult.”
  • He rejects the logic that open-source catch-up automatically means a hardware correction. “I really cannot work out that relationship… Open-source models catching up absolutely favors infrastructure providers.” As model-layer bargaining power weakens, and if the total pool of capital remains unchanged, money should flow more favorably toward infrastructure-adjacent companies — “the stock should rise, not fall.” When DeepSeek emerged last year and NVIDIA sold off, “I also did not understand how these people were thinking,” followed by his self-deprecation: “I am not suited to trading stocks.”