Pioneers Insight Method Research Author
Back to Pioneers
Lin Qiao
Founders 2 Curated Dialogues

Lin Qiao

Fireworks · CEO

Frontier Insights

Core Thesis: Enterprise proprietary data will fragment intelligence into millions of task-specific models; frontier labs will act as commodity power grids rather than monolithic intelligences. As token costs fall 10x and volumes scale 100x, value shifts from raw model size to specialized inference.

Strategy: Win via full-stack co-design—coupling inference serving, custom evaluations, and RL feedback loops to capture high-margin, private deployments over generic API resale.

Warning: Most AI startups risk scaling into bankruptcy. Nominal token price cuts mask real compute constraints; hardware depreciation, transistor limits, and task-completion inefficiency will crush companies lacking proprietary architectural efficiency.

Key Views & Dialogues

Most AI Startups Are Scaling Into Bankruptcy | Lin Qiao, CEO of Fireworks

  • 🗓️ Date2026-08-03 | 🎙️ Show:Gradient Dissent

Fireworks says it processes over 40 trillion daily prompt-and-generation tokens, with 95% from customized models, alongside a $1.5 billion Series D at a $17.5 billion post-money valuation. Its co-designed training-and-serving stack searches more than 100,000 inference options per workload, but “scaling to bankruptcy” makes value-per-task economics and proprietary data the durability tests.

View Dialogue Notes & Key Takeaways
  • Fireworks says it processes more than 40 trillion prompt-and-generation tokens daily, with 95% coming from customized models or deployments. Lin Qiao says that is “bigger than OpenAI’s API and Gemini’s API” based on what Fireworks knows, while conceding that accounting may differ. Alongside a $1.5 billion Series D at a $17.5 billion post-money valuation, the claim positions customization as scaled demand rather than an edge experiment.

  • Qiao’s central thesis is that intelligence will fragment into millions of models—“one per application, per use case”—rather than settle into a frontier-lab duopoly. Most valuable data lives inside applications and enterprises and “should never be shared,” so companies can turn customer behavior and business logic into continuously updated proprietary models. Fireworks expects specialized and general intelligence to coexist.

  • AI product-market fit no longer guarantees a durable business because AI infrastructure and inference costs can make growth look like “scaling to bankruptcy.” Open models may list roughly 10x cheaper, but their 1.5–2x greater verbosity leaves observed savings closer to 5–6x per completed task. Qiao sees the industry moving from “token maxing” to “value maxing,” particularly as public companies must defend AI features whose ROI remains TBD.

  • Reinforcement-learning fine-tuning turns model development into a product discipline built around proprietary evaluations, rewards, and feedback loops. Fireworks supports researchers controlling every parameter, as with Cursor’s Composer models, as well as an SDK for less specialized teams; Doximity combines methods including SFT, DPO, KTO, and RL for medical research. Qiao’s governing insight, prompted by a conversation with Jensen: “There’s no specialized general company.”

  • Coding was Fireworks’ dominant workload last year and is now the enabling layer for a much wider application cycle. Work that once required strong product teams and multiple quarters can, in Qiao’s telling, take one person knowing nothing about writing code only weeks. That velocity is producing general-purpose knowledge-work applications, vertical products across legal, finance, recruiting, marketing, sales, and support, and possibly a consumer search-and-recommendation unlock next year.

  • Asked whether American companies should fear Chinese models, Qiao declined to collapse geopolitics into the open-versus-closed question. Her stronger call is that American companies should release their best models, and she explicitly urged OpenAI to do so; Biewald’s question had also named Anthropic. Lukas pressed the obvious contradiction with R&D monetization; Qiao said openness should have a viable business rationale.

  • Fireworks argues its moat is not generic inference but a co-designed training-and-serving system optimized for “one size fits one.” It claims bitwise-equivalent training and inference, globally disaggregated training runs of up to tens of thousands of GPUs, and a search space of more than 100,000 inference options searched for each workload. Day-zero releases matter because customers distrust public benchmarks and must test new backbones before the next model arrives.

  • The company’s open-model advocacy does not extend to its own proprietary training and inference engines, a tension Biewald repeatedly surfaced. Qiao says rapid internal change makes community management unproductive and prefers supporting vLLM and SGLang over creating another rival project. Her broader operating formula is extreme ownership, flat teams, candid pre-mortems, fast decisions without sufficient data, and marketing that remains authentic enough to preserve technical credibility.

  • 🔗 Original source & video: Most AI Startups Are Scaling Into Bankruptcy | Lin Qiao, CEO of Fireworks

Listen to full conversation →


The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

  • 🗓️ Date2026-07-20 | 🎙️ Show:20VC

Fireworks argues that private enterprise data favors millions of specialized models over one AGI, with open weights offering control and customization once models cross the quality threshold. It processes 40+ trillion tokens daily and expects 10x lower token costs to drive 100x usage, while $800M in AR and 30-40% gross margins reflect rapid scaling and leave infrastructure bottlenecks in energy, chips, and transistors.

View Dialogue Notes & Key Takeaways
  • Lin Qiao’s founding thesis is a direct challenge to AGI maximalism: “if you think intelligence is a derivative of data,” the majority of the world’s data is private, “locked inside enterprise,” and will never be shared — so the frontier is specialized, private intelligence, and the endgame is “millions of specialized models — one per application, per use case,” not one model that rules everything. Fireworks already processes 40+ trillion tokens a day, the majority from customized models, not off-the-shelf ones.

  • On Harry’s are-OpenAI-and-Anthropic-overvalued question, Lin reframes frontier labs as “power lines” — essential infrastructure that won’t replace what’s built on top — and leaves the investor question standing: are power lines good businesses when open and closed models have both “crossed the quality threshold” and open weights can be tuned on small proprietary data to beat general models on your eval? “What I don’t want to see is there’s only one company owns intelligence.”

  • The headline economics call: token costs haven’t fallen yet “because of supply chain constraint,” but competition will deliver 10x cost reduction in the next three years, driving 100x usage — and Fireworks’ own token count could go “anywhere ranging from 20 to 100x” by end of next year. “We’re at a very early stage of S-curve explosion.” Harry’s inference — a capex bubble thesis is “ridiculous” — she endorses, with the caveat that the real bottleneck is the bottom of Jensen’s five-layer cake: energy, chips, “we’re bottlenecked by small parts — transistors.”

  • “Scaling to bankruptcy” is the mechanism pushing enterprises to open weights: PMF and durable business have decoupled because inference isn’t a commodity like CPU was in SaaS — incumbents with huge traffic can’t get AI features past the CFO, so they must own and tune their own models. Her three-year non-consensus: “every single company will own their own intelligence as a must-have. It’s not optional.”

  • On app companies training their own models (Harvey vs Legora, where Harry disclosed his Legora position): Cursor pioneered tuning and “now almost all coding companies tune their own models” — the tool-calling harness “needs to be co-trained with the model powering it.” Harry said Fireworks’ CTO Dima was embedded at Cursor for months; Lin described decoupled RL infrastructure running across five-six data center regions on scattered GPUs instead of an InfiniBand supercluster.

  • The business: $800M in AR, expecting to “at least double” by year-end with just 200 people. The 30-40% gross margin (vs SaaS’s 80%) is framed as a hyper-growth choice, not the new normal — “constraints slow down innovation.” Hard line: “we absolutely are not going to move into application layer”; data centers are “always on the table — the question is timing”; chips are out because workloads are too dynamic and hardware depreciation math has broken (“within a year from one vendor alone we have three SKUs”).

  • Chinese open models (the top six on OpenRouter) get a pragmatic answer: guardrail every model, open or closed, because every provider “will infuse their own judgment, their own taste” into training. If China restricts access, it’s “a big impact in the short term” — but the US “will be able to build that open system by ourselves, and we should” (Nvidia is training a model likely called Nemotron amid that supply-chain gap).

  • 🔗 Original source & video: The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Listen to full conversation →