
张璐
Frontier Insights
Core Frontier Thesis: Open-source AI (e.g., DeepSeek) is structurally leveling the playing field with closed-source frontier models, drastically slashing unit economics and accelerating enterprise penetration. Concurrently, compute dominance is shifting toward absolute vertical sovereignty—exemplified by Musk’s cross-entity, orbital-scale industrial operating systems.
Strategic Imperatives: Prioritize high-margin, verifiable B2B agentic workflows (healthcare, industrial automation) where proprietary data moat-building outpaces raw foundation-model hype, rather than chasing generic consumer agents.
Risks & Warnings: Reliability bottlenecks in autonomous agent execution, escalating talent wars, and structural capital concentration threaten venture returns as hyperscalers commoditize baseline inference.
Key Views & Dialogues
159: Musk’s Terafab Space Compute, Nvidia Reclaims the CPU | Talking New AI Compute Trends with Zhang Lu of Fusion Fund
- 🗓️ Date:
2026-04-07| 🎙️ Show:晚点聊 LateTalk
Terafab combines Tesla, SpaceX and xAI around compute sovereignty, targeting 1 TW of annual capacity with 80%-90% deployed in space; radiation, maintenance, launch costs and demand maturity could delay the window seven to ten years. NVIDIA’s Vera CPU and Groq extend its heterogeneous AI factory, while inference could reach 60%-70%; Blackwell and Vera Rubin may support over $1T in 2027 revenue if token demand grows.
View Dialogue Notes & Key Takeaways
Terafab is not really competing for a chip fab; it is competing for compute sovereignty across Musk’s ecosystem. The plan is to integrate Tesla, SpaceX and xAI from chip design and manufacturing through deployment and applications, with an ultimate target of producing 1 TW of AI compute annually and deploying 80%-90% of it in space. Zhang Lu described it as “a giant cross-company industrial operating system,” but expects both the rollout timeline and funding needs to far exceed Musk’s own expectations.
Orbital data centers are not a mature startup bet over the next 2-3 years. Cosmic radiation, chip packaging, launch, maintenance and repairs all present major technical and cost challenges; serving Earth-based applications would also introduce latency. Zhang’s estimated maturity window is seven or even ten years. “If everyone wants to explore building data centers in space, why not build them in nearby Canada?”
SpaceX needs the entire space economy—not just rockets and Starlink—to support a trillion-dollar IPO narrative. Zhang believes orbital deployment also reflects Musk’s deeper goal: “He doesn’t want any government to regulate him.” With jurisdiction over space, orbital enforcement and lunar resource ownership still unclear, SpaceX could become both the main enabler of the space economy and a future rule-setter.
The nearer-term opportunity than orbital data centers is serving a space industry that is “AI native, robotics native.” Microgravity factories could explore crystals, materials and protein structures difficult to produce on Earth; AI could manage satellite traffic and data trading, while robots extract water, hydrogen and oxygen from lunar regolith to build lunar “gas stations.” Genuine demand for space-based compute will emerge only after these local applications scale.
The more investable opportunities today are the efficiency layer of terrestrial AI infrastructure and other space infrastructure. Interconnect, optical switches, power consumption, memory, security, inference costs and system integration all present clear bottlenecks; startups must also decide whether they will serve future hyperscale clusters or build capital-intensive clusters themselves. Zhang’s direct conclusion: “I don’t think it’s a good time. I think it’s still a little too early.”
Nvidia is rewriting itself from a GPU company into an “AI factory” that produces tokens, while inference turns revenue from one-off demand into recurring cash flow. The training-to-inference compute mix has shifted from roughly 70%-80% versus 10%-20% toward an even split, and could eventually invert to 20%-30% versus 60%-70%; Jensen Huang therefore expects Blackwell and Vera Rubin to support more than $1T in data-center revenue by 2027.
The Agent era will lift demand for heterogeneous compute across CPUs, GPUs, LPUs and NPUs, making platform control more important than winning on any single chip. Nvidia is using Vera CPU to cover tool calls, code execution, multi-Agent workloads, reinforcement learning and simulation, while adding Groq’s low-latency, high-throughput inference capabilities to Vera Rubin; even if its CPUs do not always beat AMD’s, control of the full system, rack and one-stop purchasing creates a major integration advantage.
Enterprise AI has entered a phase in which budgets, deployments and exits are all accelerating. Regulated industries favor private-cloud or on-premises deployments of small language models to balance accuracy, privacy, latency and cost; companies in the Fusion Fund network have reached AI budgets of up to $12B, while some sub-10-person companies grew annual revenue from zero to $20M. Finance, healthcare, insurance, industrials and supply chains could be among the fastest adopters, and some portfolio companies founded less than two years ago were acquired last year at 10-20x returns.
🔗 Original source & video: 159: Musk’s Terafab Space Compute, Nvidia Reclaims the CPU | Talking New AI Compute Trends with Zhang Lu of Fusion Fund
E230 | Behind the $1T Revenue Expectation: Nvidia’s Peak and Soft Underbelly
- 🗓️ Date:
2026-03-26| 🎙️ Show:硅谷101
黄仁勋 expects cumulative Blackwell and Vera Rubin orders to reach at least $1T by end-2027, underpinned by recurring Agent inference demand rather than one-off training. Seven mass-produced chips and full-stack execution strengthen NVIDIA’s position, but CoWoS, power and component shortages, plus Groq and TPU competition, remain delivery and architecture risks.
View Dialogue Notes & Key Takeaways
黄仁勋’s core expectation is that cumulative orders for Blackwell and Vera Rubin will reach at least $1T by the end of 2027, double the previous $500B benchmark. That is even higher than 2024 global semiconductor sales of more than $600B, but 张璐 believes the demand logic holds: training is a one-off investment, while Agent inference generates recurring cash flow; the cost mix could shift from training at 70%-80% in 2023 to inference at 70%-80% in the future—“Long-term cash flow must come from inference” (“长期的现金流,那一定是来自于推理”).
The hardest hurdle to the $1T expectation is not orders but physical capacity that capital cannot buy on demand. 肖志斌 sees 3nm wafer capacity as likely to keep up, while CoWoS advanced packaging is harder to call; capacity has grown roughly 3x since 2024, Micron and Samsung announced HBM2 mass production in March, and Micron, Samsung and SK hynix are advancing customized HBM4E solutions, but production-line investment, process optimization and validation still take 1-2 years. “This isn’t software: however much demand you have, you can’t instantly turn it into an equivalent amount of sales.”
NVIDIA’s moat has expanded from CUDA into full-stack execution spanning chips, systems, supply chain, developers and customer feedback. Vera Rubin launched 7 chips simultaneously, all already in mass production; NVL72 delivers 10x Blackwell’s inference efficiency, cuts token cost to one-tenth and raises token per watt 35x. Mark says NVIDIA internally went from no one using coding agents to “100% using them” within 1-2 months, but the real challenge remains turning generated designs into end-to-end-optimal designs.
Groq’s LPU shows that inference will not be dominated by a single chip type; future data centers are more likely to be heterogeneous systems combining GPUs, LPUs, Switch and optical interconnect. Groq uses on-chip SRAM to store weights and KV Cache, sacrificing capacity and cost for ultralow latency, especially for token-by-token decoders and Agents; 黄仁勋 even suggested reserving 25% of data-center space for Groq and other inference chips. 肖志斌’s blunt view: “GPUs are actually not very well suited to agentic applications.”
TPU is a real threat, but not yet enough to overturn NVIDIA’s position as a third-party infrastructure provider. Google can vertically optimize models, chips, interconnect, power and applications, but external customers may not replicate its internal efficiency; NVIDIA relies on cross-customer software, exceptional execution and control of TSMC and CoWoS capacity to defend its lead. The more consequential gaps are private deployment, Edge AI, robotics and specialized CPUs/NPUs, while its huge market cap means it is “celebrated by capital, but also held hostage by capital.”
OpenClaw and NemoClaw are pushing token demand toward always-on Agents and shifting SaaS’s unit of competition from software seats to AI labor. 张璐 believes vendors may move from IT budgets into the larger labor budget, but using hiring standards as a benchmark, she argues an Agent would need to complete more than 90% of a role and outperform more than 90% of people to truly replace them; there is still a gap. SaaS companies without model capabilities that fail to transform quickly face the highest risk. “Future software companies … become labor suppliers.”
Data centers have already validated extremely strong demand, but $1T in revenue ultimately depends on land, power, delivery speed and operational reliability. Alex says roughly 90% of new US projects are shifting to behind-the-meter natural-gas generation, while modular construction cuts the time from greenfield to live service from 18-20 months to 6-9 months. Memory prices are up 100%-200%, shortages are emerging in SSDs, ConnectX-7, switch gear, Intel CPUs and CDU liquid cooling, and CX7 is also shifting to BlueField; the supply chain is unlikely to ease meaningfully by the end of 2027. “Capacity exists, but the price is impossible to determine.”
🔗 Original source & video: E230 | Behind the $1T Revenue Expectation: Nvidia’s Peak and Soft Underbelly
E227 | The AI Battle for the U.S. Healthcare Market: Big Bets by Giants—Can Startups Win?
- 🗓️ Date:
2026-03-04| 🎙️ Show:硅谷101
U.S. healthcare AI is shifting from experimentation to organizational integration, with coding, billing and prior authorization offering structured, multibillion-dollar entry points tied directly to reimbursement. HIPAA, data control and near-zero hallucination tolerance favor compliant infrastructure, while OpenAI and Microsoft contest hospital workflows and startups retain room through vertical data and trusted deployment.
View Dialogue Notes & Key Takeaways
The inflection point for healthcare AI in 2026 is that large pharmaceutical and healthcare companies have moved from debating “whether to integrate” to accepting that they “must integrate.” Eli Lilly and Nvidia announced a billion-dollar-scale strategic partnership, while some healthcare companies have even established “artificial intelligence universities” requiring employees, management and even boards to undergo training. This means AI demand is moving into organizational processes and infrastructure.
The first value healthcare AI delivers in the U.S. may not be replacing diagnosis, but repairing the resource mismatch between doctors and insurers and documentation systems. MGH generalists work an average of 61.8 hours a week, yet typically see only 15 to 25 patients a day for about 15 minutes each; only around 10% of denied insurance claims are appealed, but roughly 80% of appeals are overturned, showing that much of the waste comes from process rather than medical judgment. “Why have such expensive people doing low-value repetitive work?”
medical coding, medical billing and prior authorization combine huge markets, clear rules and low tolerance for error, making them the most realistic entry points for AI. Coding is essentially “translating” diagnoses and procedures into standardized codes such as I10 and CPT codes; errors can lead to denials, delayed reimbursement and even upcoding fraud risk. The work is highly structured, and many steps do not even require generative AI; Gen AI’s incremental value lies mainly in checking supporting documents and predicting denial risk.
HIPAA, data control and deployment models are not ancillary compliance items, but the entry barriers and competitive moats for healthcare AI. Claude for Healthcare is focused on infrastructure including billing, coding, HIPAA-compliant cloud deployment and connectivity APIs; federated computing allows more than 60 healthcare systems to collaborate without physically moving data. The priority order is “compliance, security and no hallucinations” ahead of general-purpose model capability, leaving room for vertical small models and on-premise deployment.
OpenAI is competing simultaneously for the consumer gateway and the hospital operating system, but control over hospital workflows remains undecided. ChatGPT Health addresses more than 200 million healthcare-related questions asked by users each week, with data isolation, no training use and explicit disclaimers; ChatGPT for Healthcare aims to connect to EHRs, summarize medical records, draft prior authorization materials and let hospitals build agents. The issue is that Microsoft products such as Office and Outlook are already embedded in hospital workflows, so “adding a feature” may be easier than OpenAI rebuilding system integrations from scratch.
OpenEvidence proves that a narrow use case can quickly create product advantages, while also exposing the risks of a high valuation, advertising conflicts and low switching costs. It constrains RAG with top journals, authoritative guidelines and traceable citations, and reportedly is used daily by 40% of U.S. doctors; roughly $100M in annual revenue corresponds to a $12B valuation, while doctors use it for free and revenue depends on pharmaceutical advertising and content promotion. 张璐’s judgment is blunt: “I don’t think it is a company with core AI capabilities”; its true moat is content licensing and physician penetration.
AI can become a tool for healthcare stratification and resource balancing, but the episode insists on “human in the loop,” with doctors remaining the decision-makers. For common colds, test results, emergency guidance and care navigation, AI may provide a “reliable first step,” with rural and underserved areas benefiting most; complex diseases still involve individual variation, unknown biomarkers, liability and treatment trade-offs. O3 scored 60% on HealthBench, with a top score of 32% in the hardest mode, which is closer to real clinical encounters than a high multiple-choice score and makes the capability boundary clearer.
Healthcare AI is unlikely to be winner-take-all: giants are better positioned for foundational platforms, while startups can still win through vertical data, customization and trusted relationships. Healthcare data is fragmented across hospitals, pharmaceutical companies and medical-device makers, and customers may not want to tie all their risk to a single technology giant; at the same time, To C opportunities are expanding from treating illness after it appears to continuous monitoring, preventive health and quality of life. Demand is real and durable, but the key variable is still who pays, while free substitutes will continue to pressure consumer subscriptions.
🔗 Original source & video: E227 | The AI Battle for the U.S. Healthcare Market: Big Bets by Giants—Can Startups Win?
New Year Livestream 1: AI in 2025 and 2026, Consensus and Non-Consensus in Tech
- 🗓️ Date:
2026-01-15| 🎙️ Show:硅谷101
Enterprise AI has shifted from whether to adopt it to “How and how much,” favoring vertical models, localized fine-tuning and cocktail deployments. DeepSeek, Thinking Machines and SSI keep model research contestable, but a 2026 GPT-3- or GPT-4-level breakthrough is uncertain; monetization requires accuracy, proprietary data and profits.
View Dialogue Notes & Key Takeaways
Enterprise AI crossed the “whether to do it” decision threshold in 2025, but budgets did not automatically flow to the most expensive, most capable general-purpose models. The boardroom question has shifted from yes or no to “How and how much?” Vertical small models, localized fine-tuning and “cocktail” stacks have gained consensus on cost, privacy and regulatory grounds. 张璐 notes that financial services, healthcare, insurance and other service industries covering over 50% of U.S. GDP offer extensive opportunities for vertical AI Agents.
DeepSeek rewrote the model race from a “four or five big labs” oligopoly into one where open-source models and new labs can still change the outcome. 徐皞 says Chinese models such as Kimi can produce strong results at relatively low cost, and a growing number of U.S. companies have begun using them; Thinking Machines, SSI and other New Labs show that core research can happen outside the major incumbents. “DeepSeek is only the beginning,” but it is still “too early” to say whether 2026 will deliver a GPT-3- or GPT-4-level huge moment.
Scaling Law is not dead; the debate has shifted to who can afford it and how to realize it through systems engineering. GPT-5’s underwhelming performance briefly led the market to conclude that pretraining had hit a wall, while Gemini later restored attention to pretraining; 徐皞 believes data cleaning, allocation, domain knowledge, cluster interconnects and fault tolerance all remain far from exhausted, making “10x room” in pretraining entirely imaginable in hindsight. 张璐’s qualification is that Scaling Law still holds, but is no longer the only path and is increasingly something only a handful of players can afford.
Meta’s investment thesis is stuck on a real fork: should it fill the gap in foundation models, or first turn model capability into user experience? The host said Meta reportedly acquired Manus for roughly $2B-$3B, with talks lasting just over ten days. 张璐 sees Manus as a strong team on execution, product and data, but questions why Meta would buy an application company when it needs model capability more; 徐皞 argues Meta can simply call OpenAI or Anthropic APIs and still fill the model gap in 2027.
If OpenAI IPOs in 2026, its billion-scale user base and user inertia give it a very large story and substantial potential, but costs, margins and retention will determine whether the valuation holds. It has advanced its for-profit structure and arrangements with Microsoft, but 张璐 says Gemini’s training costs may be under 30% of OpenAI’s, while the latest YC cohort is also relying more heavily on the Gemini API; the path to profitability remains unclear. 徐皞’s view is that OpenAI is “mainly a story”: no single business is nailed down, yet “every one of them makes me feel there is huge potential.”
Anthropic is the more focused winner in enterprise APIs and AI coding, but the guests explicitly rejected the idea that it has already built a moat. Very few Fortune 500 or Fortune 1000 companies use only OpenAI, 徐皞 says; even if Anthropic has not fully overtaken it, it is on the way. 张璐 believes Claude Code and Anthropic’s safety and To B positioning could keep attracting regulated-industry budgets. The problem is that switching costs remain low: “In today’s AI era, no company truly has a moat,” and an IPO must answer both the high-growth and cost-optimization questions.
The more monetizable application opportunity in 2026 may be vertical Agents; the bottleneck is not the demo but industry accuracy approaching “all correct.” 张璐 says companies in insurance, supply chain and other sectors have grown rapidly from early stage to tens of millions, and in some cases hundreds of millions, in revenue; she is also bullish on healthcare and space tech. 徐皞 warns that industry customers do not accept systems that are 90% or 99% correct—they want them to be “all correct.” That forces startups to build extensively around proprietary data, fine-tuning, platforms and workflows, leaving room in niches that general-purpose models cannot directly cover.
🔗 Original source & video: New Year Livestream 1: AI in 2025 and 2026, Consensus and Non-Consensus in Tech
Mid-Year Review of 2025 Silicon Valley Tech Developments | Conversation with Fusion Fund’s 张璐: Community-Driven Innovation, the Talent War, VC Transformation, and the U.S. IPO Landscape
- 🗓️ Date:
2025-07-27| 🎙️ Show:十字路口Crossing
DeepSeek shifted AI’s center of gravity toward open-source, community-driven innovation, but scarce architecture-defining talent may command valuations that price out future returns. AI is increasing capital efficiency and pushing VCs toward multiproduct finance, while M&A repairs liquidity ahead of IPOs; the public market remains largely closed, favoring vertical to-B, healthcare, and industrial AI.
View Dialogue Notes & Key Takeaways
DeepSeek shifted the first half’s focus from “closed-source models leading” to a greater emphasis on open-source ecosystems, leading 张璐 to view Nvidia’s post-selloff outlook as bullish. Simpler architectures and training paths that do not depend on the most advanced GPUs could broaden AI adoption; the capital markets sent the opposite short-term signal, but she sees innovation moving from purely corporate-driven toward increasingly community-driven.
The host described AI agents as the next general-purpose platform after the PC and the internet; 张璐 emphasized that a true agent must handle complex tasks, make autonomous decisions, and choose its own tools. Coding agents such as Cursor and Windsurf depend on Claude, GPT, and Gemini, which means native products such as Claude Code could compress their moats; OpenAI once proposed acquiring Windsurf for $3B. The host argued that Google’s roughly $2.4B deal for the team, technology license, and know-how showed that, in Big Tech’s eyes, “money is cheap; time is the most valuable thing.”
张璐 identified OpenAI as the AI company facing the greatest challenges this year: public-facing C-end data is nearly exhausted, while the next step requires high-quality industry data and a direct fight with Microsoft in the to-B market. Microsoft remains constrained in the short term by its agreements and Azure ties, but it has already added models from Mistral and Cohere, strengthened internal development, and may allow Copilot to call multiple models dynamically; she characterized this as proactively preparing for “de-OpenAI-ization.”
Google’s AI technology and full-stack cost structure may be undervalued by the market, while Meta and Apple face, respectively, an ecosystem gap and a missing product. Google has TPUs, models, cloud, infra, and applications, but must decide when to dismantle search advertising, its “cash cow”; Meta is using cash, compute, and talent to “buy time,” and 张璐 remains bullish over the long term but does not think it can catch up by year-end, while Apple still has a hardware-entry window that is closing.
The talent war is not for ordinary engineering execution but for the “brains” capable of defining model architectures and product paths in the unknown, a group that may number no more than a few thousand people globally. Reported individual offers of $100M, Mira Murati’s company raising $2B in its first round, and Ilya Sutskever’s company raising $1B at a $5B valuation all reflect scarcity premiums; but 张璐 worries that first-round valuations have already priced out returns because “it’s not as if we’re going to have many trillion-dollar companies.”
AI is increasing capital efficiency while forcing large VCs to evolve from single-product venture firms into multiproduct financial institutions. Companies that once took three to five years to reach $2M-$3M in revenue can now reach tens of millions in a year; later-stage financing needs may fall from hundreds of millions of dollars to tens of millions, prompting large funds to expand deployment through RIAs, fund-of-funds vehicles, and other asset-management products.
M&A has become the primary liquidity repair mechanism for venture capital ahead of IPOs, but talent-driven deals may not allow investors to share fully in the value created. Fusion Fund had 4 exits in the first half, 3 involving important contributors to the open-source ecosystem; conventional strategic acquisitions can rapidly recycle capital and talent, while talent acquisitions often direct most consideration to the team, leaving VCs with relatively limited returns in Windsurf-style deals.
The U.S. IPO market cannot yet be described as reopened; the opportunities generating real revenue are concentrated more in vertical to-B, healthcare, and industrial AI. More than 20 Silicon Valley companies preparing to go public in the first quarter largely put their plans on hold, while CoreWeave and Circle remain only a handful of examples; meanwhile, U.S. healthcare accounts for roughly 20% of GDP, more than 30% of human data is healthcare-related, and less than 5% is being used, making this “data gold mine,” together with industrial and space automation, the deeper value of AI’s “shovels.”
🔗 Original source & video: Mid-Year Review of 2025 Silicon Valley Tech Developments | Conversation with Fusion Fund’s 张璐: Community-Driven Innovation, the Talent War, VC Transformation, and the U.S. IPO Landscape
E200 | An In-Depth Investor Conversation: Core Moats and Investment Logic for AI Agents
- 🗓️ Date:
2025-07-17| 🎙️ Show:硅谷101
Agent financing has reached $10B valuations and aggressive talent poaching, but premium pricing increasingly reflects strategic scarcity rather than proven application-layer moats. AI Coding faces foundation-model encroachment and a compressed first-mover window, while more durable to B opportunities combine high-value workflows with private data, layered processes, and clear willingness to pay.
View Dialogue Notes & Key Takeaways
Agent financing has entered the phase of $10B valuations and talent poaching by tech giants. But pricing increasingly reflects strategic scarcity, not proven startup moats. Cursor parent Anysphere raised $900M at nearly $10B; Google paid $2.4B to bring over Windsurf’s core team, while Cognition AI pursued the remaining team; Thinking Machines Lab raised $2B at a $12B valuation. 周炜’s warning is that global capital is crowding into a handful of quality assets, and “money isn’t money anymore”(钱都不是钱了). Big Tech will build in-house and acquire at a premium.
The Agent boom is being driven by advances in models, open-source ecosystems, and the first real end-to-end products—not by a concept that suddenly appeared. 张璐 defines an Agent as a system that can make decisions autonomously, handle complex multistep tasks, and complete work end to end with support from tool libraries and vertical knowledge bases. Over the past 6 months, even traditional businesses running on little more than Excel and email have become addressable. She emphasizes that open-source models have lowered the barrier to vertical fine-tuning and small-model innovation; 周炜 believes the industry has finally moved from debating foundational technology to building Agents that help users achieve an actual end goal.
The most dangerous illusion in AI Coding is mistaking first-mover advantage for a durable moat. Claude is one of Cursor’s underlying engines and also competes directly through Claude Code. 张璐 observes that “70% to 80%” of research-oriented developers use Claude Code, with some unicorns migrating from Cursor or Windsurf. Cursor’s UI/UX, context integration, and existing integrations create migration costs, but 周炜 believes a large model can fill in the gaps with “one night of training.” Without proprietary data or a genuine technical moat, “sell while you can”(该卖就卖了). 张璐 also classifies AI Coding as primarily to C, while Fusion Fund focuses mainly on to B opportunities sold directly to enterprises.
The investability of vertical to B Agents comes from high-value workflows, not from how large the industry sounds. 张璐 cites Walmart, which issues roughly $60B in commercial paper each year. A portfolio company with fewer than 7 employees automated a process previously dependent on people and banks, generating $6B of paper in a week and charging just “0.01 to 0.02”—still a substantial business. The investment thesis is to find workflows that are frequent, important, repetitive, and boring, because “many verticals are actually very large markets.”
In highly regulated industries, the technical goal is not zero hallucinations but an error rate within a controllable range, while making cost, privacy, and deployment work at the same time. 张璐 is particularly interested in combining large language models with reinforcement learning: LLMs handle understanding, analysis, and expression, while reinforcement learning adds decision-making, drive, learning, and memory. Repetitive tasks can also reuse verified answers instead of regenerating them every time. She strongly favors vertical small models because they require less inference, training, and energy and can run locally; one model described as having “under 100M to 1B tokens” can reportedly run on a Raspberry Pi with performance similar to GPT-4.
General-purpose Agents may be technically viable, but the economics and competitive dynamics are far harder for startups. They must call powerful models such as GPT-4, Claude, and Gemini, while broad cross-industry coverage drives up inference costs. If “AI labor costs more than human labor,” enterprises have little reason to adopt them. More fundamentally, foundation-model companies will build general-purpose Agents themselves. 周炜 is still willing to bet on the category because “without dreams, how are you different from a salted fish?”(没有梦想那跟咸鱼有什么区别)—but he admits the odds of success are “extremely slim.”
More durable investment moats include layered business processes, non-public data, and a clear willingness to pay. 周炜 uses the octopus machines from The Matrix as a metaphor: General AI may break through the outer layer quickly, but industry know-how, granular processes, and private data force Big Tech to break through “layer by layer.” These companies may never reach $10B, yet could still produce a cohort of $1B businesses. US enterprises pay once they see commercial value; Chinese customers are more likely to haggle and cap usage-based fees, so Chinese teams often need to target the global market from day one.
The next to C cycle may not be another general chat interface, but products that use local memory and emotional relationships to create real switching costs. 周炜 favors companion robots with local deployment and continuous training: cloud models “each rule for 100 days,” leaving users free to switch at any time. 泓君 imagines that if a $29 monthly subscription lapses and causes a dedicated AI to forget shared experiences, the emotional connection could itself become a reason to pay. 周炜 expects it to become common for the AI-native generation to fall in love with robots, and believes a lifelike product that crosses the uncanny valley could appear “within 2 years.” Meanwhile, this cycle’s high-priced acquisitions show that startups need not bet everything on the endgame; exiting midway is a viable path.
🔗 Original source & video: E200 | An In-Depth Investor Conversation: Core Moats and Investment Logic for AI Agents
100: How Silicon Valley Sees DeepSeek: A Conversation with Zhang Lu of Fusion Fund on Open Source, Agents and “Beyond AI”
- 🗓️ Date:
2025-01-29| 🎙️ Show:晚点聊 LateTalk
DeepSeek R1 suggests open-source models can catch closed systems, while unsupervised reinforcement learning produced chain of thought without process labels. V3 used 2,048 H800s and beat Llama 3.1, trained on 16,000 H100s, on some benchmarks; lower costs may widen AI adoption and GPU demand, though proprietary data still favors closed-source players.
View Dialogue Notes & Key Takeaways
Zhang Lu sees DeepSeek R1’s breakout as a victory for the open-source ecosystem, not as a Chinese company independently catching up with OpenAI. From V2 and V3 through R1, DeepSeek has consistently published technical reports and intermediate experiments, allowing teams worldwide to share architectural exploration. That openness is especially scarce as OpenAI is mocked as “CloseAI” and Mistral and AlphaFold 3 have tightened access to details. For investors, if open source catches up with closed source, the most direct beneficiaries are startups that can call, distill and modify models at low cost.
R1’s most important industry implication is not any single benchmark, but that unsupervised reinforcement learning can produce chain of thought without process labels. Zhang Lu does not see this as overturning the scaling law; rather, it suggests that after reducing the need for expensive labeling, more compute could still lift capabilities by “another order of magnitude.” Lower costs are only the result of this technical path; the real upside is that model training, inference and self-exploration have been reopened as areas for innovation.
DeepSeek’s cost cuts triggered near-term concern about Nvidia, but Zhang Lu’s long-term view is the opposite: cheaper models will pull AI into every industry sooner and ultimately lift GPU demand. The comparison raised by the hosts was that V3 used 2,048 H800s and beat Llama 3.1 on some benchmarks, while the latter was trained on 16,000 H100s. Zhang Lu’s point is that lower per-task costs will bring AI to more use cases that previously could not afford it; this is not the end of compute demand.
Meta will face brand pressure over “why a small company did it better,” but it may also become a long-term beneficiary of the open-source boom. DeepSeek’s disclosed V3 training cost of roughly $5.57M covers GPU hours only, excluding upfront investment, R&D and personnel, so it cannot be directly compared with Meta’s full spending. Counting only the “raw ingredients for cooking” is not the same as counting the kitchen, pots and pans. DeepSeek’s new architectural direction could help Llama iterate, while Meta can continue differentiating itself from closed-source players such as Google through the open-source ecosystem.
The closed-source camp has not stopped, and competitive advantage is moving beyond model scale toward proprietary data, internal tools and commercial data flywheels. Anthropic is acquiring more high-quality B2B data through enterprise orders; xAI may tap 2D/3D industrial data from Tesla factories and supply chains, as well as SpaceX and Starlink. Gemini, Claude and ChatGPT each have advantages in accuracy and writing style. Zhang Lu heard from a founding member of an unnamed company that 70%—80% of its internal code is already written by proprietary AI tools; undisclosed applications are being used first to accelerate the company’s own iteration.
The Agent thesis is right, but today’s products look more like “very junior interns” than analysts. When Zhang Lu tested Operator, its search speed was “like an old lady,” and it still fabricated information. Enterprise customers are therefore more likely to buy narrow, specialized and accurate industry Agents than a generic assistant that costs $200 per month but tries to do everything. Compared with traditional RPA, the breakthrough is lowering the user barrier to “can you chat?”; deployment still requires workflow integration, inference optimization and cost infrastructure.
Outside AI, Tech Bio and space tech are the two areas to watch most closely, though both are being accelerated by AI as a “catalyst.” Longevity has shifted from whether people can live to 150 or 200 toward whether they can reach 100 or 120 with a clear mind and healthy body. In space, there is a possibility that over the next 3—5 years the cost of sending one person into space could fall to $50K—$100K, while launching a single satellite could fall to $10K—$20K. Zhang Lu also warned that in 2025, “uncertainty will become the only certainty”; early-stage companies will derive more of their edge from rapid iteration, industry data, commercial channels and resistance to the giants’ control of resources.
🔗 Original source & video: 100: How Silicon Valley Sees DeepSeek: A Conversation with Zhang Lu of Fusion Fund on Open Source, Agents and “Beyond AI”