Cerebras Early Investor 周楠 on the Nvidia Challenge and Baidu's US Lab
Summary
- Cerebras is not a full-stack substitute for Nvidia; it is an Option B that exploits a crack in the inference bottleneck. It turns an entire wafer into a computing engine, seamlessly interconnecting a vast number of compute cores and putting compute, memory and communications on the same piece of silicon to reduce data movement. 周楠 draws the boundary clearly: Cerebras is “a challenger on specific workloads,” while CUDA, networking, software, supply chain and customer trust remain system-level moats Nvidia cannot easily replicate.
- Cerebras went from a valuation of several billion dollars to nearly $100B at one point after its listing; the key variable was the shift in compute demand from training to inference. 周楠 says more than half of GPU demand now comes from inference, while coding and agents are driving token consumption higher. Low latency, high throughput and lower cost per token are therefore becoming direct determinants of product experience and gross margin. OpenAI’s at-least-$20B order signed in January 2026 reflects both a performance decision and a strategic effort by a frontier lab to reduce single-supplier risk.
- Above its roughly $50B market cap, 周楠’s bull-case ceiling for Cerebras is $500B, but what needs validating is delivery, not the concept. The wafer architecture works, but customers remain concentrated among a handful of large accounts, including G42; the OpenAI order has not yet been delivered, and the company still needs to prove that customized infrastructure can be operated reliably and scaled quickly. Nvidia will not simply watch share leave: its $20B acquisition of Groq is a direct counterattack on the inference market.
- Baidu co-led Cerebras’s Series C at a valuation of about $700M in 2017, when the company did not even have a physical chip. Baidu’s U.S. research team ran a nearly 300M-parameter language model based on PaddlePaddle in a simulator and examined yield, liquid cooling, power, compiler and API risks one by one. 周楠 recalls that “almost every downside case in the risk model happened”; tape-out and system deployment were ultimately delayed by roughly 1-2 years. The patience of early investors and the board helped the company survive a hardware cycle that lasted nearly a decade.
- The investment originated in Baidu’s U.S. research team seeing the empirical seeds of Scaling Law in Deep Speech 2. The team had not yet formalized it mathematically, but it had a clear working view: larger models, more data, longer training and more compute consistently improved results. Training a nearly 300M-parameter model on GPUs took more than 3 months, and each iteration required several more months. Cerebras’s claim of 100x or 1,000x efficiency gains therefore addressed a real and urgent research bottleneck.
- Baidu’s U.S. research team was both a forgotten AI talent hub and a capital opportunity cut short by geopolitics. At its peak, the roughly 250-person group produced founders, co-founders or founding members of OpenAI, Anthropic, Inflection, Adept and Meta FAIR. Baidu’s planned independent Growth Fund had OpenAI, Databricks and Scale AI on its explicit investment list, but never launched after LPs pulled back. 周楠’s lasting regret is that, had the fund closed, Baidu could have become an investor in a cohort of frontier labs and AI infrastructure companies.
- The non-consensus window in early AI investing has compressed from years to 1-2 months, turning VC into a game of using large funds to hammer the winners. 周楠 sees flywheels already forming around coding agents, frontier labs and parts of infrastructure. The remaining cracks may be in Physical AI, inference optimization and new CPU architectures for agent scheduling. He believes robotics has more fragmented use cases, but unlike autonomous driving, it does not need to reach “99.9%” accuracy everywhere; its “aha moment may arrive earlier than we think.”
Deep dive
1. Cerebras Sells a Computing System, Not a Replacement GPU
周楠’s core definition is wafer-scale architecture: conventional wafers are cut into many chips, whereas Cerebras minimizes the cuts and seamlessly interconnects the vast number of compute cores across the full wafer to create a giant AI computing engine.
The company covers chips, servers, liquid cooling, power, the compiler and the software stack. The host accordingly revises her own description: rather than calling Cerebras a chip company, it is better understood as searching for “a new way to deliver AI compute.”
The architectural goal is not simply to add more arithmetic units, but to bring compute, memory and communications closer together, reducing data movement across chips and to external HBM. 周楠 compares it to “a very large brain.”
2. It Challenges Inference Workloads, Not Nvidia’s Entire Ecosystem
Ten years ago, Baidu researchers were not trying to “beat Nvidia”; they wanted to avoid a future with only one supplier. Their view was that roughly 96% of a GPU’s die area was not optimized for deep learning, leaving room for a more efficient dedicated architecture.
Nvidia is now what 周楠 calls a “$5T company,” and its moat extends well beyond the chip: CUDA, the developer ecosystem, networking, systems software, customer trust and the supply chain together form a full-stack advantage.
That is why 周楠 describes Cerebras only as “a challenger on specific workloads.” When a task is constrained by memory bandwidth, communications latency and response speed—particularly in inference—it can be an excellent choice. Full replacement of Nvidia is not the conclusion today.
3. The Shift from Training to Inference Changed Cerebras’s Valuation Equation
Two or three years ago, when ChatGPT had just emerged, scarce compute was mainly being used to improve model capabilities and training remained the center of the compute narrative. Cerebras was then valued at only several billion dollars, while the market’s first bet was still Nvidia.
周楠 says more than half of GPU demand now comes from inference, while coding and AI agents have sent token consumption soaring. Inference has moved from an ancillary step after model completion to the core cost that determines user experience and commercial gross margin.
That is the point at which Cerebras’s low latency, high throughput and inference speed became genuinely scarce. 周楠 describes the valuation surge as “a matter of course,” rather than a re-rating driven only by market sentiment.
4. OpenAI’s At-Least-$20B Order Bets on Both Performance and Supply Autonomy
Sam Altman invested in Cerebras personally in 2016, before Baidu’s 2017 investment. The timing came soon after OpenAI was founded in late 2015, leading 周楠 to conclude that Altman understood early on that model companies could not rely on just one type of chip.
In January 2026, OpenAI signed a contract worth at least $20B with Cerebras. 周楠 sees two drivers: compute has become the primary bottleneck to further model scaling, and every frontier lab needs diversified supply and strategic autonomy. He also expects companies such as Anthropic to seek alternative or customized compute solutions.
When the host asked whether the deal constituted a related-party transaction, 周楠 stressed that the early shares were held personally by Sam Altman and that Altman was not a particularly large shareholder, so it “doesn’t count as a related-party transaction.” The host nonetheless retained her doubts about his multi-track capital dealings.
5. Cerebras Cloud Removes Adoption Friction; It Is Not Just a Move into Cloud
In the early years, the customers able to buy new training systems were mainly large platforms such as Google, Microsoft, Amazon and Baidu. As application companies and use cases multiplied, deploying the hardware itself became a constraint on growth.
周楠 calls this adoption friction: after buying entirely new hardware, customers still have to build data centers and modify their software stacks, making the process lengthy. Cerebras Cloud wraps the underlying complexity and lets customers access training, applications or API services directly through an API.
The host’s “dark suspicion” was that, once the system is wrapped in a cloud layer, it may no longer matter whose compute sits underneath. She also cited Andrew Feldman’s blunt statement: “We will work with everyone except Nvidia.” 周楠 sees the cloud as a one-stop solution for a compute-constrained market.
6. $500B Is the Imagination Case; Customer Concentration and Delivery Speed Set the Floor
Cerebras approached $100B after its listing and is currently worth roughly $50B. 周楠 says he “wouldn’t be surprised” if the company eventually reached $500B as inference demand continues to expand.
The floor is not whether the wafer works—that milestone has already been cleared—but whether customers can expand rapidly beyond a handful of large accounts such as G42, and whether the systems can be delivered and operated reliably at scale over time.
The host noted that the OpenAI order has not yet been delivered, so the company cannot declare victory in inference ahead of time. 周楠 believes the current team is capable of completing the customized infrastructure, but acknowledges that the next question is “how quickly they can deliver this solution.”
7. Wafer Scale Trades Manufacturing and Systems Complexity for Communications Efficiency
Conventional GPU clusters connect multiple independent chips through PCIe, NVLink and similar links. Cerebras keeps compute, memory and networking on the same piece of silicon, reducing cross-chip coordination and external-memory transfers and lowering the cost of communications and data movement.
The risks therefore extend beyond a single chip: wafer yield, packaging, thermal management, power, localized shorts, compiler mapping and whole-system stability all have to work simultaneously. A delay at any one stage can hold back commercialization.
The episode’s postscript frames the inference era as a window for heterogeneous chips: Cerebras is betting on wafer scale, while Nvidia has responded with its $20B acquisition of Groq and subsequently incorporated Groq into its own compute platform in its GTC offering.
8. A Failed Search-Company IPO Led 周楠 into AI Investing
Before joining Baidu, 周楠 worked in investment banking at Barclays Capital and moved to Hong Kong, where he worked on listings for mobile-internet companies including Alibaba and JD.com. He also worked on an IPO for a chip company.
The real turning point was another search company that failed to list. Its prospectus told an AI story that made 周楠 feel for the first time that “this is the future.” That failure instead paved the way for his subsequent career move.
After hearing 吴恩达 speak at a Silicon Valley lecture, 周楠 became intensely interested in Baidu. When Baidu began recruiting AI investors globally at the end of 2015, he applied and moved from the sell side to early-stage investing at Baidu’s U.S. research institute in 2016.
9. Baidu’s U.S. Research Institute Had Elite Talent, GPU Budget and Research Freedom
When 周楠 joined, Baidu’s U.S. research institute had about 250 people at its peak and was already based in Sunnyvale. He describes the atmosphere as “exciting every day,” with research spanning speech, vision, retail, fintech, healthcare and autonomous driving.
吴恩达 had both strong talent magnetism and a substantial GPU budget. 周楠 believes his early public insistence that GPUs were suited to training AI models played an important role in building Nvidia’s understanding of deep learning.
Researchers who later left became founders, co-founders or early members of Inflection, Adept, OpenAI and Anthropic, while others helped create Meta FAIR. Former colleagues remember it as “an era when a whole group of immortals were battling it out,” with a talent density rarely seen even at other top laboratories.
10. Deep Speech 2 Had Already Seen Scaling Law, Before It Had a Formula
The host identified Dario Amodei as the first author of Deep Speech 2. 周楠 confirmed the paper’s importance to Baidu and said Greg Diamos was also one of its authors. He emphasized that the Baidu team had not yet formalized Scaling Law mathematically, but had observed a stable empirical relationship.
The relationship was straightforward: larger models, more data, longer training and stronger compute systems continued to improve model performance. What is now theorized as Scaling Law was, in 周楠’s view, already emerging around this paper.
This was not a retrospective label pasted onto old research; it directly shaped the investment mandate at the time. If AI progress required models, data and compute to scale together, investors had to find a computing system better suited to deep learning than existing GPUs.
11. Training a Model Once Every Three Months Made 1,000x Efficiency a Real Need
Baidu’s language model at the time was approaching 300M parameters, an enormous scale a decade ago. Researchers told 周楠 that fully training it on GPUs took at least 3 months, while tuning and iterating required several more months.
Cerebras’s pitch at the time was to make deep-learning training 1,000x more efficient, compressing months into days or weeks. Even though the result was still based on simulation, it targeted the researchers’ most urgent pain point with unusual precision.
周楠 also studied Graphcore, Wave Computing and several ASIC companies. He viewed ASICs primarily as inference solutions, while training was the immediate priority. Wave was eliminated first because of team issues; Graphcore improved efficiency but was less radical than the Cerebras approach.
12. Andrew Feldman Did Not Design Chips, but He Assembled a Rare Chip Team
Andrew Feldman previously founded SeaMicro, which was acquired by AMD, and then built Cerebras with nearly the same core team. A little more than a year after its founding, the company had about 80 employees, nearly 70 of them PhDs, with a combined several hundred years of industry experience.
The host raised the obvious concern: Feldman himself was not a chip engineer. 周楠’s response was that a systems company needs product definition, customer understanding and organizational ability, while Feldman already had several technically strong co-founders with whom he had worked for years.
What ultimately convinced 周楠 was Feldman’s conviction. He did not just present a grand vision; he answered questions about yield, thermal management, power, the compiler and customer adoption one by one, spending roughly 2 hours a day for 4 consecutive weeks walking a “chip novice” through the risks.
13. Baidu Used One of the World’s Largest Language Models to Validate a Machine That Did Not Yet Exist
Baidu researchers were willing to participate in due diligence because slow training was already a firsthand problem. Greg Diamos was particularly important: he had been one of the core developers of Nvidia’s CUDA ecosystem and was also an author of Deep Speech 2, giving him a working understanding of both GPUs and models.
Cerebras had not yet taped out a chip and had only a simulator. 周楠 says Baidu was the only company then able to validate the system using what was at the time the world’s largest language model. The team ran a pre-Transformer model based on PaddlePaddle on the simulator.
The result could answer only what performance would look like if yield, the compiler, thermal management and packaging all worked; it could not replace testing on physical hardware. 周楠 did not disclose specific data, saying only that the results were “really good,” enough to validate the architectural concept.
The collaboration served more than the investment decision. Some Baidu researchers maintained close cooperation with Cerebras, and others later joined the company, showing that the diligence process also became an early product and talent bridge.
14. Quantifying the Worst Case Turned Disruptive Risk into Investable Risk
周楠 brought in a Stanford professor, chip experts and Baidu’s autonomous-driving hardware team to review manufacturing risks. The team estimated that a failed tape-out could consume another 6 months and add $5M-$10M in costs, then checked whether the company had enough cash to absorb it.
On cooling, Cerebras had already demonstrated a liquid-cooling system far larger than the wafer itself. Localized chip shorts could be identified in software, with workloads rerouted to redundant units.
Software diligence focused on how models would map onto the architecture and how the system would connect with TensorFlow, Caffe, PyTorch and customers’ data centers. 周楠 believes that without Baidu’s researchers, the compiler and API risks “could not have been predicted.”
15. After 4 Weeks of Diligence, Baidu’s Investment Committee Approved the Non-Consensus Bet in Under 2 Days
The lead came from Coatue founding partner Thomas Laffont. He had studied semiconductors and, after hearing Baidu’s team discuss its compute thesis in a meeting, introduced them to Cerebras, an early investment of his.
周楠 completed a bilingual investment memo of more than 10 pages in roughly 4 weeks. He recalls that the investment committee included Jennifer Li, 陆奇 and 李彦宏; the deal was approved in less than 2 days after the memo was circulated—“an instant, painless approval.”
吴恩达 left Baidu in April 2017 and therefore did not participate in the decision. The investment closed around August or September 2017, with Baidu co-leading the Series C. Management did not ask to wait until after tape-out.
The valuation was about $700M, already close to the unicorn threshold. 周楠 forecast the 2025 AI training-chip market at only $22B and expected Cerebras to take roughly 20%, implying a 3-5x return. He now says the more important question is “how wildly wrong I was,” because the overall market turned out to be far larger than forecast.
16. The 2017-2019 Trough Proved the Risk Register Was Not a Paper Exercise
Tape-out, wafer and compiler issues were delayed one after another. 周楠 recalls that “almost every downside case in the risk model happened.” The company eventually cleared those hurdles, but the process ran roughly 1-2 years later than the ideal schedule.
At the IPO celebration, early investors repeatedly brought up that trough. 周楠 believes the continued support of early investors and the 3 board members for Feldman was critical to helping the company survive a hardware cycle that lasted nearly a decade.
Follow-on financing was not effortless either. The company’s valuation moved through stages of more than $2B, $4B and $8B, while leading mainstream VCs at one point pulled back because of market uncertainty and the risks of wafer scale. G42’s dual role as investor and customer also created single-customer concentration risk.
17. A 1-2-Year Tape-Out Delay Locked the Training Market into Existing Clusters
The host proposed a counterfactual: had Cerebras taped out earlier, the company shipping chips to OpenAI might have been Andrew Feldman rather than 黄仁勋. 周楠 acknowledged that “perhaps history would have been rewritten more substantially.”
They immediately pulled back from the romantic scenario together. Transformer emerged in 2017, while a new chip would still have to pass through manufacturing, mass production, compiler and delivery validation; no frontier lab would put core training workloads on an immature system.
Once GPU clusters and data-center investments are in place, switching training chips becomes a systemic risk. 周楠 therefore believes Cerebras should focus less on continuing to fight for the training market and instead capture inference workloads, giving customers a reason to adopt a new architecture for new demand.
18. Baidu Wanted to Spin Out Its AI Investments, but Geopolitics Cut the Plan Short
Baidu’s strategy at the time went beyond Cerebras, covering autonomous driving, lidar, L2/L3 chips and related hardware and software solutions. Teams that emerged from the Baidu ecosystem, including Horizon Robotics, were also within its observation and investment scope.
周楠 helped prepare an independent Growth Fund intended to build a portfolio around frontier AI full-stack systems, data warehouses and data engines. Databricks, Scale AI and OpenAI were all on the explicit deal list, and OpenAI was willing to accept Baidu investment around 2017-2018.
His position was unequivocal: if the fund had closed, “we definitely would have invested.” But LPs who had initially shown interest withdrew one after another for geopolitical reasons, and the fund never launched. 周楠 calls this one of his enduring regrets.
He later took the same thesis to Insight Partners, Sequoia and Benchmark, but was told to keep investing in SaaS and avoid high-risk AI. After failing to find consensus, he moved to Qualcomm Ventures as Baidu’s U.S. investment team contracted.
19. When Anthropic Was Taking Shape, Financing Compute Was Harder to Sell Than the Model Idea
周楠 recalls receiving a call in the summer of 2020 from several researchers who had worked with him at Baidu and later joined OpenAI. They were excited that GPT-3 was nearly trained: “The thing we wanted to do at Baidu is about to be done.”
He asked whether the model had reached the desired level of generalization. They answered honestly, “Not yet—it still hallucinates.” They also expressed distrust of Sam Altman and the safety approach, and discussed leaving OpenAI to start a new company.
周楠 gave the same advice he had given in the Baidu years: they needed to find someone willing to commit enormous sums to compute, and mainstream VCs at the time were not going to write that check easily. With the pandemic at its worst and no vaccine yet available, he did not meet them in person. Jaan Tallinn later engaged with and supported the team repeatedly as an individual, and the story became Anthropic.
20. Baidu’s U.S. Research Institute Was a “Military Academy,” but Could Not Hold on to the Future It Saw Early
On the joke that “Dario was treated badly at Baidu,” 周楠 says it was only a joke, then relays a colleague’s account: Greg Diamos identified and recruited Dario Amodei. Dario came from a mathematics, physics and biology background rather than traditional computer science or AI training; there were internal doubts, but Greg saw his ability in model training and his ideas about AI. 周楠 stresses that this is only the internal version he heard and that it is unclear whether Dario himself would agree with it.
Researchers were free to pursue their own ideas, run experiments and publish papers. The atmosphere was “extremely free.” 周楠 considers Baidu a very important part of Dario’s career, not a negative episode.
Asked whether Baidu “got up early and arrived late,” 周楠 first attributed the outcome to geopolitics: as a Chinese company continuing to conduct frontier model research in the United States, Baidu faced an increasingly difficult environment. The host added that even among Chinese models, Baidu was not the leader; 周楠 acknowledged that he had spent much of his career in the United States and was not in a position to make a fine-grained assessment of China’s domestic landscape.
His broader formulation is that the United States is better at disruptive architectures and original conceptual innovation, while China is better at engineering catch-up and application rollout. The host noted that Baidu’s remaining achievements included autonomous driving and Kunlunxin; 周楠 agreed, but said the lab’s talent dividend spread primarily into Silicon Valley.
21. 周楠 Went from Edge AI Back to the Cloud and Put Inference Infrastructure at the Center
At Baidu, his core themes were compute, data warehouses, data engines and autonomous driving. At Qualcomm, he spent roughly 6 years focused on edge AI, edge chips, IoT and robotics.
One lesson was that domestic hardware and IoT companies in the United States have difficulty scaling quickly, while smart-home and early robotics applications are more abundant in China. Comparable investments therefore have an easier time gaining traction in China.
The 2022 GPT moment arrived several years earlier than Baidu’s own projections. 周楠 began shifting his attention back to cloud-based infrastructure, concluded internally at Qualcomm in 2023-2024 that enterprises would adopt AI faster, and identified coding AI as a major use case before Cursor’s breakout.
Open gaps in the large-model era include RAG, model deployment, inference optimization and multimodal optimization. 周楠 cites a company called AIG, acquired by Nebius this March and founded by a student of 韩松. He encountered the company soon after its founding but missed the investment because it was sold so quickly—an outcome that itself highlights the strategic value of combining inference optimization with the cloud.
22. The Non-Consensus Window Is Closing, but Physical AI and CPUs May Still Offer Openings
In a private investment letter dated February 2025, 周楠 listed general-purpose AI agents including Manus and Genspark, as well as video models, multimodal models, multimodal infrastructure, Figure AI, Physical AI and vertical AI. The time from discovering these opportunities to seeing consensus form is often only “1-2 months.”
In response, leading funds are shifting toward large follow-on checks into category winners. 周楠 cites Benchmark’s announcement of a $2B fund that was fully raised within 24 hours. Anthropic may already be approaching the $1T level, and reaching $500B in the future would not be surprising; this year’s theme is “using big money to hammer the winners.” He also cites early flywheels forming at infrastructure companies such as Fireworks, fal.ai and Baseten.
Asked whether this is still venture capital, 周楠 calls it “a variant of venture capital.” AI remains in the early stages of adoption, and winners may continue expanding through flywheels, but concentrated bets create much greater single-name risk when the judgment is wrong.
He remains willing to study Physical AI. Robots face more diverse tasks than autonomous vehicles, but not every task needs to reach autonomous driving’s “99.9%” accuracy threshold; as models scale, generalization gains could suddenly steepen. The resulting opportunities include low-latency, power-efficient edge chips and new CPU architectures that handle agent-scheduling tasks and serve as inference chips.