Pioneers Insight Method Research Author
Back to Pioneers
曹卿云 Qingyun Cao
Founders 4 Curated Dialogues

曹卿云 Qingyun Cao

AI Pioneer

Frontier Insights

Frontier Thesis: AI scaling has decoupled from standard systems: a 10x compute-to-optical interconnect deficit and an impending memory wall necessitate immediate architectural evolution toward open CPO standards and multi-tier, semantically managed KV caching.

Strategic Imperatives: Capitalize on the optical supercycle beyond copper limits while shifting memory pipelines from static caches to dynamic cross-hardware reuse (CPU/SSD/GDS). On talent, aggressively secure tier-one researchers as automated AI workflows disrupt traditional high-skill roles like quant research.

Key Risks: Ecosystem inertia—namely lagging enterprise adoption of advanced cache layers, integration hurdles in emerging optical standards, and volatile, consensus-lacking model performance.

Key Views & Dialogues

【Special Variety Show—Part 2】 Silicon Valley AI Researchers | The Model Guessing Game and Prompt Battle

  • 🗓️ Date2026-07-19 | 🎙️ Show:硅谷坐标 Silicon Valley Vector

Blind testing found Gemini 3.5 Flash, Claude 4.6 Sonnet, and DeepSeek outputs highly similar, challenging OpenAI researchers to identify their own model. Voice agents have entered real workflows such as Chase credit-card closures, but replacing most daily tasks remains distant and may depend on robotics.

View Dialogue Notes & Key Takeaways
  • A blind test exposed just how homogeneous model outputs have become: even OpenAI researchers repeatedly failed to recognize their own model among the major labs’ offerings. Six guests relied on stylistic tics—parentheses, em dashes, line breaks, and bullet points—to identify the models. Gemini 3.5 Flash, Claude 4.6 Sonnet, and DeepSeek were repeatedly misidentified; Bessie’s unexpected takeaway was that “it’s actually very hard for everyone to reach consensus on which one is their own model… many models are just incredibly similar.” For investors focused on model differentiation, it offered a direct look at how much outputs are converging.

  • The claim that older models are quietly being made dumber before a new model launches received a qualified internal rebuttal. Wang Sen said, “Based on my personal understanding, that should not be the case.” The deeper explanation is that users’ own expectations are unstable: “Even if it gives you an answer, you might not be equally satisfied with it yesterday and today.” Teams do look for clues in social media and user feedback, and after confirming bad behavior, try to eliminate it through algorithmic iteration; users are encouraged to click thumb down.

  • Voice agents are already operating in the real world; AI replacing most of daily life is still a long way off. Zhou Yichao, who recently joined a voice-agent startup, said that when you call Chase to close a credit card, “there’s a very high probability an agent is already taking the call, and it very likely can complete your task quite well.” But AI replacing most things in daily life “is still relatively far off” and “may still depend on the arrival of robotics.”

  • At the frontier, Chinese and US research has more in common than not. “The frontier model companies doing well are basically in China and the US,” and there is broad consensus on how to approach scaling laws and RL. The real differences may lie in resources—“with fewer resources… people work more meticulously”—and in the communications environment: Chinese teams have more Chinese members and more concentrated communication, while US teams are more diverse and require more alignment.

  • The guests’ AI workflows center on sub-agents, documentation, and starting fresh chats frequently. For problems they cannot solve, they first ask the model to lay out a plan, revise it, then “spin up multiple sub-agents to carry out” the work. Long conversations can send a model “suddenly into a dreamlike alternate reality” after compaction, so they keep records in documents. Rather than maxing out the context window in one chat, they start a new conversation—“faster, and sometimes even more accurate.”

  • The panel split openly on whether to sell after an employer goes public. One camp argues that “at current valuations… they can all be held for the long term”; the other says, “Of course you sell… people have waited so long in private companies… you can’t be all in one place,” selling part of the position to buy other companies. FIRE targets ranged from $2M—enough to return to Chengdu or a wife’s hometown, provided one stops comparing oneself with others—to $10M.

  • The Prompt Battle strategy: use verifiable tasks and as many tool calls as possible. The prompt that got ChatGPT 5.5 Thinking High to think as long as possible was: “Search for the 90 most-cited papers for each year since NeurIPS was founded and summarize them in a doc”—it thought for 6 minutes 24 seconds. The competing prompt, “Query at least 100 webpages and produce a research report on how much money it takes to achieve financial freedom,” took just 1 minute 23 seconds. The core strategy is to set a clear, verifiable objective so the model cannot take shortcuts, produce plausible-sounding nonsense, or turn the question back on the user.

  • 🔗 Original source & video: 【Special Variety Show—Part 2】 Silicon Valley AI Researchers | The Model Guessing Game and Prompt Battle

Listen to full conversation →


[Special Variety Show · Part 1] “From Hot to Flop” | A Model Ranking Game with 7 Silicon Valley AI Researchers

  • 🗓️ Date2026-07-18 | 🎙️ Show:硅谷坐标 Silicon Valley Vector

Model quality rankings showed no consensus: GPT-5.5 was not consistently “hot,” while Gemini’s reception ranged from “elite” to “total flop.” Compensation showed a clearer signal: Meta TBD was deemed “universally hot,” AI researchers’ incomes are rising, and auto research could expose quant research to more efficient replacement, with GPT-5.5 Pro demonstrating a capability leap.

View Dialogue Notes & Key Takeaways
  • There was no consensus on model quality: the ranking judge put GPT-5.5 at the top, but when the cards were revealed, its label was “top-tier,” not “hot,” so the round ended in “perfect failure.” The answers included Gemini, GPT-5.2, GPT-5.5 and GPT; one guest ranked Gemini “elite” (“The Big Three is always the third one”), while 周奕超 went straight to “total flop”: “What model in the US is worse than it?” He added: “The biggest flop is the one nobody remembers.”

  • The consensus “hot” pick on pay was “Meta TBD,” which the host called “the universally agreed answer”; the discussion centered on a widening compensation split: AI researchers make far more than before, while pay growth in non-AI roles may be constrained. The panel also noted that talent is already moving from quant into AI; if models can do auto research, “they can do quant research too, and may do it much better than humans.”

  • The job-security round was a clean sweep: Apple was ranked “hot,” while XAI was a “total flop.” The host used Apple’s status as the only company that had never laid people off as the deciding factor, though a guest then pointed out that Apple had in fact done layoffs. XAI was dismissed for “cutting too many people,” alongside the rule of thumb: “When in doubt, ban it first.” OpenAI only made “elite,” on the grounds that “there doesn’t seem to be any large-scale turmoil, but everyone just quietly disappears.”

  • On work-life balance, 王森 said he had originally planned to put today’s OpenAI and Anthropic in the “this generation’s total flop” bucket, but switched to Netflix to avoid a collision. The closing reflection was more telling: automation has spent the past 200 years trying to trade more work for more free time, yet “we are doing more work, but we haven’t gained more time for ourselves.”

  • The AI-era skills ranking was the panel’s most internally coherent: empathy was “hot,” physical health “top-tier,” cooking “elite,” Photoshop “NPC,” and coding “a total flop.” 王森’s test was that coding is “highly verifiable”—write good tests and you can check the result—while Photoshop also involves harder-to-verify aesthetic judgment. He said that since January or February this year, he has not opened a code editor and has used agents to write all his code.

  • The strongest capability signal surfaced in the banter: a guest’s card-playing friend sent GPT-5.5 Pro an open conjecture in probability theory he had considered during his PhD nearly 30 years ago, and the model “then proved it.” Friends cross-checked the result and concluded it was correct; a second version is now being posted. The panel called it “a pretty hot weekend activity.”

  • The AI-bubble verdict is a reflexive argument. “What is a bubble? It is only a bubble when nobody thinks it is one. You now encounter someone talking about an AI bubble every day—so how could it be a bubble?” The corresponding insider view is that frontier-lab employees see the progress of the next generation of models before outsiders do, and may glimpse changes 10 days or 1 month ahead—a kind of privileged experience in this era.

  • 🔗 Original source & video: [Special Variety Show · Part 1] “From Hot to Flop” | A Model Ranking Game with 7 Silicon Valley AI Researchers

Listen to full conversation →


Silicon Valley Coordinates x TensorMesh 江鋆晨: AI’s Memory—A Three-Layer Understanding of KV Cache

  • 🗓️ Date2026-06-30 | 🎙️ Show:硅谷坐标 Silicon Valley Vector

KV Cache is moving beyond reusable black-box state toward semantic manipulation of attention, while most industry remains at layer one. TensorMesh combines CPU, SSD, or GDS reuse with LMCache and CacheBlend to lower costs and improve speed; a “KV Cache moment” may arrive before the end of this year, but large-company adoption remains uncertain.

View Dialogue Notes & Key Takeaways
  • 江鋆晨 lays out the episode’s core framework: KV Cache has three layers of understanding, and most of the industry is still stuck at the first. The first treats it as reusable “black-box data” that can be stored; the second treats it as a white box containing semantic (attention) information, enabling lossy compression, non-prefix reuse, and cross-model reuse; the third directly changes the semantics—“pay more attention to this, pay less attention to that”—and can even improve output accuracy. “Very few people in industry understand this layer,” which is precisely why TensorMesh was founded.

  • He thinks the “KV Cache moment” may arrive sooner than expected—“possibly before the end of this year,” with a wave of companies entering the space. He is not intimidated by Big Tech: the advantage is cross-disciplinary. Infra engineers who can work with GPU data cannot handle layers 2 and 3, while ML researchers who understand semantics lack the engineering capability required for layer 1. He also cites Sam Altman’s view that large companies have “people’s attention spread across too many places,” meaning a team might work on something for 2 months before moving on.

  • He dismantles the pricing illusion that prefill is worthless because input tokens are cheap: the true compute cost per input token is roughly the same as for an output token, while pricing does not reflect cost. Prefill is highly parallel and, in his example, accounts for roughly one-tenth of the user’s wait time. But in real use cases, inputs running to tens of thousands of tokens are often more than 10x longer than outputs of a few thousand tokens. Because agent applications are stateless, input will only keep getting longer. That is the foundation of TensorMesh’s cost-reduction logic.

  • The economic logic is “trade storage for compute,” but TensorMesh is building software on top of hardware, not buying hardware. When KV Cache sits on CPU, local SSD, or GDS, retrieving it is both cheaper than recomputing it and faster; only colder storage layers create a cost-versus-latency trade-off. He also stresses that the company is decoupled from the storage-hardware cycle: price spikes reflect a supply-demand imbalance, with orders booked 2 years out, and should ease over the long term. The calculation customers should make is how much they spend on hardware versus software.

  • He rejects the TAM question outright: “total addressable market is a misleading word,” because most people do not yet realize this market exists. With the model and prompt fixed, helping the model understand the prompt better is a “third path” to improving agent quality. The best use case is enterprise shared knowledge—codebases, policy and legal documents—reused across coding, chatbot, and RAG applications.

  • Differentiation rests on ecosystem and technology: LMCache has the best ecosystem support among comparable projects, while CacheBlend solves non-prefix reuse—KV Cache storage “is not merely Storage; it is a Service” with compute inside. Asked about similar projects mentioned by investors, he simply says, “They really are similar, good luck.” His response to architecture risk is that KV Cache is only a temporary name; fundamentally it is model-native data, and it will remain relevant in the Mamba era—Qwen 3.5, for example, still retains roughly one-quarter Transformer layers.

  • Two analogies are worth remembering: KV Cache is “big data for the AI era,” and TensorMesh wants to build the Databricks for “data only models can read,” with Ion Stoica as an adviser; in distribution terms, it is the CDN of its era—“Akamai was OpenAI back then.” The company is currently building a mini-CDN inside the data center. Once edge computing and distributed inference take off, it could become an internet-scale knowledge delivery network: “We may have been looking at 6 years into the future 3 years ago.”

  • He bets the foundation-model endgame looks more like online video than search: a handful of companies consolidate, but sovereign AI and enterprise private deployments will support many independent services. His open-source view is blunt: OpenAI and Google’s open-source models “will find it difficult to compete with these Chinese open-source models,” and the open-source market may still be China’s to watch over the next few years. xAI’s acquisition of Cursor is one example of the consolidation wave, though he says he does not know much about that specific case.

  • 🔗 Original source & video: Silicon Valley Coordinates x TensorMesh 江鋆晨: AI’s Memory—A Three-Layer Understanding of KV Cache

Listen to full conversation →


Silicon Valley Coordinates x Innolight’s 于让尘: The AI Optical Interconnect Supercycle

  • 🗓️ Date2026-04-24 | 🎙️ Show:硅谷坐标 Silicon Valley Vector

AI optical interconnects face a 10x gap: compute has grown roughly 300x since 2022, versus 30x for optical bandwidth, while copper is nearing its limits. TeraHop is advancing open 12.8T XPO and Open CPO-X standards, but 400G technologies, OCS adoption and semiconductor integration remain key variables.

View Dialogue Notes & Key Takeaways
  • AI optical interconnects are in a supercycle driven by a 10x gap: compute has grown roughly 300x since 2022, while total optical-connection bandwidth has grown only about 30x. TeraHop Vice President 于让尘 compares compute with gray matter and connectivity with white matter: the two occupy nearly a 1:1 ratio in the human brain, while the white-matter-to-gray-matter ratio in cats and dogs is roughly 10%-20%; “we may still be at the mouse-and-kitten stage.” Optical interconnects currently account for only a single-digit percentage of data-center capex, but he argues that connectivity investment must grow faster than compute: “this gap has to be filled, and the entire industry has to work on it together.”

  • Scale up is the industry’s biggest structural opportunity: bandwidth is roughly 10x scale out, while copper is nearing its physical limits. Nvidia’s all-copper NVL72 rack “is already constraining the brain’s development”; the industry is moving from the 576-chip systems now coming online, and 1,000-chip systems, toward Google’s interconnection of 9,000 TPUs. Copper can currently run only 1.5-2 meters; in the 400G/lane era, “even 1 meter will be a stretch,” pointing to a long-term shift toward optical replacing copper.

  • The economics need to be calculated backwards: compare output per token, not the cost of optics versus copper. 于让尘 cites the industry rule to “use copper when you can, use optics when you must,” but stresses that end customers care about returns per unit of compute and per token. GPUs cost $30K-$50K, versus several thousand dollars for an optical module; “if you save on interconnect, you are sacrificing your compute.”

  • The technology roadmap is not a religious either-or: “Internet companies are practical; they do not make religious choices.” A single company may choose pluggable optics, NPO and CPO at the same time—“adults don’t do multiple-choice questions” (“成人不做选择题”). Google has said publicly that “CPO is always two years away, and every year it is still two years away.” The major cloud companies are all developing proprietary chips—TPU, MTIA, Trainium, Maia, OpenAI Titan and xAI—which makes them receptive to an open optical-interconnect ecosystem. CPO’s failure to reach scaled commercial deployment is partly because customers dislike vertical integration and suppliers gain too much bargaining power, while the technical bottlenecks remain formidable.

  • TeraHop is positioning itself around open standards: its 12.8T XPO and Open CPO-X multi-source protocols can deliver 4-8x bandwidth growth, with about 100 companies already participating. The case for the technology rests on silicon-photonics maturity: more than 1,000 eight-channel products are already in use, making 16, 32 and 64 channels a natural progression. Nvidia’s $4B investment in Lumentum and Coherent in early March was interpreted by 于让尘 as concern that the supply chain could not keep up with 10x growth—“locking down the supply chain with real money.”

  • OCS is a greenfield market Google began developing roughly 10 years ago, while other major players remain at zero or very low adoption—but electrical switching is not going away. Optical switching cuts power, latency and cost, but it switches slowly and is suited only to predictable workflows. 于让尘 is explicit in his caution: it “cannot completely replace” nanosecond-scale electrical switching over the next 5 years; “never say never,” and the two technologies will coexist over the long term. OCS’s 2-3dB loss makes coherent light a good pairing, and Google’s strategic investment arm has backed the coherent-light startup Celero.

  • The industry’s biggest bottleneck is that semiconductor integration remains early-stage, while the value chain will soon become a trillion-dollar question. Lasers are still made on 3-inch wafers and are working toward 6-inch, versus 12-inch wafers for silicon photonics; quantum-dot lasers could eventually put the light source directly on silicon. Over the next 5 and 10 years, the direction is semiconductor-based optical interconnects and deep optoelectronic integration, including TSMC’s COUPE platform. Interconnects should be treated as part of compute itself and ultimately become an equally important component.

  • 🔗 Original source & video: Silicon Valley Coordinates x Innolight’s 于让尘: The AI Optical Interconnect Supercycle

Listen to full conversation →