168X War Room: A Conversation with Herman Jin—AI Shorts Can Shut Up Now!
168X War Room: A Conversation with Herman Jin—AI Shorts Can Shut Up Now!
Summary
- In June, Herman Jin was “probably among the first to turn bearish”; over the past few weeks, he has shifted relatively bullish: after Anthropic dropped the major bearish catalyst of $65B ARR, there was no deleveraging, mass forced liquidation, or rise in volatility, and CDX did not move sharply higher (CDX IG rose only modestly), suggesting leverage and funding-market contagion did not spread further. His biggest disagreement with the bears is where we are in the cycle: “The bears think this is 2000; we think the worst is 1998.” After this selloff, the market could still deliver a “what the hell happened to 1999 and 2000” rally, though valuation risk and a potential turn in next year’s cash flow remain sources of volatility.
- The key variable behind his bullish turn is open-source models: CSPs can run post-training or further distillation on open-weight models, escape the squeeze of being trapped between two model vendors at the front end and storage suppliers at the back end, improve cash flow, and potentially justify CapEx through Enterprise AI revenue. Anthropic and OpenAI are the wrong anchors—they accounted for only 27% of compute deployed over the past two years; the rest could also be monetized through CSP Enterprise AI. Herman said Microsoft could somehow add $50B, while Enterprise AI could add more than $20B; he also said Microsoft has $20B of revenue this year, equivalent to roughly $40B on an ARR basis.
- “Can the people selling cheap compute make money?” is the more certain trade in his view: DeepSeek still has an 80% gross margin, Alibaba estimates a 2.2-year payback, Microsoft’s Enterprise AI teams underwrite ROI on a 6-to-10-month payback, and V100s are still in use. He cited Jensen Huang’s view that the apparent ceilings in token usage and AI coding are supply-driven, not demand-driven.
- Storage is the area where he would allocate more cautiously: Herman believes storage vendors are expanding gross margin mainly through price increases, not capacity additions or product iteration; gross margin is already around 85%, even 87%, leaving less room to move higher. Storage and Anthropic could both become pressure points for the industry, with everyone incentivized to cut prices. Technically, KV cache scales with the size of the context being read; splitting work across multiple agents shortens each context, so KV cache could fall in an approximately square-root relationship while GPU calls increase, allowing some storage demand to be substituted by compute. He personally prefers NAND over DRAM and does not think portfolios should be concentrated in storage.
- Intel foundry is “certain to succeed” in Herman’s framework: the core of a foundry is the data package tuned jointly with customers, while Apple needs a non-AI-oriented backup and Google needs alternative capacity because it is not a traditional TSMC customer. He believes talent linked to 刘德音 and the current factions, along with 利普·布坦 and 陈立武, will help Intel close the gap; but Intel’s short-term sentiment has already run too far, and fundamental safety and valuation safety are separate questions.
- His recurring framework for AI users is that “the AI era has only Protoss and Zerg”: he says 0.1% of people consume 99% of the tokens. If someone cannot use up even a $200 Claude Max subscription in a week under a normal workload, that may indicate the job has little room for extension and is relatively easy to replace. He said he wrote 450,000 lines of trading-grade code in 3 months, equivalent to what a 20-person professional trading team would write in 1 to 1.5 years; the system is close to zero bug, but not absolutely bug-free.
- At the macro level, he believes a bubble is inevitable—it has simply flowed to the wrong places for now: nominal GDP is growing roughly 6% to 6.5%, about 190 basis points above the roughly 4.7% long-term interest rate, while Starbucks trades at roughly 36x forward PE versus about 12x for NVIDIA. The bubble may truly rush into semiconductors when CSP cash flow turns positive and the market reprices semiconductors from cyclical stocks as a non-cyclical industry. The terminal risk is “AI cancer” absorbing debt, equity, and other resources until unemployment crushes carbon-based consumption; but he judges the market to be in phase 1 or 2, not phase 4—“it’s not today.”
- He is more cautious on crypto: in his view, Binance controls the trading price of every coin, while on-chain markets have no pricing power; stock tokens on Binance are already siphoning away liquidity, and market-maker profits have fallen sharply. He said his per-trade high-frequency taker profit fell from roughly $100 to $25; if market makers eventually stop quoting, the market could “fall apart completely.” He kept buying as BTC fell below $65,000 and sees $200,000, $500,000, or $1,000,000 over the long term, but the near-term question is still where the funding comes from. He called Hyperliquid a “Pi Xiu temple” and dislikes its slow, centralized, post-trade on-chain mechanism.
Deep dive
1. Let the Bears Make Their Full Case: AI Is Systemically Short of Money
- Herman Jin said he was “probably among the first to turn bearish” in June. His initial thesis was that AI was “systemically short of money”: CSPs, hyperscalers, and others would need financing, while CDS and debt pressures could rise. Five- and seven-year debt issued in the low-rate environment 5 years ago, along with longer-dated investment-grade bonds, is now entering its refinancing cycle.
- That refinancing cycle draws from the same pool of capital used by Blue Owl, Apollo, and others to finance data centers, creating a collision. Tech-company financing in July and August this year “seems” to have reached $218B, versus roughly $80B for all of 2025. CSPs, semiconductor expansion, and equity issuance all add asset supply, so he sees capital markets as a supply-and-demand market, not merely a valuation market.
2. His Original June Short: Can the Math Be Justified?
- Herman’s most important judgment at the time was that Anthropic would eventually collapse, forcing the market to question the return on AI investment. If companies spend $800B this year and $1T next year, how much revenue would next year’s $1T need to generate to justify today’s investment?
- Whether total revenue can validate total CapEx is the core accounting exercise linking semiconductors and AI. If expectations fail that test, semiconductors may not receive a valuation; if the math works, the market could produce another rally like May’s. Dollar debt, dollar interest rates, and inflation add macro pressure to valuations, but Herman believes parts of the bear case are valid, parts overstated, and parts insufficiently understood.
3. Why He Turned Bullish, Part 1: Bad News Could Not Move the Market
- Anthropic’s $65B ARR announcement was a “very large bearish catalyst,” yet there was no deleveraging, large-scale forced liquidation, or rise in volatility afterward, and CDX did not move sharply higher. CDX IG rose only modestly over the prior 2 weeks, against a backdrop of heavy financing activity.
- The absence of follow-through means funding-market leverage was already close to exhausted; there was no continuing wave of liquidations or contagion. On June 2, momentum funds were still adding to positions and those positions were vulnerable to a squeeze. At current valuation levels, the case for continuing to short has weakened materially.
- His long-term framework still allows for volatility. Over the next 2 years or longer, the market may repeatedly ask whether improving fundamentals will receive a valuation, and how much of one. He sees no change in corporate performance, AI demand, or token demand; NVIDIA and subsequent quarterly results are more likely to improve than deteriorate.
4. Why He Turned Bullish, Part 2: Open-Source Models Are a Game Changer
- The emergence of open-source models could make CSPs profitable. The most direct evidence would be positive free cash flow, or cash flow that is less negative than before. CSP revenue could then justify CapEx internally, allowing valuations to reset higher.
- NVIDIA’s latest results were excellent, but Herman sees them as a catalyst rather than the fundamental issue. The real questions are how large and durable the front-end demand market is, and whether there is enough capital to keep funding investment.
5. This Is 1998, Not 2000
- Herman’s biggest disagreement with the bears is this: “The bears think this is 2000; we think the worst is 1998, or even some point before 1998.”
- During the internet bubble, he believes there was a great deal of circular activity inside the system: one party bought advertising and another received the advertising spend, but no meaningful money entered from outside the system. Today, OpenAI is “nearly profitable,” Anthropic’s operating business is already profitable, and CSP Enterprise AI profitability and cash-flow data are improving.
- The other difference is that today’s CapEx forecasts are being made by trillion-dollar TMT companies that lived through the 2000 crash, not by startups that can forecast freely. Herman believes these companies keep their own internal books and will not bet their entire future on Anthropic simply because Anthropic has a backlog.
6. The Bottleneck Is Supply, Not Demand
- Herman cited Jensen Huang’s view that apparent ceilings in token usage and AI coding are caused by supply, not demand. Data-center utilization is only around 70%, but that is also a supply constraint, not evidence of insufficient customer demand.
- Anthropic’s ARR growth has slowed amid internal competition, business diversion, and lower token prices, but it has not turned negative. Herman’s explanation is insufficient capacity and model supply—not a lack of people who want to code.
7. What Exactly Is the $65B ARR? Definitions, Price Cuts, and a Supply Ceiling
- Mr. Z asked whether the move could be designed for an IPO—to push the number down first and then drive it higher. Herman said that was possible, but the disclosure could also reflect regulatory requirements or an effort to standardize reporting before a listing.
- The second explanation is price cuts. At his peak, Herman used roughly $3,000 a day in usage beyond his membership allocation; at the low end, the figure was around $1,000 to $2,000. After prices were cut, his own usage did decline.
- The third explanation is a supply ceiling: providers cannot supply as many tokens as customers want. Microsoft, through Enterprise AI, and Adobe already have backlog orders, but not enough compute to fulfill them.
8. The Wrong Anchor: Compute Outside the 27% Can Also Be Monetized
- Herman said Anthropic and OpenAI together accounted for only 27% of compute deployed over the past 2 years. They are currently the model companies with the clearest commercial monetization, but they should not be treated as the sole anchor for the entire AI industry.
- Roughly another 70% of compute is being used by CSPs to develop Enterprise AI. Herman said Microsoft could “somehow add $50B,” while Enterprise AI could add more than $20B. That revenue does not come from users calling Anthropic or OpenAI APIs; it comes from Microsoft deploying compute and serving enterprise customers directly.
- He also said Microsoft has $20B of revenue this year, equivalent to roughly $40B on an ARR basis. Of the original 100, 30 are making money; once the other 70 are developed, they may divert revenue from the original 30. But Anthropic collapsing does not mean all of AI collapses.
9. The Economics of the Cards: A String of Relatively Certain Numbers
- Herman said DeepSeek’s costs are already far below Anthropic’s, yet its gross margin remains 80%. Alibaba’s report puts the payback period at 2.2 years. 梁文锋 has also said that if he had the money and could buy the cards, he would convert all of it into cards.
- Microsoft’s Enterprise AI teams underwrite ROI on a 6-to-10-month payback. V100s are still in use, and H100s can run tasks originally requiring higher-specification cards after engineering optimization.
- Buying Anthropic at its current PE could still lose money, Herman said, but whether the sellers of cheap compute can make money is a more certain question. Gross-margin allocation among GPUs, CSPs, and model vendors may change, but compute assets still have a defensible return profile.
10. Trillion-Dollar Companies Keep Their Own Books; They Are Not Betting on Anthropic
- Google, Amazon, Microsoft, and Meta are all pursuing their own OEM and productization strategies. Herman does not believe they are waiting to see whether Anthropic delivers results before putting the fate of trillion-dollar companies on Anthropic.
- Even if Anthropic went bankrupt tomorrow, these companies could still build around their own users, sales forces, resources, and enterprise data. Mr. Z added that several CSPs are expected to spend roughly $725B of CapEx this year. Herman believes they must have seen their own internal data and found profitability and high ROI in that data.
11. Jensen’s $500B Infrastructure Deal: Opening the Door to Wall Street Debt
- On the $500B infrastructure transaction linked to NVIDIA, Blackstone, Apollo, and BlackRock, Herman said he believes Jensen Huang’s judgment. One core objective, in his view, is to extend the GPU depreciation cycle and make GPUs financeable as productive assets, like cars or aircraft.
- His Wall Street friends have not yet accepted that argument. Herman does not think the transaction automatically proves the bubble is large, but it will certainly bring more capital into semiconductors. If the funding vehicle eventually fails, Jensen could be blamed in the same way participants in the CDO structure were blamed.
- He compared it with Chinese real estate: early on, developers relied on their own equity capital; later, government capital, local-government financing vehicles, and local-government debt entered the system. Data centers still rely mainly on equity financing today. Infrastructure financing opens a door for Wall Street debt to enter more easily—not all at once, but potentially with far-reaching consequences.
12. Why GPU Leasing Works: Cash Costs Are an Order of Magnitude Below Rent
- Mr. Z questioned whether GPU infrastructure financing could rely on stable cash flows in the same way as highways or wind farms. Herman’s answer was that the cash costs of GPU power, hosting, maintenance, and related services are an order of magnitude below the rent the market is willing to pay.
- A100s are fully rented. Around 95% of the capacity of H100s and GPUs released in 2020 and 2022 has already been booked. V100s are still being used, which Herman sees as the simplest evidence that the infrastructure-financing logic works.
- He cited Andrew Ng’s estimate of a $30T brain-labor market. Labor substitution is a process, not an overnight event in which everyone is fired; 10% to 25% could be displaced each year. The terminal risk remains that all capacity comes online, the unemployed cut advertising and consumption, and the economy freezes in crisis—but “definitely not today.”
13. CSPs Fight Back: From Being Squeezed on Both Sides to Holding Pricing Power
- In the past, CSPs were unsure whether to train their own models and could only rent compute to 2 model companies. Model companies could have gross margins near 90%, while OpenAI and Anthropic bought compute with little regard for price and might even accept $50 per watt. CSPs were squeezed between model vendors and storage suppliers.
- With open-weight models, Microsoft and other CSPs can run post-training or further distillation. Herman believes this path differs legally from directly distilling closed-source models in the United States; under his assumption, the former is permissible.
- Once CSPs have open-source models, they do not need to sell compute only to model vendors; they can sell a full service directly to enterprise customers. They could push the original $50-per-watt price down to $30-$35 and pass the pressure to downstream suppliers. The main targets would be suppliers charging excessive prices; NVIDIA and optical vendors charging normal rates would not necessarily be affected.
- The 2 buyers that previously cared relatively little about price were Microsoft and OpenAI. Once their ability to pay is no longer unlimited, more bargaining power returns to the CSPs. A model company without comparable users, sales, security infrastructure, and compute resources will struggle to compete with a CSP.
14. The Dario Proposition: Who Has the Most Efficient Inference?
- Herman cited Dario’s framework: the key questions are “who has the most efficient inference” and “who trained the best model.” As models become more equal, the market will focus increasingly on inference efficiency and compute scale, with more bargaining power ultimately accruing to CSPs.
- He drove home the enterprise advantage with a series of questions: model companies do not have Microsoft’s sales force, user base, resource infrastructure, security systems, or compute. Qwen is currently one of the more widely deployed models in enterprise. DeepSeek creates political and other concerns for some users, while Qwen creates fewer.
- That does not mean he is bullish on every SaaS company. Falling token prices and improving model capability make the SaaS business better, but whether any specific company can beat Microsoft remains a separate question.
15. Storage: The Gross-Margin Story Built on Price Hikes Is Nearing Its Ceiling
- Mr. Z mentioned Micron’s roughly 86% gross margin and the possibility that Apple products could rise in price by 30% to 40%. Herman reiterated that storage vendors are increasing gross margin and revenue mainly through price hikes, not capacity additions or product iteration.
- Storage gross margin is already around 85%, even 87%, in his view. The room to move toward 90% is limited; once gross-margin growth peaks, valuation has a harder time moving higher.
- He described front-end Anthropic and back-end storage as AI’s 2 pressure points. Excessive storage prices reduce the share of GPUs and optical components in the overall solution and lengthen CSP payback periods, giving the entire industry an incentive to push storage prices lower.
- Storage demand may indeed be unlimited, but so are demand for compute, TSMC, and CPUs. “Unlimited” does not mean any one link in the chain can charge a 90% gross margin at will. NAND is his personal preference, though he acknowledges that DRAM demand is very strong.
16. KV and Task Decomposition: Compute and Storage Can Partially Substitute for Each Other
- Herman said KV cache scales with the size of the context being read. Once work is split among agents and each agent’s context becomes shorter, KV cache falls approximately with a square-root relationship rather than linearly with the number of agents.
- He said he uses Ultra Code with Opus to call more than 100 agents for coding and review, and believes the results are no worse than Fable running as a single model—possibly stronger. The model that truly needs high intelligence is the one directing and decomposing the work; many smaller models can handle the remaining tasks.
- This approach increases GPU calls but can reduce context and KV cache, creating some substitution between GPU and storage demand. If the work is not decomposed properly, agents may call only a small number of parameters, leaving GPUs idle while storage is used heavily.
- He personally prefers NAND. As agent counts and inference calls grow, demand for CPUs and NAND may be more certain than demand for DRAM, with a potentially longer cycle. That does not mean NAND will necessarily replace DRAM.
- He also criticized serious internal conflicts of interest in Samsung’s sales organization and asked 崔泰源 to explain his stake in SK Telecom, as well as why he would swap his own HBM for NVIDIA cards while also building his own CSP. Korean storage stocks could rebound and even break their previous highs, but portfolios should not be allocated entirely to storage.
17. Chinese IPOs and Optical Communications: Low Pricing and a More Certain Optical Thesis
- Mr. Z asked whether YMTC might IPO in a few weeks. Herman said CXMT’s DRAM performance is very strong, though it may not yet be able to make HBM. YMTC is also very strong, but he does not know what valuation the Chinese capital markets will assign it.
- Chinese IPOs are often priced low initially, so companies such as Unitree and CXMT could rise 300%, 400%, even 500% or 600% on debut. That differs from Korean-market financing, where the offering is priced against an existing valuation.
- On optical communications, he believes delays in CPUs or Vera Ultra would not change the long-term direction. Optical components will represent a larger share of the overall solution, demand will expand, and the main bottlenecks will ultimately be volume production and packaging supply. As demand rises, optics itself could become the next bottleneck.
18. Power Semiconductor Timing and the Broader Shortage-Cycle Framework
- Power semiconductors should be viewed through the delivery schedule for Rubin. Demand will show up around the third quarter, later than optical and storage demand. This cycle is better approached through optics and storage first, with power semiconductors coming later.
- In GFS’s optical-communications path, Herman does not think monolithic integration is necessarily the best route, though he does not see it as the most important issue.
- His overarching framework is that major semiconductor shortages come from major shortages of token capacity. A shortage of tokens ultimately converts into demand for semiconductors. During a shortage cycle, whoever owns the factories and capacity has pricing power; beyond TSMC, IDMs including Intel could benefit.
19. Foundries Are Data Packages; Major Customers Need Backup Capacity
- Herman believes Intel foundry will “definitely succeed—you have no doubt about this.” In his analogy, a foundry is fundamentally a data package jointly tuned by the customer and the process. A conventional fab using its own package might produce 14nm; loaded with TSMC’s data package, it could produce results above its previous level.
- TSMC’s advanced processes and capacity have been tuned to a significant degree with Apple. Apple cannot accept a foundry that is entirely AI-oriented. Once NVIDIA becomes an important customer, Apple will need Intel as a backup. Google is also not a traditional TSMC customer; if it needs roughly 11M to 12M wafers next year and TSMC can provide only part of that capacity, Google could turn to Intel for tuning and orders.
- Herman used a “king of Causeway Bay” analogy: TSMC remains the big brother, but major customers cannot depend on a single supplier and need another enforcer to strengthen their bargaining position.
20. Intel: Factions, 陈立武, and an Underappreciated Foundry
- Herman believes TSMC has a 刘德音 faction and a C.C. Wei faction. 利普·布坦 also has a Berkeley background, so people associated with those factions may feel little psychological resistance to moving to Intel. He summarizes foundry execution as “people and data packages”: the data lives in people’s heads, which is why Intel’s yield improvement could come faster than expected.
- 陈立武 took over a company whose stock and internal power structure had both reached a low point. Herman believes the Trump factor is behind him and that multiple rounds of internal faction-cleansing have already taken place. After the stock rose from roughly $18-$20 to near $100 and briefly around $140, his credibility on the board became very strong.
- Intel will not easily replace TSMC, which retains incomparable accumulated advantages in products, know-how, and customer relationships. But Intel has asymmetric advantages in some tape-outs, chiplet packaging, and advanced packaging, along with greater political resources.
- Herman cited roughly $10.5B of foundry-related investment at the end of last year, $15B this year, potentially even $20B, and a related figure of $12.5B to $13B by the end of next year. At a 1x valuation or 2x PB, he considers the current valuation relatively safe. Short-term sentiment, however, has run too far; no fundamental problem does not mean no short-term risk.
21. Sentiment as a Contrarian Signal and the 7709 Leveraged-ETF Lesson
- Herman recommends using sentiment indicators in reverse. Put the KOLs whose calls you consider most accurate on a blacklist and exclude their opinions; when most KOLs are uniformly bullish, consider shorting, and when they are uniformly bearish, consider buying.
- On CSOP 7709, he believes the 2x leveraged ETF combines leveraged equity exposure with short-volatility exposure, making it dangerous at a bottom. Bottoms are usually choppy rebounds, where volatility and path decay erode value.
- 7709 is better suited to a view that prices will rise in a straight line while volatility falls. When buying the bottom, buy the underlying directly—for example, 000660—instead of a 2x leveraged ETF. Broad consensus buying of 7709 can itself signal overheated sentiment.
- Intel’s short-term sentiment is dangerous, but Herman believes its packaging, CPU, advanced packaging, and potential storage business with Tower Semiconductor are not fully priced in. If price increases are incorporated into CPU estimates, he believes next year’s EPS could easily reach $6.
22. TeraFab and Grok
- TeraFab wants to work with Intel on storage, but Herman does not think Elon Musk currently has that much cash. He does not know where the funding would come from, but Musk frequently manages to raise money and also benefits from the Trump factor, so the possibility that he finds the capital cannot be dismissed.
- One reason Grok has improved is SpaceX’s ability to deploy compute quickly, sell compute, and form compute clusters. Herman believes the relevant variable is not just today’s model-training performance but also the ability to deploy compute.
- A second reason is that Cursor provides substantial post-training data and front-end traffic. Herman’s current view is that Grok may still trail GLM, but it is improving quickly and is inexpensive. If its valuation fell to $1T, he believes an entry opportunity could emerge.
23. Model Rankings in Actual Use and the Four Key Inputs
- Herman stressed that his rankings come from hands-on use: OpenAI’s o3 is strongest, followed by Fable, then DeepSeek and GLM; GLM may be better than Kimi. In actual use, GLM and DeepSeek are no worse than Opus; Herman even thinks DeepSeek is stronger. Grok handles simple tasks well, but still struggles with complex ones.
- He believes model usability depends on 4 things: the model itself, the compute behind it, the data, and how the AI agent organizes and decomposes tasks. Claude Code’s advantage comes to a significant degree from the fourth factor.
- He distilled Claude’s task-decomposition method into skills in Codex, which is why his own Codex experience is so strong. But this is personal usage experience, not a universal ranking across all models or all tasks.
24. Protoss and Zerg: Failing to Use Your Tokens May Mean You Are Being Replaced
- Herman’s framework is that “the AI era has only Protoss and Zerg.” People who use models constantly and know how to use them are Protoss; most people are Zerg. He repeated the claim that 0.1% of users consume 99% of tokens.
- If someone cannot use up even a $200 Claude Max subscription in a week under a normal workload, that suggests the work has little room to expand and that the person may be among the first to be replaced.
- Herman said he wrote 450,000 lines of trading-grade code over the past 3 months, while the Shenzhen team wrote roughly 60,000 lines in total over the past year. He estimates that producing 450,000 lines would require a professional trading team of around 20 people working for 1 to 1.5 years. The system is close to zero bug, but not absolutely free of problems.
- In high-frequency trading, 2 milliseconds is “zero” to him. Reducing latency from 2 milliseconds to 1.7 milliseconds requires coordinating many clouds, synchronizing clocks, and handling internal-network and security issues; token usage can rise exponentially. In his view, people saying AI coding has reached its limit often have not used it at this depth.
25. AI Replaces Organizational Structure, Not Just Programmers
- Herman believes AI will not only replace code farmers; it will also replace the organizational structure built around them. Front-end, back-end, product, and testing functions may no longer need to be separate teams—one person could call multiple AI agents to complete an entire workflow.
- He used the difference between a battleship and an aircraft carrier to describe organizational efficiency, and believes 2 people doing high-frequency trading could potentially compete with large institutions.
- Programming moved from assembly to C and then to natural language. The scarce resource is not the executor but the person who can pose complex questions, possess know-how, and control proprietary data. AI may not be able to teach that person how to trade, but it can implement the method already in their head.
- Mr. Z cited Garry Tan’s line that “Don’t boil the ocean” used to be the rule, whereas in the AI era people can “boil the ocean.” Herman said a friend’s children study literature and history; work that once might have required a postdoc 5 years to complete can now be done by a high-school student in 2 months and published in a specialist journal.
26. Why Microsoft Can Sell a Token for More Than GPT: The Meter
- Model vendors have solved only the I in ROI. A programmer may go from working 15 or 16 hours a day to completing the task in 10 minutes, but the boss may not be able to see the resulting output.
- Microsoft, Amazon, and other enterprise products have a meter that can measure how many tokens produce how much additional sales or reduce how many returns. In cross-border e-commerce, for example, an API can track after-sales service, customer conversations, and return rates, allowing the company to calculate the AI-generated return.
- If spending RMB0.5 produces RMB2 of revenue, the enterprise can see the RMB1.5 difference and will keep investing. Herman believes this makes Enterprise AI growth more durable.
- He said Microsoft has roughly $20B of revenue this year, equivalent to about $40B on an ARR basis. If several more enterprise use cases work, 3 of them together could reach the scale of one Anthropic.
- He rejected semiconductor suppliers’ attempts to infer token demand directly. The supply side can question supply, but it cannot conclude that demand has peaked solely from orders. At minimum, it should ask enterprise customers such as Google. Google, Amazon, and Microsoft are still chasing one another, and Google itself says there is too much demand to fulfill.
27. Healthcare: WuXi AppTec Is the TSMC of Pharmaceuticals
- Responding to Mr. Z’s comments on Moderna’s mRNA progress, Herman recalled the case of “an Australian treating a dog,” saying the direction was already visible then, though he does not know which pharmaceutical company will ultimately break through.
- AI will make DNA screening much faster, but screening accounts for only around 5% of costs; drug testing and the subsequent pipeline account for roughly 95%. More precise genetics should lower the cost of repeatedly searching for methods and data in Phase 2 and Phase 3 and make the process more efficient.
- Herman sees WuXi AppTec as the TSMC of pharmaceuticals and believes it would be difficult for anyone to beat the company on cost. He said the cost of one case could differ by 20x between China and the United States.
- His broader view is that AI’s original capability was parallel computation over large datasets to identify patterns, which suits healthcare, the military, and trading. Coding and conversation came later. More tailor-made medicine could emerge, with healthcare advancing through large-scale processing of genetic data.
- Software will also benefit from cheaper tokens and model parity, but Microsoft is the main competitive concern for SaaS. In the near term, companies such as Snowflake may still make money.
28. The AI Cancer Theory: How the Cycle Ends
- Herman calls AI “the cancer of the economy”: its ROI is so high that it absorbs debt, equity, returns, and other resources. That is not necessarily bad immediately, but if AI replaces large numbers of white-collar workers, the economy will need a new arrangement for the unemployed.
- He compared it with the social transformation after 1848 and said AI could ultimately produce severe unemployment, an economic downturn, and insufficient consumption. Silicon-based consumption still requires carbon-based consumption; without jobs and income, no one buys hamburgers or coffee or places advertising.
- His timing call is that AI is still in phase 1 or 2. The terminal outcome above would arrive in phase 4—“it’s not today.” Governments may eventually restrict AI or provide support through mechanisms such as UBI, but he does not know which solution will prevail.
- He believes Microsoft and Amazon’s financing needs and CDS could rise next year. The real danger is debt costs pushing operating companies toward regional-bank levels, leaving some regional banks unable to finance themselves and creating a run-like problem.
29. 190 Basis Points: The Bubble Is Inevitable; It Has Simply Flowed to the Wrong Place
- Herman believes high long-term dollar rates are not driven solely by the amount of U.S. debt or a problem with the dollar itself. Nominal GDP growth of roughly 6% to 6.5% exceeds the roughly 4.7% long-term interest rate, and that roughly 190-basis-point gap creates asset bubbles.
- Excluding AI, the real economy cannot absorb higher interest rates. AI allows another part of the economy to tolerate them, placing Protoss and Zerg in the same world. Policymakers will try to prevent long-term rates from rising further; otherwise, real businesses such as pizza shops and coffee shops will be destroyed.
- Herman believes the current bubble is concentrated in areas such as Walmart and Starbucks. Starbucks trades at roughly 36x forward PE versus about 12x for NVIDIA. The real risk arrives when Starbucks can no longer sell coffee—not when the market labels NVIDIA a bubble today.
- The bubble may flow into semiconductors once the market decides semiconductor demand is non-cyclical and CSP free cash flow turns positive. He does not think the rally is limited to the November midterms or already over. But if CSP cash flow deteriorates further next year, the sector could still sell off once before recovering.
30. Rules for Retail Investors and His Own Accounts
- Herman advises retail investors not to trade the market if possible and not to fight quantitative robots. They can feed all their past trades into AI and check whether they were actually profitable.
- He admitted that when the $65B news came out, he knew it was “fake” and not especially meaningful, yet still could not resist liquidating his entire position with one click. Even if the trade is right once, repeating that behavior over the long term may be no better than a coin toss.
- Most of his savings are in Google, with additional positions in Meta and Goldman Sachs, held for the long term without watching them closely or caring much about price. His semiconductor trades are mainly in his trading account. He also mentioned continuing to dollar-cost average into 3M.
- His Meta position is smaller than his Google position because he once sold Meta to buy an unclear “Topic[?]” position, which he now thinks may perform terribly. On the joke that manually placing orders makes him a retail investor, he admitted that he sometimes even forgets his account passwords.
31. Crypto: Pricing Power Sits on Binance, and Liquidity Is Leaving
- Herman says trading price discovery for every coin sits on Binance; on-chain markets have “no pricing power, none at all.” Ninety-nine percent of crypto demand and trading comes from exchanges rather than on-chain usage.
- Stock tokens on Binance are taking substantial liquidity. In the past, coins outside Bitcoin, Ethereum, and Solana could regularly post daily volume in the billions; today, tens of millions is already considered large. Herman used to trade roughly $2B to $3B a month on Binance and has fallen from VIP 7 to VIP 5. Many friends who were VIP 9 have also fallen to VIP 6.
- His high-frequency taker profit has dropped from roughly $100 per trade to $25. Even on better days recently, it was only around one-quarter of the old level, indicating insufficient displayed depth and overall capital. Market-maker profits are extremely low; if they stop quoting, takers will have no orders to hit and the market could “fall apart completely.”
- He kept buying as BTC fell below $65,000 and has not sold. Long term, he can see $200,000, $500,000, or $1,000,000, but short term the market first needs to answer where the funding comes from, because exchange capital is being pulled into equities.
- He calls Hyperliquid a “Pi Xiu temple.” He says orders take roughly 900 milliseconds, the system is centralized, and trades are put on-chain after the fact. Before execution, the platform first hedges on Binance or checks the price; when a fill would benefit the platform it withholds execution, and when it would hurt the platform it fills the order. He does not want to trade there.
32. Listener Q&A: Sivers, AXTI, and How to Buy Optical Exposure
- On Sivers Semiconductors, Herman believes the company does not currently have enough capacity. European suppliers would struggle to expand quickly; Sivers’ Scottish facility is small, and if it sends wafers to Win Semiconductors, indium phosphide may not even have a PDK available. Win itself also lacks sufficient capacity.
- He sees a clear distinction between stock speculation and real capacity. When supply is genuinely short, the market can substitute expected capacity for actual capacity, but that does not prove a company truly has capacity. Lumentum, Coherent, Nokia, and other U.S. names are expanding at roughly 10x the scale and have received funding support from Jensen Huang.
- On AXTI, he agrees that InP wafers are genuinely in short supply and are a bottleneck, but bottlenecks cannot be bid up indefinitely. He recommends making large-cap names the core of optical exposure rather than chasing extremely high-beta stocks; in optics and testing, larger companies such as AEHR are worth considering.
33. BE Valuation and Congressional Holdings as a Signal
- Asked about BE’s valuation, Herman recalled that during deleveraging, companies with especially strong results are sometimes sold last; they then surge when results are announced and are shorted again. He cited same-day moves of 13% for Intel, 20% for BE, and 50% for AEHR. His strategy is to buy after deleveraging ends rather than get hung up first on the company’s name.
- The listener argued that BE’s power solution could be accepted by both U.S. political parties and capital markets, reducing political risk. That was the listener’s view, not Herman’s independent confirmation of Nancy’s or either party’s position. Herman mainly addressed the valuation framework: look at forward PE. He also said he continuously monitors the holdings of Republican and Democratic lawmakers.
- The listener mentioned that Intel’s EPS could reach $6 next year. Herman used that to illustrate that even if the current PE looks above 100x, the forward expectation must be included. He recommended congressional-holdings monitoring tools such as Quiver Quantitative, which costs around $200.