Silicon Valley Coordinates x Alan Du, Microsoft's Strategic Investor: Software Moats in the AI Era
Silicon Valley Coordinates x Alan Du, Microsoft's Strategic Investor: Software Moats in the AI Era
Summary
- Alan Du’s core judgment: In the AI era, software’s real moat is enterprise data and workflow know-how behind the firewall—not the model itself. His data anchor: GPT-4 used roughly 1.5 petabytes of training data, while companies such as JPMorgan and PayPal hold 15, 20, and 30 petabytes of data behind their firewalls—“the truly useful data is actually behind the firewall, and companies generally building agents cannot access it easily.”
- The software most vulnerable to disruption is easy to identify: “data-hauling” software, RPA, single-function software, and consulting/market research. He cited an MRI CD: a doctor’s scan required several hundred dollars of specialized software to read, but a friend simply dropped the file into Claude and “it built software in real time to read it.” By contrast, “natural focal points” for enterprise information flows such as ServiceNow and Atlassian will not be replaced by agents; they will enable them.
- Microsoft’s moat “has become very strong”: global enterprise data and accumulated know-how, combined with procurement inertia and sales channels as nontechnical barriers. “Selling to enterprises and selling to SMBs are completely different—you even speak different languages”; he also acknowledged that GitHub Copilot still has “some room for improvement” on user experience.
- The “bubble” debate requires separating value from price: AI’s value is unquestionable, but FOMO is unmistakable in pricing. AI compute and early-stage investment have surpassed or are approaching $1T, while the largest AI companies, including OpenAI and Claude, may have no more than $10B in combined annualized ARR; over the next 3–5 years, they will need to generate roughly $800B in revenue to justify the upfront investment.
- Improving utilization of the existing GPU fleet matters more than building new data centers—the market cites 30–40%, while “we may think it is lower.” Training and inference have high failure rates, with roughly 10–25% of failures attributed to data-cable connections, while over-provisioned backup capacity sits idle; potential fixes include VM-style GPU virtualization and replacing copper and fiber with millimeter waves. This is something the market needs to figure out over the next 12–18 months.
- Agent payments are “basically still at the narrative or demo stage,” and real adoption will require entirely new infrastructure. PCI compliance bars agents from holding personal credit-card information, while Stripe virtual cards and Visa/Mastercard tokenized credentials still borrow a human identity. Traditional rails may not support one-to-many-to-many payments, sub-cent micropayments, or streaming payments; stablecoins have “some potential” in these complex settings, but buying a pair of running shoes does not require them.
- Security is shifting from passive defense to active defense: agents create a new attack surface, insiders have become a source of threats, and we still have not seen runtime security truly operate at scale. AI-generated code remains concentrated at the application layer because underlying code must be 100% reliable, which conflicts with the probabilistic nature of agents. The durable opportunity is in software design and planning: “context is king; whoever captures context accumulation may ultimately be the one who wins.”
- The actionable advice for founders: prioritize speed over valuation, break $30M–$50M rounds into smaller tranches to build a POC first, and differentiate on industry know-how rather than tools. His direct advice to Chinese founders: build diverse teams and leave the comfort zone—“you are not selling to Chinese companies alone.”
Deep dive
1. M12: Microsoft’s $150M-a-year, balance-sheet-funded strategic investment arm
- Alan Du opened by introducing M12: it invests roughly $150M a year directly from Microsoft’s balance sheet, with a typical ticket size of $5M–$10M. The fund focuses on the space between Series A and Series B, spanning compute-chip hardware, horizontal software and databases, vertical applications, cybersecurity, gaming, and frontier technology.
- His candid division of labor between lead and follow-on investors: financial investors are “more professional” at pricing and setting terms, while strategic investors bring their own commercial considerations, making financial investors the more natural leads. Most strategic investors write only $500K or $1M checks; strategic funds capable of investing $5M–$10M in a single round remain relatively rare. M12’s larger contribution is the commercial value of Microsoft’s enormous downstream ecosystem.
- His advice to founders raising capital is to speak with both sides. Deep-tech AI founders face high technical barriers, and “if you only talk to financial investors, the depth of the conversation may be limited.” Once a financial investor is identified as lead, bring in a domain-savvy strategic investor to round out the financing. “That combination is very powerful and sends a very strong signal to the market.”
2. “Why can’t Google do this?” is an overblown fear
- Alan’s first-principles answer: Large tech companies may be enormous, but “when it comes to a specific product, the team actually working on it is very small,” with limited resources at its disposal. More important is inertia: existing products are too profitable, and “there is resistance to disrupting yourself.”
- Startups have structural advantages in speed and the cost of failure. When Amazon, Google, or Microsoft makes a mistake, customers are “very unwilling to forgive you.” Startups, by contrast, often sell to other startups that are more willing to tolerate experimentation, allowing them to iterate quickly.
- The way out is positioning. Large platforms build general-purpose technology for 80% of users; startups should find a niche vertical they understand exceptionally well. “Just do the 1% or 2% of vertical software—you can actually build something very large.”
3. The trillion-dollar question: Which software is most vulnerable to disruption?
- Asked the “trillion-dollar question,” Alan declined to lump all SaaS together but offered a clear profile of the most vulnerable businesses. First is shallow software that merely moves data around. Second is RPA: the shift from “record and repeat” to an agent that “reasons and then acts” could replace the vast majority of companies providing RPA services.
- Single-function software is also highly exposed. After a friend underwent an MRI, the doctor sent over a CD whose contents required several hundred dollars of specialized software to read. He simply dropped the file into Claude, which “built software in real time to read it.”
- Consulting and market research may also be replaced quickly. Consultants primarily read secondhand material; “you simply cannot outrun an agent. It can read through the limited information available online in a very short time, and there is no way to compete with that.”
4. What will not be disrupted: the “natural focal point” of enterprise information and the economics of accountability
- Companies such as Atlassian and ServiceNow are “the natural focal point of the entire enterprise information flow.” From IT ticketing to Slack conversations, information converges there. “It will not be replaced by an agent; it is what enables the agent. It gives the agent the information it actually needs.”
- He used an automotive assembly line to argue that enterprises often should not build agents entirely in-house. A healthcare company, for example, is not differentiated by assembling an AI agent; doing so requires hiring expensive specialists and handling testing and compliance. More important is the transfer of accountability: if something breaks, the enterprise can ask ServiceNow to roll back and patch it. “If you build everything yourself, the person responsible is you… Once an enterprise does the math, it quickly realizes that building the entire system itself may not make sense.”
- The hardest data point in the discussion was the comparison in scale: GPT-4 used roughly 1.5 petabytes over its full training run, while companies such as JPMorgan and Alan’s former employer PayPal have 15, 20, and 30 petabytes of data behind the firewall. “Intelligence at the level of GPT-4 used only that much data. Once enterprises make proper use of their own data, the performance gains can be substantial.” On the Atlassian CEO’s claim that the moat is decades of process experience, Alan’s verdict was unequivocal: completely correct.
5. Traditional software plus AI and AI-native software can both thrive—by customer segment
- The dividing line is the buyer. Large enterprises have complex procurement processes, strict regulation, and failure costs that are “even intolerable,” making them more likely to buy mature service systems such as Microsoft’s. Customers serving emerging industries and startups are more willing to adopt AI-native tools: the employees new companies hire “have been using these new tools since school,” while the cost of experimentation is low and iteration is fast.
- On how startups avoid being swallowed by big-tech features, the answer remains the same: “You simply have to move faster than everyone else, and be willing to try new tools, new platforms, and new models.”
6. Microsoft’s moat “has become very strong”: data, inertia, and distribution
- Alan returned to workflow know-how and proprietary enterprise data. Microsoft serves every enterprise in the world and most individuals; “few companies can match us in the data we can see, or in the know-how accumulated by the people we see during the production process.” Microsoft may not move as fast as others, but it can build an agent experience that works exceptionally well.
- The more important barriers are nontechnical. Onboarding new software at an enterprise is highly complex and governed by inertia. “Whether a sales team knows how to sell to enterprises and whether it knows how to sell to SMBs are completely different things—they even speak different languages.” Existing channels and distribution are “an extremely difficult moat to cross.”
- His self-criticism is worth recording: “We acknowledge that GitHub Copilot has some room for improvement on user experience.” But Microsoft’s huge user base and ecosystem mean that when it needs to improve quality and performance, “we can do it very well.”
7. “Bubble” is an imprecise label: trillion-dollar investment versus less than $10B of ARR
- He insisted on separating value from price. AI “is absolutely valuable; there is no way to question that.” But FOMO is clearly present in pricing: companies with impressive founders can raise hundreds of millions of dollars, even $1B, in seed rounds while their products are immature and their technical development has barely begun. That suggests the market may be overheated and exhibiting some FOMO.
- The gap can be quantified. Investment in AI compute and early-stage companies has surpassed or is approaching $1T, while “the largest AI companies, including OpenAI and Claude, may have no more than $10B in combined annualized ARR.” The estimated gap is $800B–$1T, implying that these companies—or emerging AI companies—will need to generate roughly $800B in total revenue over the next 3–5 years to justify the investment already made. “That gap is still quite large.”
- Adoption is not moving in lockstep. AI-assisted coding is advancing extremely quickly—“even within Microsoft, the vast majority of the coding process has been replaced by agents”—while traditional manufacturing and healthcare are only beginning to move. The bottlenecks are regulatory and privacy concerns, not technology.
8. More urgent than building new data centers: GPU utilization may be below 30%
- Alan framed the market’s next question: Some people estimate GPU-cluster utilization at 30–40%, but “we may think it is lower.” After more than $1T has been invested in data centers, should the industry keep building new ones or focus on improving utilization of what already exists? “This is something the market needs to figure out over the next 12–18 months.”
- The mechanics behind low utilization are clear. Training and inference have very high failure rates, with failures arising from software as well as installation of the links between GPUs and between boards. Then there is over-provisioning: large pools of backup GPUs sit idle, ready to take over when workloads are shifted.
- The solutions run on 2 tracks. On the software side, GPU virtualization could apply the same logic as VM, Docker, and Kubernetes to improve hardware utilization. On the hardware side, of roughly 70% of failures, perhaps 10–25% occur at data-cable connections. Fiber is expensive and has physical limitations of its own: a 1°C temperature change can affect performance, and bending can attenuate the signal. New technologies are therefore testing millimeter waves as a replacement for copper and fiber, aiming to lift utilization from 30% or lower to 60–70% while reducing energy demand.
9. M12’s portfolio and priority sectors: 100-plus investments, 15 unicorns, and 6 IPOs
- The host introduced M12’s portfolio: more than 100 investments, including 15 unicorns and 6 companies that have gone public. Alan then outlined its priority areas. In hardware, the fund favors optical approaches that replace traditional hardware and conventional chips, with Neuro Flux and D-Matrix as examples. In data, it has backed companies such as MicroOne that help enterprises obtain first-party, clean, high-quality data—“data is one of the ultimate moats.”
- Its security portfolio includes AI-native cybersecurity and model-security companies HiddenLayer and Rich Security, as well as Arise, which evaluates model performance. In vertical applications, its latest investment is Datarail, which uses AI to replace traditional supply-chain workflows.
- Two areas are reopening or newly emerging. Gaming previously attracted very little strategic investment, but “with so many world models or physical models being developed, gaming has become a very natural use case.” The other is the intersection of blockchain and AI—agent management, identity, and payments—which “we are also exploring gradually.”
10. Agent payments: a narrative-stage market facing use cases traditional rails may not support
- The sober assessment of the current market: “It is basically still at the narrative, experimentation, or demo stage.” Large-scale payments by agents are barely visible. The first obstacle is legal: allowing an agent to hold personal credit-card information violates PCI compliance.
- He is not satisfied with the transition solutions. Stripe’s one-time virtual card, preloaded with $100 and disabled once depleted, “is not a sustainable long-term solution.” It is slow and expensive, still borrows a human identity, and changes the dispute from “you versus the merchant” to “you plus the agent versus the merchant.” Visa and Mastercard tokenized credentials are relatively viable for now, “but this cannot be called pure agentic payments; ultimately, it is still using a human identity.”
- The deeper problem is an agent-native identity. In a procurement workflow, an agent may buy data, purchase compute, and pay a model provider at the same time: “Whose identity is it using?” The system needs a verifiable, auditable identity owned by the agent, along with one-to-many and even one-to-many-to-many transactions, micropayments priced in fractions of a penny, and streaming payments. Traditional payment rails may not be able to process the entire flow. Giving agents their own identity and permissions and enabling them to interact with multiple agents requires entirely new infrastructure, and the relevant startups are not yet mature.
- The stablecoin-versus-traditional-rail debate is use-case dependent. Having an agent find a pair of running shoes “does not require any stablecoin at all”; traditional rails are cheaper and more reliable. But in complex one-to-many scenarios where compute is priced in real time per call, “stablecoins still have some potential.”
11. A new security paradigm: agents are a new attack surface, and passive defense must become active
- Agent identity is currently one of the hottest topics: moving from agents borrowing human identities to call APIs toward agent-native identities, with real-time visibility into what an agent is doing, post hoc auditability, and rollback remediation when something goes wrong. Security AI and Vesa are relatively successful examples, and some may already have been acquired.
- Model security is a leading form of active defense. M12-backed HiddenLayer and Protect AI, which was acquired at a high price, conduct red teaming and penetration testing to ensure that models are not exposed to prompt-injection attacks or hackers after deployment, and do not leak customers’ private data.
- His verdict on traditional security companies was blunt: “Many of them simply do not have” the ability to solve the problem. Traditional cybersecurity is passive, “gatekeeper” defense, while the speed and potential risks of agents are beyond what passive defenses can anticipate. The threat actor has changed as well: threats once came almost entirely from outside, whereas the inside can now become a source of threats through prompt injection or agent hallucinations.
- The biggest gap is runtime security. After granting an agent access, “it is very difficult today to truly see in real time what it is doing, whether it is doing what I asked it to do, and whether it is compliant. Many companies say they are working on runtime security, but we have not yet seen anyone truly achieve it at scale.”
12. The limits of AI coding—and the endgame where “context is king”
- The capability boundary is clear. AI is “still some distance away” from writing low-level code and remains concentrated at the application layer. Agent-written code contains a probabilistic element, but 99% reliability is not enough for infrastructure such as a database engine—it must be 100% reliable. A failure there can be catastrophic, while application-layer errors cause limited damage and can be tolerated.
- That creates 2 opportunity areas. The first is AI code review, driven by the surge in machine-generated code: CodeRabbit, Qodo, and Graphite have raised significant capital and are growing quickly because there is still “a certain level of doubt” about machine-code quality. The second is the upstream planning and design layer, where compliance and risk controls are addressed at the design stage. Clover Security, Prime Security, Clearly[?], Sezzle[?], and Depth First[?] have already gained some traction.
- The core idea behind these companies is that front-end models such as Opus and Codex are already very good at code generation. The durable position is to centralize memory and context at this layer: the design layer can see Jira tickets, Slack, and Notion boards, along with company preferences, personalization settings, and industry requirements, all “concentrated at this point.” He also flagged a risk: if Codex one day decides to become a fully closed system like Claude Code, it “could threaten the survival of many open-source platforms.”
- His one-line summary of the endgame: “Context is king… In the entire process, whoever can capture context accumulation may ultimately be the one who wins.”
13. Deployment in regulated industries: a know-how problem, not a technology problem—and the FedRAMP intermediary
- On hospitals, financial institutions, and other heavily regulated industries, Alan argued that the bottleneck “is not at the technology level.” The key is knowing how to execute complex compliance processes. M12 recently invested in an as-yet undisclosed software platform—“possibly to be announced next week”—that helps other software companies, especially cybersecurity vendors, sell to the government.
- The numbers explain the need for intermediaries. Completing the US federal government’s FedRAMP review process costs roughly $5M–$15M and takes 3 years. “The vast majority of companies want nothing to do with a highly regulated, high-risk business like this.” A model similar to an MSSP (Managed Security Service Provider) shifts the compliance burden from the software company to an intermediary with years of accumulated process expertise and distribution. That is also why customers proactively approach M12 portfolio companies after investment: they have been struggling with this problem for years.
14. Advice to founders: speed over valuation, smaller rounds, and leaving the Chinese comfort zone
- The democratization of barriers has moved differentiation elsewhere. Early cloud deployment meant software startups no longer had to build their own data centers; now Lovable, Replit, Cursor, and Codex let people build relatively complete products without learning to code. “The barrier is lower for you, but it is lower for everyone.” The final differentiator is domain know-how: “What are the things only someone who has spent time in this industry would know? That is where real differentiation comes from.”
- On fundraising, he advised against chasing valuation and marquee institutions. “In the early stage, speed matters more: finding a reasonably fair valuation, reasonable terms, and an investor who can genuinely help you is more important.” The entry cost for training foundation models may have fallen, but training can still require $30M–$50M. His recommendation is to split that into 1 or 2 smaller rounds: “Raise $5M first, build a demo-ready proof of concept, set smaller milestones, and you will ultimately move faster.”
- He offered Chinese founders 2 blunt observations. First, many technically excellent founders “still have some gaps in productization and commercialization” later on, so they should consider early whether the team includes people who understand go-to-market and other industries. Second, diversify the team: “Many Chinese founders prefer to start companies with other Chinese… but you are not selling to Chinese companies alone.” Build beyond the comfort zone from the start and find people from different backgrounds who complement the team’s skill gaps.