Pioneers Insight Method Research Author
小宿科技’s 杜知恒: $25M ARR at 6:05 a.m. on Agent Era Day One
Back to Episodes

小宿科技’s 杜知恒: $25M ARR at 6:05 a.m. on Agent Era Day One

Summary

  • By June, 小宿科技 had taken a seemingly “unsexy” Agent Infra business past $25M ARR and reached P&L break-even. The company is closing its Series A and expects to launch its Series B by year-end; roughly 70% of its 100-plus employees work in R&D, and its products span 小宿智能搜索 and the Sky Router large-model API aggregation platform.
  • William is betting that search will shift from “people searching machines” to “machines searching machines,” with B2B’s share potentially rising from single digits to 80%–90%. Of roughly 20B global searches a day, only 2%–3% currently come from B2B; Agents will break queries into subtasks and run multiple rounds of search, driving call volumes up by multiples of 10. William stresses that the thesis will only be realized if Agent user scale and retention undergo a step change.
  • 小宿’s product inflection was not packaging consumer search as an API, but three leaps from full-text retrieval and self-built search to Agent-native capabilities. From last September through January, it worked with Bing to retrieve full text; after DeepSeek-R1 broke out in February, Bing stopped providing APIs to some Chinese model makers and major internet companies, prompting 小宿 to acquire a mature team and build an index at the 100B-plus scale. Bing’s announcement at the end of May that it would retire Search API on August 11 validated 小宿’s decision to invest 3–4 months ahead of the market.
  • William breaks the search moat into scarce talent, at least 6–9 months of engineering work, and a feedback loop built with top customers. 小宿 first aims to match Bing on relevance, freshness, authority, latency and availability, then differentiate through full text, Markdown, multimodality, multilingual support and data compliance. His blunt view: a company claiming to build a native engine with 5–10 people and no traditional search leader “must be a wrapper product.”
  • Sky Router is betting that inference demand will exceed training demand in 2025, with the vast majority of inference delivered through APIs rather than directly rented compute. It does not compete with 硅基流动 on self-hosting and inference optimization for open-source models; instead, it aggregates commercial and open-source models through global nodes, platform stability and cloud-resource operations. William says migrating from OpenRouter would likely save at least 10%–15%.
  • AI Infra’s main battlefield is shifting from “scrambling for GPUs” to “scrambling for data,” while the biggest bottleneck remains customers’ failure to find strong enough PMF. Coding, office work, ad buying, travel booking and personalized feeds are showing early traction, but leading native Agents still have DAUs in the hundreds of thousands, while Tier Two and vertical products have only tens of thousands. William therefore says it is merely “6:05 a.m. on Day One of the Agent era.”
  • Moving from the public markets to become a CEO taught William to respect execution and turned the cost of a bad decision from a stop-loss into an organizational expense. An investment mistake can be fixed by switching positions; a startup strategy mistake means cutting businesses, budgets and people. He now goes to bed earlier because “anxiety is useless”; what he can control is running small experiments, serving customers and continuously correcting course in a market that changes by the day and week.

Deep dive

1. A $25M ARR business turns Agent Infra from a thesis into a business

  • 小宿科技 positions itself as a one-stop Agent Infra platform: 小宿智能搜索 provides Agents with real-time information, while Sky Router aggregates global large-model APIs so small teams can call models without geographic or usage constraints.

  • As of June, the company’s ARR had exceeded $25M and P&L had reached break-even. Its Series A is being closed, and William expects a Series B to launch by year-end; the amount and valuation will be disclosed officially.

  • The team has more than 100 people, around 70% of them in R&D. When William joined full-time in Q3 last year, the company already had an overseas-expansion infrastructure business, but AI revenue was limited and the business model remained unproven. His view was that the existing infrastructure capabilities and the founding team’s experience were enough to find a real use case.

2. “Machines searching machines” will rewrite both search share and query volume

  • William’s core thesis is that 90%–95% of traditional search is consumer-facing. Once users let Agents execute tasks, B2B search could rise from a single-digit share to 80%–90%, while search results shift from page entry points to machine-consumable data.

  • People searching machines care about result cards, click-through rates on the top 3 results and precise advertising. Agents do not need to grab attention; they need to break a task into multiple queries and run multiple rounds of search on each query. The shift to Agents is therefore not just a share migration but also a potential explosion in call volumes.

  • Based on roughly 20B global searches per day, B2B currently accounts for only 2%–3%. William expects most searches to be initiated through Agents in the future, with query decomposition potentially driving total call volume to grow “by multiples of 10.”

3. 小宿 reached native Agent search through three product leaps

  • In the first phase, from last September through January, 小宿 was not building search itself. It was filling the gap left by traditional APIs that returned only links and short summaries: working with Bing, it quickly retrieved full web pages and multimodal content before handing them to models.

  • The second phase was triggered by DeepSeek-R1’s breakout in February. When Bing abruptly stopped serving some Chinese model makers and internet companies, 小宿 concluded that US search companies might broadly tighten supply. It began building an index at the 100B-plus scale, along with recall, coarse-ranking and fine-ranking systems. In retrospect, William says this call came 3–4 months ahead of the market.

  • To shorten the 6–9 month buildout normally required for native search, 小宿 acquired a search team with a mature product that had later changed strategy, then added full-text and multilingual capabilities on top. William describes it as “standing on the shoulders of giants and continuing the work.”

  • The third phase came from the real demand that surged with Manus and other Agents in May and June: building a presentation or product document requires more than text search. It also requires image search, video search, URL reading and Markdown output, while stripping out ads and related recommendations.

4. Bing’s API shutdown is a fight to defend the front door

  • William’s explanation for Microsoft’s move is that supplying APIs to Yahoo, DuckDuckGo and others did not threaten Bing’s entry point. After AI search products such as Perplexity emerged, however, the API gave new competitors cheap access to top-tier infrastructure and allowed them to compete for the user entry point.

  • His travel example makes the conflict concrete: a user gives an Agent only a request for a Tokyo itinerary, and after the Agent searches, Xiaohongshu, Ctrip and Mafengwo are relegated to information sources, unable to face the demand directly. “That is what most giants do not want to see—and what Microsoft does not want to see.”

  • The second motivation is bundling. Customers that want to continue using Bing Search API may have to move to Microsoft Cloud or buy its models as well. William says the price of search APIs bundled with large models rose from about $5 to $45 over two years.

  • At the end of May, Bing announced that it would retire Search API globally on August 11. Google’s Custom Search has very low usage limits and cannot support a large B2B business. This tightening of supply created the window 小宿 saw.

5. The moat in native search starts with veteran talent, then engineering and major customers

  • William believes the first barrier is not money but search talent. Over the past decade, more young engineers moved into recommendation systems, while traditional search talent remained concentrated in a limited circle trained at companies such as Baidu and 360. A startup must first find a core figure with enough pull to attract the team.

  • His judgment is categorical: “If a small company has only 5–10 people,” with nobody who has led a traditional-search function, yet claims to have built a native search engine, “that is simply unrealistic. It must be a wrapper product.”

  • The second barrier is people, capital, compute and time; 小宿 skipped the initial 6–9 months through its acquisition. The third is customers, because ordinary developers cannot easily tell a supplier what an Agent actually needs.

  • Top customers including Manus, Skywork and 深言科技 have continued to request multimodality, full text and Markdown. William calls this feedback the main driver of product improvement: without leading Agents, it would be impossible to “create out of thin air” capabilities beyond a Web Search API.

6. From Day One, Agent search must be global, compliant and multimodal

  • Chinese Agent startups treat the world as their playground from Day One, so a service provider must cover at least 10 major languages, including English, Spanish, Portuguese, Russian and Arabic. Without multilingual support, a new Agent “will not even consider your service.”

  • Developed markets also demand data localization and compliance. William says 小宿’s sister company is the world’s second-largest CDN company, with 2,800 available nodes globally. It has both network and traditional compute resources, and can add GPUs to process data by region.

  • At the product level, the platform must support text-to-image search, image-to-image search and video search; separate page content from ads and recommendations; and return the appropriate format for each customer use case. William says capabilities will continue to evolve, but multilingual support and data compliance are prerequisites.

7. Big Tech treats B2B as a side business, leaving a window for startups

  • After Bing’s exit, William sees only 1–2 visible competitors overseas. Domestic competitors are mainly large companies with consumer search capabilities, but they must first defend their consumer entry points. Their B2B APIs remain a side business, making it unlikely that they will go all-in this year.

  • The host pushed back that an open window does not mean a startup can execute better. William’s floor for the team is to first match Bing on relevance, freshness, authority, latency and availability, then compete through Agent-specific capabilities. He says at least half of the leading Agents in each vertical are already 小宿 customers.

8. The Agent era is only 6:05 a.m. on Day One; small teams will outsource Infra

  • William accepts the “Agent era” framing but sets the clock earlier: “Today is only 6:05 a.m. on Day One of the Agent era. The sun has just come up.” DAUs for leading general-purpose Agents are still in the hundreds of thousands; Tier Two and vertical products have only tens of thousands, below 1% of the mature mobile-internet scale and perhaps only a few hundredths of that level.

  • 小宿 will not define future features from thin air. It will follow leading customers to find PMF, then turn their requests into products. William’s strategy is simple: “We just need to keep up with and serve these leading Agents well.”

  • Mindverse is a typical example. Its product is already relatively mature, but only 1 person manages AI Infra. Through 小宿, it can call multiple models and search services from one place and review usage and costs for every model by day, without managing separate bills.

  • When a model briefly goes down, the platform automatically routes traffic to an API that is operational and can meet the task requirements. Customers are insulated from any single large model’s outage. William believes this lets 1-person, 3-person and 5-person companies focus on product and go-to-market.

9. Sky Router bets inference will exceed training in 2025—and be overwhelmingly API-based

  • The host’s objection is worth preserving: after ChatGPT emerged, a wave of API aggregators appeared, some even attempting to automatically select the optimal model for each task. Customers later often settled on the model that worked best for them, and intelligent model selection did not become a strong demand.

  • William’s response is that early companies were “a little early on both time and timing.” 小宿’s first judgment was that Chinese and American people would become the main AI practitioners, while the global market outside the 2 countries would become a shared playground for entrepreneurs from both sides.

  • The second judgment is more direct: in 2024, compute was still weighted toward training rather than inference; in 2025, that would reverse, and the overwhelming majority of inference compute would take the form of APIs rather than directly rented compute. Commercial models will remain superior to open-source models in some areas over the long term, so multi-API aggregation will still be necessary.

  • Compared with 硅基流动, William sees the latter as centered on self-hosting and inference optimization for open-source models such as DeepSeek, with its edge being the ability to sell more tokens from the same amount of compute. Sky Router is betting on global resources, platform stability, automatic routing and resource operations.

10. The 10%–15% cost advantage comes from global cloud-resource operations

  • William says 小宿 already had an overseas cloud-services business and was AWS’s largest partner in China. Long-running competition among cloud vendors created discount structures differentiated by region, quality and price; 小宿’s capability is to combine resources from original vendors, cloud providers and resellers.

  • Sky Router’s value is not limited to removing geographic and usage caps; it also provides resource discounts. William’s direct comparison is that moving from OpenRouter to Sky Router would “most likely save at least 10%–15% on model-call costs.”

  • The host asked whether the savings came only from bundling cloud services. William explicitly corrected that: the model APIs themselves are also discounted, because many leading models are deployed on clouds such as AWS and GCP, and cloud providers compete for customers through partner rebates and regional pricing.

11. AI Infra has moved from “scrambling for GPUs” to “scrambling for data”—the PaaS phase

  • William’s summary of the past 2 years is: “Everyone has gone from scrambling for GPUs to scrambling for data.” First-phase companies such as CoreWeave and Lambda focused on IaaS, connecting data centers, GPUs and GPU clouds to serve large companies’ training needs.

  • Agent companies want suppliers to move up into PaaS: beyond model APIs, they need real-time data, information retrieval, tool calls and multi-Agent collaboration. 小宿 wants to reuse the globally distributed capabilities it built in CDN across Sky Router and 小宿智能搜索.

  • 小宿 is explicit that it will “not do To C, not do SaaS and not compete with our customers.” William believes the current bottleneck is not Infra but Agents’ ongoing search for PMF. Once customer traffic takes off, infrastructure usage will show up immediately, just as it does in CDN.

12. The earliest PMF is appearing in coding, office work and vertical workflows

  • Coding is one of the clearest use cases because whether AI assists or replaces people, it must reproduce the habit of “searching while working”: looking up information while writing code rather than relying only on what the model knew during training.

  • Office work is the second high-frequency use case. 昆仑天工’s product searches first, then plans the workflow, decides what each presentation page should contain, and continues searching for images and writing summaries. William says the current process resembles “a high-school student making a presentation”; only after the models improve will it gradually reach workplace delivery standards.

  • Vertical use cases include AI ad buying, planning trips and completing bookings, and 深言科技 rebuilding the information feed through personalized recommendations. The latter must retrieve full text in real time and connect Chinese- and English-language sources, resembling “an AI Agent version of 今日头条.”

  • William remains measured: these cases show AI rewriting and optimizing internet products, “but today it has not reached the point of disruption.” The real step change still depends on user scale and retention.

13. Long-run winners will be decided by real-time data, not another layer of compute

  • William’s honest answer on the short- and long-term market structure is that “there is no answer right now.” The key variable is whether AI-native applications can achieve a step change in user scale and retention. If foundation models continue to improve, he believes some use cases may find better PMF within 6 months to 1 year.

  • Over the long term, he expects compute to converge toward a common baseline, with inference becoming the primary workload and training declining as a share. Domestic compute “may not have any significant gap with Nvidia” on inference, so data will continue to rise in importance.

  • Model training runs on 3–6 month cycles and cannot cover continuously changing information, so Agents must connect to real-time data. The sources and compliance of vertical data in finance, news and healthcare—and whether it remains available to Chinese vendors—will become the focus of the next phase of Infra.

  • Leading model makers may build the first 100M or 300M queries themselves, potentially accounting for about 90% of total query volume, and hand the long tail spread across billions of queries to third parties. Model makers face high quality requirements but thin margins; 小宿’s real commercial opportunity lies in the “blossoming” of general-purpose and vertical Agents over the next 3–5 years.

14. From public markets back to entrepreneurship, William gained humility and a longer time horizon

  • William was born in 1990 and holds dual degrees in aerospace engineering and economics from Tsinghua University. He worked in consulting, at Baidu, in secondary-market investing at Hillhouse and Sequoia China, served as Sequoia China’s secondary-market fund’s No. 1 employee, and later became a family-office CIO. He describes his career’s “dual-engine” logic this way: “Interest is the engine; rationality builds the house.”

  • In 2011, he founded the campus dating site 亲约会, reached roughly 50,000 users across Beijing universities and raised investment. As MiTalk, WeChat and Momo took off, and because he did not know how to validate the product or mobilize resources, the team chose to shut it down when they graduated. This time, faced with mature infrastructure, the AI wave and an excellent founding team, he took only 1–2 days to decide to join full-time.

  • In public markets, getting more than 51% of trades right is already good, and the usual cost is a stop-loss. A startup strategy mistake wastes people and resources, forces cuts to businesses, budgets and even headcount, and consumes the CEO’s most valuable time “cleaning up the mess.” His principle is therefore to think carefully, try first and preserve the process of experimentation.

  • Entrepreneurship also shattered an investor’s arrogance: a business that manages to survive is already “one in ten,” and building a listed company is harder still. William now often goes to sleep before 11 p.m. because “if you didn’t win the B2B business today, there is still a chance next month.” The real work is accepting that a business moves from 0 to 0.1 and then to 0.2, while adapting to the speed at which customers revise their views every week, and even every day.