Mizzen AI’s 孙克强 & 李一豪 on AI’s User Research Opportunity
Summary
- Mizzen AI is not betting on a “cheaper interview tool”; it is targeting the key link in the AI product loop that has yet to accelerate. 孙克强’s own words: “AI has dramatically accelerated product iteration, but user research is lagging far behind R&D.” A traditional project takes roughly 60 person-days and costs about 100,000; Mizzen Insight claims to deliver 100x the speed at one-tenth the cost, compressing the research cycle from months to hours—ask in the morning, get the report after lunch. The team currently has just 8 people. Traditional user research is labor- and knowledge-intensive; adding headcount increases management overhead, while ROI declines when delivery quality stays unchanged.
- 李一豪 sees a large aggregate market with no dominant player, entrenched structures, and extreme fragmentation—a clear window for substitution. He estimates the global traditional research market at roughly $80B-$90B and emerging UX tools at $3B-$5B, while even the leading companies generate only $2B-$3B in annual revenue. His comparison is the $400B traditional legal market and $60B-$70B legal-tech market: fragmentation does not prevent the No. 1 player from going public and generating substantial returns. More importantly, these traditional, semi-digital industries change slowly, giving new entrants room to build an AI-native advantage.
- The real product moat is not replacing interviewers with ChatGPT, but training a questioning model that can get to the crux. 孙克强 concedes that general-purpose models are “very good answerers, not good questioners,” while most high-quality interview processes remain closed-source. The team is securing exclusive interview data that has passed its confidentiality period and using reinforcement learning to build its own environment and reward model, evaluating moderator questions through dimensions including answer authenticity, detail, subjective analysis, and consistency. 李一豪 sees this as the core moat of vertical AI: domain benchmarks and an edge-data flywheel, rather than simply calling a stronger base model.
- The biggest value of lower costs may not be taking existing large projects, but turning user research into small experiments that can be launched every day. 孙克强 observes that product-building costs have “fallen dramatically, even below distribution costs,” pushing customers from one large report a month toward “continuous, incremental user research.” Consumer-electronics customers are already testing small questions such as whether to adopt “try before you pay” and whether handwriting and screen typing should be separated. The platform charges by project with no platform fee, with the goal of “democratizing user research”; Koji believes new-market opportunities are more startup-friendly.
- The counterintuitive advantage of an AI moderator is reducing interpersonal performance, but response quality still depends on mechanism design. 孙克强 acknowledges that respondents initially feel uncomfortable, but argues that AI reduces interference from a moderator’s gender, emotions, and chemistry. The platform discloses its scoring criteria—authenticity, detail, subjective analysis, and consistency—in advance and links scores to incentives, making “more truthful, more willing expression” the optimal strategy. Its respondent supply comes from 2 domestic and 2 overseas vendors plus partner expert pools, with total capacity claimed to be in the tens of millions.
- The team explicitly rejects the shortcut of “interviewing AI” for now, while preserving a long-term path toward respondent simulation once enough real data has been accumulated. 孙克强 says the gap between current models and real human context remains “huge”; fixed training contexts can reproduce only past and incomplete information, while general reasoning and brand trust are still insufficient. The long-term plan is not to generate personas from scratch, but to distill a respondent’s past statements into a preference graph and run evidence-based, low-cost experimental simulations.
- This early-stage bet is simultaneously on a long-term technical thesis and the founder’s ability to pivot. 孙克强 moved from visual preference research and HPS through GPT-3.5 and RLHF to multimodal human preference, then abandoned the art-market direction after studying galleries, artists, and traders. What 李一豪 values is that “real sampling,” rapid pivots, and resilience are strengthened rather than crushed by negative feedback. The endgame is for Mizzen to become “Iron Man armor for every PM,” with Web coding and web user research as its two hands: humans contribute intuition, aesthetics, and “strokes of genius,” while AI completes the research-to-iteration loop. Asked how he would invest $3M across 3 people, 孙克强 said the first 2 checks would go to Quick Stone’s 2 longtime shareholders, and the third to Silicon Valley company Krea.
Deep dive
1. The faster AI R&D moves, the more user research becomes the hard bottleneck in the product loop
孙克强 uses 3 questions to screen directions: Is the demand real? Are the technical variables sufficiently disruptive? Does it lead “toward greatness”? Nielsen’s history of more than 100 years shows that the need has persisted, yet the industry remains “very slow, very expensive, and highly inefficient.”
When exploring ideas internally, the team could interview only a little over 10 people a week; traditional projects typically require about 60 person-days and roughly 100,000 in project spend. AI can run the 20 to 30 interviews that were previously conducted sequentially in parallel, leading Mizzen Insight to claim “100x speed, one-tenth the cost.” The team currently has just 8 people.
The broader thesis is that products must start with user insight, return to users after R&D, and keep every link moving; any lagging link limits iteration. 李一豪 adds that a model’s implicit understanding of feedback could ultimately affect interaction, experience, performance, and even pricing.
2. A near-$100B slow, expensive, fragmented market leaves room for a new species to consolidate
李一豪 estimates the global traditional research market at roughly $80B-$90B and emerging UX tools at $3B-$5B. Although the industry has entered a phase of structural entrenchment, M&A, and consolidation, even the leading companies generate only about $2B-$3B in annual revenue.
His comparison is the roughly $400B traditional legal market, where legal tech has already reached $60B-$70B. Before China’s “double reduction” policy, K-12 education grew from roughly RMB260B to RMB460B, approaching RMB500B; even the top 3 companies could each reach RMB10B-RMB15B in annual revenue.
These markets are inherently fragmented: specialization and deep service allow niche companies to persist, but the enormous total market can still support a No. 1 player large enough to go public. AI may expand the leaders—or create an entirely different “new species.”
李一豪 argues that in deeply digitized markets, ecosystems are often controlled by internet giants. Traditional research evolves more slowly, creating an opening for startups. He sees Chinese teams’ combination of product iteration, user operations, and complex To B delivery as highly competitive.
3. The “alignment tax” of human moderators means AI is doing more than saving labor
孙克强 found that running multiple moderators in parallel creates a costly “alignment tax”: insights are diluted during information synchronization, while human fatigue and mood on a given day directly change interview quality.
Even when the same person conducts every interview, “cognitive marginal returns” begin to decline. The first respondent is exciting; by the second, the moderator’s interest and density of follow-up questions can fall sharply.
AI’s value is therefore not merely replacing hours of labor, but preserving consistency—treating every respondent as if it were the first interview, maintaining curiosity, and probing details that human moderators might overlook.
4. User research shifts from big projects to continuous small experiments; greenfield demand is the real battleground
孙克强’s contrarian view is that AI has made product-building costs “fall dramatically, even below distribution costs.” One-off, expensive research reports that take a month are consequently shifting toward “continuous, incremental user research.”
Koji uses consumer products to illustrate the value. A bad software requirement can still be rolled back; if 10,000 physical products are stocked and fail to sell, the company is left with inventory and must also pay to destroy it. But traditional research’s researcher-hour rates were so high that it was generally affordable only to well-funded companies such as Dyson and Nio.
Consumer-electronics customers have expanded from core projects to small questions: whether consumers are willing to “try before they pay,” and whether handwriting input and screen typing should appear in the same product form. 孙克强 says the goal is to ensure user research is no longer “a right reserved for large companies, or only the core projects of large companies.”
The platform already has paying customers. It charges by project with no platform fee. The host restated the pricing logic as “interview 10 people, at about RMB1,000 per person, for a project costing roughly RMB10,000”; 孙克强 answered “correct,” while disclosing no standardized price list.
5. Mizzen breaks one research project into 4 agents while preserving an evidence trail
Using a large made-to-order food brand as an example, the interview-creation agent receives a natural-language requirement and organizes the interview guide in real time as the conversation proceeds.
Recruitment requires only a description of the target user profile. The platform then screens candidates and uses a 5-to-10-minute interview agent to verify identity and eligibility, ensuring respondents match the brand’s requirements.
The AI moderator asks respondents to turn on their cameras, takes them directly back to specific situations, and follows up. The report agent then aggregates all interviews and produces overall conclusions plus qualitative and quantitative analysis for each question. 孙克强 emphasizes that every report finding can be supported by source transcripts, rather than delivered as a single summary that cannot be audited.
6. AI may reduce social performance, but response quality depends on incentive design
The moderator’s challenge is straightforward: once respondents realize that the other side is AI, will they become perfunctory, less emotional, or unable to maintain attention? 孙克强 admits there is “a little of that at first,” but says product design has already addressed the issue.
The first layer is making respondents understand that they are helping a real brand improve its product. The second invokes the “Hawthorne effect”: a human moderator’s gender, emotions, and chemistry can trigger self-presentation, while AI may help people “drop their guard” and speak more truthfully.
The third is a benchmark that uses multimodal behavior to assess authenticity and checks detail, subjective analysis, and consistency over time. Respondents are told the scoring criteria before they begin, and scores are linked to their incentive payments. “The best solution is for you to become more truthful and more willing to express yourself.”
The underlying sample pool is connected to 2 domestic and 2 overseas suppliers, with combined capacity claimed to be in the tens of millions. The platform then uses agents for screening and interviews, while also connecting to expert pools at consulting and research firms and building its own community.
7. “Interviewing AI” is not credible today; real people remain the data source
Asked why it does not simply interview AI simulations of people, 孙克强 gives a clear rejection: the gap between current models and real human context remains “huge,” and even small changes in the environment can alter a person’s attitude and choices.
Fixed training contexts provide only past and incomplete information, while general intelligence is not yet sufficient to reliably infer real product experiences. The more immediate obstacle is that brands remain deeply concerned about the authenticity and validity of AI respondents.
The team still preserves a long-term path: accumulate an individual’s expressions and preferences through real interviews, then simulate specific respondents for efficient, real-time, low-cost experimental research. The prerequisite is not a generic persona, but verifiable historical data.
8. General models answer well but ask poorly; reinforcement learning is the moderator moat
Clients independently converge on the view that research quality first depends on what kinds of questions a moderator can ask. 孙克强’s candid assessment is that current large models are “very good answerers, not good questioners.”
One reason is that model companies’ post-training primarily optimizes answer preferences rather than how to ask a question that gets to the heart of the matter. Another is that much of the high-quality interview work done by companies such as Apple is not public, making these critical samples difficult to find online.
Mizzen has partnered with multiple domestic and overseas research and consulting institutions to obtain interview data that has passed its confidentiality period, signing exclusive agreements with them. The team also spent half a year designing its training system. The goal is not a model that can merely “talk,” but a professional AI moderator that can get to the crux.
Its reinforcement-learning environment simulates responses using real project contexts and respondent information. A reward model then evaluates simulated respondents’ answers across multiple dimensions; the better the answer, the better the question. 李一豪 believes this domain benchmark and the continuing generation of edge data form the long-term flywheel of a vertical AI company.
9. Preference graphs turn one-off interviews into a long-term asset
孙克强 argues that traditional personalization and recommendation rely on discrete, crude labels that cannot represent a person in 3 dimensions. Language can carry far richer information; what was missing was a low-cost way to integrate and retrieve it.
Mizzen’s “human preference graph” attempts to aggregate expressions users have left across different platforms and retrieve the historical preferences most relevant to the current question. Its first use is to move beyond limited labels and identify people who more precisely match research requirements.
The second use is cross-interview validation: if a respondent contradicts an earlier statement, the moderator actively confirms it. Only the third use is simulating a respondent’s views. The long-term asset, 孙克强 emphasizes, is converting real human preferences into reusable, structured data rather than letting each project end with the data effectively reset to zero.
10. The founder’s pivots have all revolved around understanding people; the endgame is Iron Man armor for every PM
孙克强 is 30, holds a master’s degree from Tsinghua and a PhD from CUHK’s MMLab, and previously spent 5 years at SenseTime and 6 months at Meta. His entrepreneurial credo is “YOLO, You Only Live Once”: “I will eventually die; before that, I want to leave something behind in this world.”
He sees GPT-3.5’s use of RLHF to show the world that “human preference is a model amplifier” as a formative insight. During his PhD, he also worked on HPS, a visual-preference model he describes as the world’s first systematic attempt to extract differences in human visual preferences. His own explanation is: “This wasn’t jumping from computer vision to user research”; it was expanding from visual preference into the more complex, multimodal motivations and reasons behind human choices.
The team initially explored human-preference benchmarks and arenas in the art world. After studying galleries, artists, and traders, 孙克强 concluded that art supply was no longer scarce; channels, capital, and distribution controlled value. The team therefore pivoted quickly. From their first rainy-day meeting, planned for 90 minutes but lasting 3 hours, 李一豪 says he recognized the founder’s “real sampling,” resilience, and instinct to keep widening the boundary.
Musical.ly founder Louis once told 孙克强 that the “productivity leap” created by the spread of front- and rear-facing cameras transformed both creation and consumption. 孙克强 wants Mizzen to unlock user-research capacity in the same way. The endgame is “Iron Man armor for every PM”: Web coding and web user research become the two hands, humans contribute intuition, extreme aesthetic standards, and strokes of genius, while AI completes the chain from research to product iteration.
Asked how he would invest $3M across 3 people, 孙克强 said the first 2 checks would go to Quick Stone’s 2 longtime shareholders, and the third to Krea, a Silicon Valley company, because it had polished product-development speed to an extreme at a very early stage.