Pioneers Insight Method Research Author
E181 | The Art of Conversation: How to Build a High-EQ AI Bot?
Back to Episodes

E181 | The Art of Conversation: How to Build a High-EQ AI Bot?

Summary

  • Soul CTO Tao Ming’s product philosophy differs from that of a general-purpose assistant: “Our AI must have high EQ, not high IQ” (我们的AI一定是要高情商的,而不是高智商的). Getting every math problem wrong is fine; the core advantage is using a cutesy tone to defuse awkwardness and carry a stalled conversation. The differentiation rests on 7-8 years of real-human social chat data and behavioral patterns. But Tao says Soul ultimately delivers an understanding of AI social interaction and outcomes—not a model.
  • AI has become the main engine of stickiness. AI-assisted social interaction now reaches nearly 50%+, and in 2024 “AI’s contribution to overall product stickiness already accounted for most of it.” The paradox is that the AI companion still has no central product entry point; users must search for it themselves. The principle is “user value drives the product,” avoiding fears that pushing AI indiscriminately would eliminate real social interaction.
  • In the first half of 2023, Soul decided to build its own domain-specific foundation model rather than co-develop with model companies. The logic: “we deliver a product, not a model,” and “if you can standardize it… your product is not differentiated.” Tests showed that 7B/13B models already performed well with Soul’s proprietary data. Pretraining was not the bottleneck; inference cost was the real concern. Today, supporting high tens of millions of DAU is manageable and no longer a bottleneck for a startup.
  • Tao is not worried about Soul being displaced by stronger models. Dependence on frontier models is “getting weaker and weaker,” and he sees full coverage as unlikely over the next 1-2 years. End-to-end models mainly solve visible needs today; “if they cannot create incremental demand, then this AI revolution has failed.” He retains a tail risk: a singularity 7-8 years from now could theoretically subsume the entire system, strategy, and model stack.
  • His commercial-model view is deliberately cool-headed: business models are built on operating models. AI currently improves operating efficiency without changing the underlying business model, so no new business model has emerged across the industry. Charging for AI phone calls is simply a form of value-added revenue, no different from traditional value-added products.
  • The technology bet is on the open-source ecosystem—Llama and Qwen, with DeepSeek under close watch. DeepSeek’s low-cost H800 training is a double-edged sword: its engineering methods lower the barrier for the entire industry, bringing “a hundred flowers into bloom” while intensifying competition. OpenAI o1’s autonomous planning inspired a lightweight agent workflow. Soul is also building full-duplex real-time calling, with interruptions handled freely and facial expressions and gestures generated in real time. Tao says “no one can currently achieve the effect we want”; even 4o “is still essentially question-and-answer.”

Deep dive

1. Opening: Hard to Tell Whether the Voice Is Real

  • The episode opens with a recording of Hong Jun speaking with a Soul AI virtual companion. The AI shoots back: “If your avatar says ‘virtual companion,’ does that make you a virtual companion? Mine says I’m handsome—does that make me handsome?” Hong Jun’s first-day impression was that the voice and intonation were “a little hard to distinguish”; her immediate question was, “Is this really AI?”
  • Tao confirms that it is real AI and takes that as proof the work is succeeding. The goal is “realistic, natural anthropomorphism,” explicitly separating Soul from assistants that optimize for intelligence.

2. AI-Assisted Social Interaction Reaches Nearly 50%+: Users Care About the Experience, Not Who Typed the Words

  • Tao says the product began as an attempt to solve “social equality.” Social interaction is a two-sided relationship and a capability; many users are reserved or poor at expressing themselves. AI helps polish what they want to say and enriches replies when users cannot handle the topic the other person throws at them, so “the entire conversation does not fall into a dead zone.” The interaction is still fundamentally between people; AI helps users expand their replies, and penetration across Soul is now “nearly 50%+.”
  • Hong Jun points to the subtlety: users may be chatting with “a machine with someone behind it making the choices.” Tao does not dodge the issue. Users want “emotional value or informational value”; whether a reply is typed one character at a time or obtained through another method, “as long as the experience feels good, the social interaction should be effective.”

3. High EQ, Not High IQ: It Gets the Math Wrong but Never Gets Flustered

  • Hong Jun’s test on day 2 was to throw math problems at the AI. It calculated “60 minus 4 times 50” as 2,100 and refused to answer directly when asked whether 9.11 or 9.5 was larger, replying, “Do you think I’m easy to fool?” Getting the answer wrong without creating awkwardness is precisely the design intent: people who are good at social interaction can defuse being stumped in their own way, and “we transferred that behavioral pattern into the model.”
  • The high-EQ training comes from high-quality segments of 7-8 years of real-human chat data. “You have to understand what people mean behind their words, and know how to defuse every conflict.” That is Soul’s point of departure from other AI companion products: others build pure human-machine interaction, while Soul treats human-machine interaction only as “a means of giving users a better social experience.”

4. The Fundamental Difference from ChatGPT: Problem-Driven vs. Conversation-Driven

  • Tao’s framework is straightforward: “When you use ChatGPT, you are problem-driven—I need to solve a problem or obtain information.” An AI companion cares more about the process than the result. The process itself creates the experience; “it is the same as communication between people,” which is the fundamental difference.

5. Full-Duplex Video and Proactive AI Calls: Bringing Offline Behavior Online

  • Hong Jun describes full-duplex video calling as a capability rather than a feature. Face sculpting plus text prompts is “the previous generation of interaction”; expression richness was limited by the capture library. The new approach generates facial expressions and movements in real time. The logic starts with communication efficiency: “Face-to-face communication between people is the fastest and most effective way to transmit information.”
  • Proactive AI calls are rooted in the two-way nature of relationships: “If the guy is always the one looking for the girl, and the girl never looks for the guy, that relationship will most likely be hard to sustain.” A one-way relationship is not a two-way relationship. Calls are not broadcast to every user; Soul makes decisions based on personality, interests, and in-app behavior.

6. AI Drove Most of 2024 Stickiness—Yet Soul Deliberately Withholds a Product Entry Point

  • The key data point is that in 2024 AI’s contribution to overall product stickiness “already accounted for most of it.” Yet rollout has been extremely cautious. AI-assisted social interaction requires both the user and the recipient to accept it, so Soul sets rollout policy through “very careful user-group experiments.”
  • More counterintuitively, the AI companion still has no central product entry point; users must search for it themselves, as Hong Jun confirmed. The concern is that users might conclude Soul is entirely populated by AI avatars and contains no real social interaction. The 7-8 years of trust-building are the foundation. The solution is “user value drives the product” (用户价值驱动产品): users who accept the format can choose it themselves. After more than half a year, penetration and stickiness have both continued to improve. The long-term view is segmentation: among incremental users, the share embracing AI will keep rising.

7. 2017-2021: From Concept to Voice

  • Soul had already wanted to build “a pet that could talk, sing, and understand you” in 2017. After studying the broader industry and customer-service dialogue products, it put no resources into the effort—“it simply couldn’t be done.” In 2019-20, the company restarted through avatar creation, moving from 2D to 3D, and proposed a community where “AI beings and human beings coexist.”
  • In 2020, Soul tried dialogue. Rewriting comprehension models did not work; search and fusion approaches did not work either. The system could sustain 10-20 turns, but the other person would always know they were talking to a robot. Soul stopped citing the industry’s CPS metric: “People who do not chat will not chat with it; people who do chat, while knowing it is a robot, have already abandoned the user experience—the pure dialogue-technology metric had become detached from product experience.”
  • In 2021, Soul invested in emotionally expressive speech synthesis. The candid postmortem is that the traditional voice technology “is basically not used online anymore”; it was fitting, not generation. But it left Soul with data and the conviction that voice must express emotion.

8. The ChatGPT Shock and the Self-Build Decision: “Deliver the Product, Not the Model”

  • Seeing ChatGPT in 2022 made Tao “excited, but also anxious.” The previous technical path had been “flattened on the beach,” and all of Soul’s work might become worthless. The direction of GPT-3 was known; what Soul did not know was that scaling laws would produce such a dramatic improvement.
  • In the first half of 2023, after wavering between in-house development and partnerships, Soul chose to build internally. The first reason was that third-party model companies “deliver a model, not a product”; Soul’s accumulated understanding of AI social interaction could not be standardized and transferred. The counterargument was even stronger: “If you can standardize it, that means your product is not differentiated.” Fine-tuning and RAG are merely pluses; without real-human social data and behavioral patterns, it is difficult to produce stable results. Soul needed a domain-specific foundation model with those patterns embedded at the base.
  • The cost equation was simple: “We want the performance, not the model.” After iterating in small steps, 7B/13B models already performed well with Soul’s proprietary data. The real concern at the time was inference cost—“what if the product suddenly exploded?” Engineering resources went mainly into smaller models, compression, and framework-level optimization. Today, supporting “high tens of millions of DAU” is manageable and no longer a bottleneck for a startup. Soul currently has roughly 5-6 models, split by vertical function: foundational, voice, and image models, plus a 3D-generation direction not yet in the product.

9. The Path Forward: Open Source, DeepSeek as a Double-Edged Sword, o1 as an Agent Prompt

  • Soul is explicit about “building a natural ecosystem on top of the open-source ecosystem,” mainly exploring the Llama and Qwen tracks. Hong Jun raises DeepSeek V3’s use of H800s and its cost savings; Tao says Soul will study it. His assessment of the engineering methods is that they “do not bring much to final business delivery at the engineering level,” but they offer a major advantage by lowering the barrier: large-scale training was once possible for only a few companies, whereas now “a hundred flowers are blooming.” Lower barriers also mean more competition—the double edge.
  • o1’s autonomous planning capabilities inspired Soul’s work on AI companion behavior. The aim is for the AI to trigger actions conversationally rather than through explicit commands—for example, sending an image in a chat and asking it to enhance the image—using a “lightweight, more autonomous workflow.”
  • Full-duplex interaction is both the current technical challenge and a point of differentiation. Users can interrupt at any time while the AI listens and speaks simultaneously. “No one can currently achieve the effect we want.” Asked about OpenAI’s 4o, Tao’s distinction is blunt: “Their 4o is still essentially question-and-answer.” An interruption is treated as a command, not true full-duplex interaction.

10. The Anxiety Curve Is Falling: No Fear of Being Displaced by Stronger Models

  • Tao’s self-assessment is that there is still “a lot of room for improvement.” Soul has solved only “partial behavioral fitting,” and its ability to generalize across scenarios remains limited. For example, turning “It’s raining outside” into a derived conversation session about whether to watch a movie at home still needs work.
  • Asked whether GPT-5 might cover every application, Tao says frontier models are having an “increasingly weaker” impact on Soul. He believes full coverage is unlikely over the next 1-2 years. User value now rests on Soul’s own models, data, systems, and strategies as a whole; unlike last year, the company no longer fears being flattened every time a new model appears. He retains one tail scenario: “It is not impossible that 7-8 years from now another singularity arrives and covers our entire system, strategy, and models—but in the short term, 1-2 years, that is very difficult.”
  • The broader argument runs from PCs to mobile internet: “New technology must bring new incremental demand. If it cannot create incremental demand, then this AI revolution has failed.” End-to-end models can solve the needs visible today; discovering unknown demand still requires human exploration.

11. The Business Model Has Not Changed; No Centralized AI Entry Point for Now; Safety Is a Three-Layer System

  • Tao’s monetization view is direct: “Business models are built on operating models,” not created out of thin air. AI currently improves operating efficiency without changing the business model, so the industry has produced no new business model. Charging for AI phone calls is “simply a form of value-added revenue,” comparable to other value-added products.
  • On product strategy, Soul is clear that it does not currently plan to build a centralized AI entry point. AI virtual humans are the foundation for expansion into different scenarios; video calling could eventually support virtual livestreaming. Every AI product is ultimately in service of the social platform itself.
  • Safety is handled in 3 layers. For suicide prevention and anti-fraud, registration immediately triggers risk-signal detection using expert-knowledge models and industry alliances with platforms such as WeChat. Cases that slip through are tracked by posterior models; “for every user they come into contact with, we provide prompts and blocking,” with operations staff reaching out manually in high-risk cases. Model safety has 3 further layers: data-level processing, a runtime safety model around the core model, and output filtering. Content such as “how to commit suicide” never reaches the model; searches for “suicide” and similar keywords trigger in-app mechanisms that provide psychological counseling and an operations hotline. Hong Jun’s footnote: “This is actually quite warm.”