Where Do Video Agents Go After Seedance? OiiOii's 闹闹 on Video Models
Where Do Video Agents Go After Seedance? OiiOii's 闹闹 on Video Models
Summary
- 闹闹’s core product thesis is that multiple reference images plus natural-language description better match how creators think, but video models have not fully converged on that paradigm; Seedance is “basically an upgraded Sora 2.” After Sora 2 launched, she removed every first/last-frame model from her product and now believes that was the right call: all of her earlier prompt-engineering work carried over to Seedance, which supports true multi-reference inputs and natural shot transitions; chasing ever-longer videos “doesn’t make sense,” as long takes can be little more than a flex.
- “Video models are simply not a startup business.” The bottleneck is not just algorithms but the organizational machinery for data labeling: with tens of thousands of annotators and standards that must be set and iterated, incumbents can get to at least 80-90 points while startups manage only 20-30; that is why Seedance leads in China, with Kling next, and why short-video platforms inherited their edge from special-effects AI Lab and Kuaishou Y-tech creative-algorithm teams.
- China’s short-video ecosystem is the world’s leading breeding ground for models, and Sora’s missed opportunity was the ecosystem, not just the technology. “OpenAI had Sora 2 but failed to build and run the ecosystem around it; if it had launched an API first instead of that app, things might have worked out differently.” The video-model race is not over: she expects at least 3 players, and says incumbents have an advantage only when a project is not the first to be abandoned and continues to receive attention.
- The case for the Agent layer is straightforward: models deliver only 30-40 points of the finished capability, while an Agent can bring the application and product close to passing—and the layer itself is a black box. Even coding, as powerful as it is, is not something that can truly be delivered as-is; Midjourney’s survival against larger models shows that multimodal product layers need not be winner-take-all. The Agent’s job is to “understand how to get the best out of a model,” using routing and stage-by-stage benchmarks, with model-product, algorithm, and quality-design roles working together.
- Seedance has changed the cost curve, driving UGC away and pushing the market toward Pro C and small B, but she does not think foundation models can keep raising prices in a video ecosystem. “If raising prices means more people can no longer make money, it is unsustainable”; short-drama margins are already thin, and ByteDance cares more about the vitality of UGC creation on Douyin—cutting prices may actually maximize value. Consumer access cannot remain free against inference costs: “Even ByteDance-sized, Doubao has started charging,” because AI generation is an “unlimited cost” and may cost more than buying equipment to shoot video.
- Her verdict on the current video-Agent boom is blunt: “I think it’s a delusion.” Only 2 or 3 products may be genuinely solid; many are wrappers or routing stations cashing in on the Agent buzzword before disappearing. Every product markets itself as a 100, even though some are at 30-40 and others 70-80—“none is a 100, including us”—and users have no shortcut: they have to try them.
- The hard multimodal problem is “structuring the subjective,” not porting the language-model playbook directly into vision. Language is explicit and standardized, which is why coding got traction first; visual expression is implicit. Subjective quality “cannot be quantified, but it leaves traces, like a color band,” so benchmarks vary by person and product layers will necessarily “bloom,” rather than collapse into winner-take-all.
- Her restraint under investor questioning is notable: she assigns only a 70-80% probability that the Agent layer will avoid being absorbed by the model, and calls that objective rather than conservative. “People speak with far too much certainty; ‘the wolf is coming’ has been cried too many times.” Seedance 2.0 taught her less about technology than posture: “not flashy at all, not sci-fi at all—just honestly and solidly doing what needed to be done.” What AI startups lack most are people willing to sweat the details and do the dirty work.
Deep dive
1. The “Can and Can’t” of a Female Founder: Not an Epiphany, but Creating the Environment
- 卫诗婕 opened with the question, “When did you first start feeling that you could do it too?” 闹闹 sidestepped the motivational narrative. She sees her temperament as gender-neutral and herself as an “observer” since childhood, with a habit of self-reflection. There was no single moment when she decided she could or could not do something; instead, she asks whether a project requires her to create an environment in which it can succeed—or in which its odds of success are materially higher.
- The time marker is this: the thought began around this time last year, while her first startup dates back to 2014.
2. Three Years at Tencent and the First Startup in 2014: “Illusory Bigness”
- After 3 years in product at Tencent, she left to start a company, prompted by extreme-sports films from the Banff Mountain Film Festival. At the time, “spending a year doing product made you a veteran”; with an employee number in the tens of thousands, she was already saying, “Tencent has more than 10,000 people now—there probably isn’t much room left, right?”
- Her retrospective on that first startup is unsparing: excitement outweighed analysis. “I had never taken a specific business from small to large, so I had no way to make a reasonable projection… That bigness was an imaginary bigness. Looking back, it was just wishful thinking.”
- She does not romanticize what 3 years at Tencent gave her: “It planted a seed, but it never sprouted. I don’t think it actually improved me.” Part of the entrepreneurial impulse came from watching the seasoned lives of team leads and directors and realizing they were not the lives she wanted.
3. A Firmer Second Startup; Zhang Xiaolong’s “View from Outside Humanity”
- OiiOii is different because she feels she has a base to stand on. In the worst case, the company can cover its own costs and gradually build the product; she also grew up in healthy competition and received enough positive feedback to understand the ingredients that determine whether something can work.
- Her most profound experience at WeChat was 张小龙’s 8-hour internal product talk. “He was observing humanity from a perspective outside humanity, then building something for humans to use—putting so many of human nature’s strengths and weaknesses into an amplifier to let them ferment, and then pruning them like a gardener.”
- Asked what she would most like to ask 张小龙, she said there was nothing. “Everything has to be done by you personally.”
4. What She Wanted to Preserve from the First Startup: A Chance for Ordinary People to Express Themselves Through Content
- The goal she had 10 years ago—to give ordinary people the ability to express themselves through content—was eventually realized by Douyin. “That means the underlying judgment was not very far off. My ability was limited, the timing was wrong, and I simply wasn’t the person who made it happen.”
- Why is she so fixated on the idea? Her self-portrait is that of “a kind of soil-type person”: seeing other people take root, grow, and flourish on top of her gives her more satisfaction than achieving the same thing herself.
5. Tencent vs. ByteDance’s Product Philosophies: Opposites, Then One
- She compares WeChat’s and ByteDance’s product philosophies to the left and right hemispheres of the brain. “You see many things in this world that appear to be opposites, but they are actually one. Before you realize they are one, you assume they are in opposition.”
- She also flags the darker side of data use: “Knowing exactly how to make an experiment’s data turn positive and doing it for the sake of a positive result does not mean the thing will remain positive over the long term.” It may be considered correct in one environment and wrong in another.
- The product standard after combining the two approaches is to observe users while remaining sensitive to data. “Data becomes credible only after it reaches a certain scale; below that, it is biased.” Content products must also find ways to classify and quantify subjective judgments of quality. “That is hard.”
6. Few Wow Moments in AI: Marketing Power ≠ Deliverable Power
- A veteran product person’s candid assessment: AI products have produced very few moments that truly made her say wow. “The force of the marketing makes everyone feel that it is extremely powerful, but when you actually use it, a lot of things cannot be completed. Its delivery capability is actually very weak.”
- 卫诗婕 asked whether that gap was not precisely where the product opportunity lay. 闹闹 agreed, but stressed that the industry is still in the early phase of figuring out how to convert technology into a sensible product. Fine-grained product differences may be erased by major model iterations, “but that does not mean they have no value.”
7. The Product People AI-Native Companies Need: Five Roles and the Scarce “Product Architecture”
- At her current product scale, she needs only 5 product roles: user experience, language models, multimodality, commercialization, and data strategy. But AI products require each person to understand the others’ domains. “If you don’t understand what the others understand, the cost of conversation becomes very high.”
- “Product architecture” is an important and scarce role. From day one, someone must know where the product should go in 2-3 months, reserve feature modules, and distinguish what model iteration will absorb from what it will not. “Many product people who understand models extremely well are very young and have not yet lived through many product changes and iterations. That is also why so many products now look alike, copying one another.”
8. The Pivotal Route Call: From First/Last Frames to Multi-Reference Inputs, Seedance Validates the Thesis
- Before Sora 2, the industry was built around first-and-last-frame control: defining a tight range for the opening and closing frames. Vidu was the first to introduce the concept of multiple reference images, which she found closer to how creators actually think. “They don’t have a specific picture in their head. They need the model to fill in and complete their imagination.” The elements are characters, props, scenes, and environments.
- Sora 2 is effectively a “single-reference” model: one image paired with a much more detailed text description. The image does not necessarily define the first frame; the model infers the scene from the combination of image and text. “That means your model has sufficiently rich imagination.” She consequently removed all first/last-frame models and kept only Sora 2 and parts of Vidu. “That decision was basically right.”
- The implementation validated the call: “Seedance is basically an upgraded Sora 2. All the PE work we previously did to optimize models applies to Seedance.” It is also genuinely multi-reference, with stronger controllability.
- Her practical advice is specific. For a character, upload “a slightly three-quarter portrait plus a headshot”; a three-view sheet can instead be misread as 3 different people. Then use positive prompts to specify a live-action human and negative prompts to exclude animation. “Under those 3 constraints, it will probably work, but there is no guarantee.” She also suspects Seedance is weaker than Sora 2 at preserving style because “the proportion of live-action data is too high, while stylized data is sparse and has been diluted.” “The algorithmic approaches are broadly similar, but data is extremely, extremely important.”
9. Models Diverge: Kling’s Cinematic Polish vs. Narrative Cutting
- Each company’s data strategy produces a different character. Kling focuses on professional film and television, with fine-grained output in science-fiction and advertising. “But it is less refined on narrative films, because its shot cutting is weaker.” Advertising wants close-ups, camera movement, and visual impact; narrative requires shot-reverse-shot structure and an understanding of character relationships.
- She criticized the early Sora-style fixation on long tracking shots at the time. “A story does not need that many long takes. Sometimes a long take is just showing off… Chasing an extremely long video for its own sake does not make sense.” Following one person continuously from a third-person view “looks like a documentary, but it gets boring.” Narrative shots work like psychological cues: to convey a crime boss, “you need a dark, high-angle shot so the audience can feel his menace.”
- Sora 2 initially took off for 2 reasons: synchronized sound effects and “very natural shot cutting,” both aligned with expectations built by years of watching filmed content. Her methodology is simple: “A person is like a model. If you consume enough content, you can feel whether something will go viral—you have been feeding yourself a lot of data.”
10. “Video Models Are Simply Not a Startup Business”
- She thinks the primary market’s reluctance to fund a “China Sora” was justified. The core issue is the organizational capacity required for labeling. Data has to be purchased; “the people labeling the data number in the tens of thousands.” Who sets the standard, how do you make sure everyone interprets it consistently, and how do you control the quality of outsourced work? Startups struggle to solve those problems.
- Large companies already have experience in graphics and imaging algorithms. Even if their initial results are inaccurate, they can iterate toward accuracy with massive data volumes. “They can get the labeling standards and the labeling itself to at least 80-90 points. A startup can only get to 20-30.”
- Why could Kuaishou produce Kling? Its video-understanding algorithm team was “extremely strong, even stronger than ByteDance at the time,” at least in 2021, with good output per employee. “None of these things appeared from nowhere. They were built on past accumulation or a foundation, and that is a real moat.” Kling emerged first through professional film and television data; ByteDance “did not favor any one type of data,” and its enormous user base pushed it toward a general-purpose foundation model—though that general model was not ready at the time.
11. Ecosystem Lead and Data Lead: Sora’s Missed Opportunity
- She rejects the equation of technical leadership with video leadership. “Video is an ecosystem.” Many of the trends now popular overseas had already taken off in China years ago, and ByteDance built that ecosystem into the world’s most advanced. “To launch a video model, you need an ecosystem as soil for it to flourish.”
- Her verdict follows directly: “It is a shame that OpenAI had Sora 2 but did not build and operate the ecosystem around it. If it had launched an API first instead of that app, it might have worked out differently. But it was still a relatively great model.”
- There is another layer of leadership: data. Short-video companies already know how to use video data in recommendation systems, and those methods can transfer to model training. Her example of usable data: have someone with 3D-modeling experience annotate a video, and professional vocabulary enters the labels. When users type those domain terms, the model can respond. “The more specialized and granular the domain coverage, the more precise and detailed the response.”
12. Why Short-Video Platforms Win: Inherited Creative-Algorithm Talent
- 卫诗婕 asked whether it was a coincidence that the best video-generation models came from ByteDance and Kuaishou rather than YouTube, Tencent, or iQIYI, all of which are major video companies. 闹闹’s answer was no. “AI capability itself sits on the creation side, not the playback side.”
- The lineage is clear. Douyin’s special-effects AI Lab once used GANs for aging, childification, and cartoon effects. “The core people behind Seedance were originally creators from AI Lab, while Kling’s core algorithm team came from Kuaishou Y-tech.” Inside these companies’ comfort zones, “it was simply a technology upgrade.” The same applies to products: “I don’t deny that there is something new, but the methods and approaches are inherited. They did not suddenly appear from nowhere.”
13. The Three-Way Trade-Off and What Comes Next: Editing Iteration and Multimodal Harnesses
- Video models face a 3-way trade-off among quality, generation time, and generation cost. She sees 2 next directions: continued quality improvements—current systems work mainly for professional and semi-professional users, while “an ordinary user cannot produce content with any information value using one sentence”—and iteration in editing. Today, changing one local element requires regenerating the entire sequence, losing both editing precision and cost efficiency.
- Why can’t models even modify a local detail in an image? Diffusion first diverges and then converges, which makes hallucination intrinsic; it cannot constrain itself to a specific frame. Traditional algorithms will have to be integrated. “Almost everyone needs this. It is an obvious pain point, and the model has to absorb it.” She believes that will eventually happen.
- Language-model harnesses, or constraint layers, “will definitely” appear in multimodal and video systems. But content has no single standard answer. The field will have to split into many categories and gradually converge within each one.
14. Structuring the Subjective: The Color-Band Metaphor and Multimodality’s Core Difficulty
- This echoes the previous guest’s view that language is inherently structured, while visual expression varies widely and is difficult to align or converge. 闹闹 cautions against directly comparing multimodality with language models: “Language is too explicit and direct. Multimodality is a form of expression, something very abstract—camera techniques rely on implication. It is doing the work of another human brain, the subjective work. Using rational methods to handle subjective things is extremely difficult.”
- 卫诗婕 asked how something that cannot be quantified can be evaluated. 闹闹 offered the metaphor worth preserving: “Subjective feeling cannot be quantified, but it leaves traces, like a color band—red, orange, yellow, green, cyan, blue, violet. From a distance they look separate; when you zoom in to each pixel, they seem to have boundaries and yet no boundaries. So subjectivity can be structured, but the granularity of the structure is coarser.”
- The implication goes straight to market structure: each company’s benchmark is different and highly dependent on the people involved. A good commercial film has common features that can be summarized; whether an art film is good is “highly personal.” That is the underlying reason the product layer will not be winner-take-all.
15. From Jianying and Douyin to Bilibili: Why She Ultimately Had to Create Her Own Environment
- After her first startup, she defined herself as a “senior employee,” wanting to learn how to see from both the layer above and the layer below. She also learned what kind of leader not to become: first, the leader who thinks they are too capable, becomes blind, and cannot hear input that would help the work; second, the perpetual nice person. “A leader has to get things done, and getting things done means making trade-offs among interests.”
- Why leave after becoming head of Jianying? “It was already big enough—so big that it had to take care of the entire user base, with AI possibly serving only an auxiliary role.” She wanted to build something that could start small, grow, and pass through an entire lifecycle. During her transfer to Douyin, she wanted to work on “emotional incentives” for creators. Existing monetization models “would gradually change the original motivation for creating content”; she wanted to combine spiritual and material incentives into a “three-way win,” but thought large companies were still too rough in their execution at the time.
- Her conclusion is final: “After trying many times, you realize that no environment will let you charge forward without restraint. It depends heavily on whether people share the same aspirations. I have to create such an environment myself.” As for the “Jianying of the AI era” label, she accepts only the shared principle—“help ordinary people express what they want to express with technology”—not necessarily the tool layer itself.
16. OiiO’s Product Form and Its “Small Fire”: Simulating a Real Production Crew
- OiiO entered through animation and anime. “Before AI, creation was extremely expensive. After AI, the cost suddenly flattened.” Its original positioning was to let everyone produce their own animated work. The product simulates a real production crew, dividing the work into 7 roles with personalities—“the art director is drinking tea; the screenwriter wipes their glasses”—so someone with no animation experience can start with an idea and produce a short film step by step.
- Before launch, 100,000 people were already waiting for beta codes. She attributes that 50% to coincidence and 50% to getting several things right. At that point, the ability to make a video from one sentence did not yet exist; experiencing it created an aha moment. It was also free. “When your purpose is not to make money, there are genuinely many interesting things inside.”
- 2 UGC examples matter most to her. One OC creator who could only draw turned static comic frames into a moving, emotionally resonant story—a task that would require extensive collaboration in the traditional industry. Another was a highly introverted Beijing office worker who did not want to appear on camera and used OiiO to make a vlog of going to work, coming home, and feeding her cat. “That is what it means to give ordinary people a chance to express what they want to express.”
17. Seedance Changes the Cost Curve: Higher Costs Kill UGC
- With Seedance, model costs became relatively high and the team had to start charging. “Once the price went up, that UGC disappeared.” The user mix shifted from primarily UGC to PGC: “The people willing to pay for a tool still have to be semi-professional users who can earn income from the work.” She admits to “a little disappointment” because costs prevented many creative ideas from being made.
- The response has 3 parts: do not subsidize without limits—“People who come to harvest subsidies will eventually leave for a place with cheaper subsidies; there will always be somewhere cheaper, and the model providers are sitting right there”—shift toward Pro C and small B users who care about costs and can circulate in a commercial economy, and, most importantly, build independent value above the model. “It has to be a black box. Others cannot be allowed to surpass it explicitly.”
18. The Agent’s Technical Depth: Routing, Benchmarks, and the Lesson of MJ’s Survival
- In response to investors’ classic question about applications being absorbed by the model, she maps every stage of a video Agent: understanding and decomposing a user’s sentence with a language model, writing the script, defining characters, scenes, and props with image models, and splitting the storyboard into shots. Each stage needs its own benchmark to select the best model; the best model in the market is not the best model for every use case.
- MJ is her key evidence. “MJ’s hand-drawn style is extremely good. Even larger and more capable models have not eaten MJ.” Its character consistency is very poor, however, so other models have to fill the gap. “All of that requires routing.” Each image model has different prompt rules, and the language-model output beneath the routing layer also differs.
- 3 types of specialists have to work together: model-product people who build the chain and understand the logic and standards, quality designers who run large numbers of visual tests to determine under what conditions a style can be preserved, and algorithm engineers who ensure chain stability. “It does not depend on headcount. It depends on expertise. Very few people can do all 3.”
19. Sora’s Shutdown and the Field: At Least Three Players, on the Condition They Are Not Abandoned First
- On Sora’s shutdown earlier this year: “It is a shame. It really was a good model, and for a long time it had no competitor in video—it was far ahead.” She suspects a strategic explanation: OpenAI may have been using its stronger language models to build capabilities like Claude Code. In a large company with intense internal competition, a project can retain an advantage only if it is not the first one abandoned.
- The Seedance launch did not end the video-model contest. She sees at least 3 players remaining, with Kling as one of them, because video models require major resources, organizational capacity, and sustained commitment. OpenAI has the organizational and technical resources, but Sora can still be the first project cut in the company’s overall priority stack.
- Her comparison of Seedance 2.0 and Sora is nuanced. Seedance has stronger overall capability but weaker live-action performance: “Seedance’s live-action output has an obvious AI feel. During the period when Sora could make live-action footage, you could not see any AI in it—it was extremely realistic.” Sora is also better at preserving style. But for audio-visual synchronization, lip-sync, and motion smoothness, “Seedance is unquestionably far ahead.” Both technology and data may matter, “but data matters more.”
20. The Video-Agent Boom Is a “Delusion”: Wrappers, Scams, and 100-Point Marketing
- 卫诗婕 asked whether the explosion of video Agents reflected a burst of demand or capital manufacturing a market. 闹闹 did not soften the answer: “I think it is a delusion. Everyone thinks there is money to be made here, but in reality there may not be any real profit.” Many products are routing stations with an Agent label slapped on top; some claim to offer free Seedance access while actually using something else, collect a round of money, and disappear. Others put a foreign wrapper on the product and claim access to unreleased Kling products to defraud users.
- Marketing is inflated across the board. “Everyone advertises a 100. Some deliver 30-40, some 60-70, some 70-80, but there is definitely no 100, including us. Once it is a probability game, you can still draw a shot you do not like.” Her guidance to users is practical: there is no fast way to tell. “You just have to try. You have to try them all.”
- She believes there are “probably only 2 or 3” companies genuinely building video Agents. “It is hard—really hard. You need past experience and an openness to learning from existing technology.”
21. Anxiety and Calm When Seedance 2.0 Arrived: There Is Still Work to Do
- She went through a period of being “a little anxious but relatively calm.” After watching the promotional video, her first reaction was, “Wow, this is amazing—we’re finished.” Once she calmed down, she realized it could not be produced from one sentence. Xiaohongshu submissions surged when Seedance launched and then fell back, “because ordinary people simply cannot use it. It is still a model constrained by strong prompts.”
- The gap between strong prompts and natural language is clear in a fight scene. Natural language says, “You and him have a Wing Chun fight.” OiiOii’s prompt may require “thousands of words for every shot” to produce a good result; the Agent sits in the middle and decomposes the request. “When something is simple, there is no distinction between professional and non-professional creators. No one wants creation to be a complicated process. That is the most fundamental point.”
- Her conclusion: “It is fundamentally not a model that can eat every step of the process, so there is still work for you to do. The longer I work on video generation, the more I feel this will not be completely eaten. The product layer will definitely bloom, because the benchmark for what is relatively beautiful or relatively ugly is different for every person.”
22. Why OiiO Beats Using Seedance Directly: 40-Way Concurrency, Less Trial-and-Error, Entry at Any Stage
- OiiO users are a subset of Seedance users, but they value different things. Using the same Seedance model, OiiO lowers the odds of a bad draw, can generate storyboards concurrently, and does not require professional prompt skills. Members can run up to 40 concurrent seats—“40 shots can come out at once.” The most vivid validation is that “students use us when they are rushing to finish homework.” That is a sign the product is usable.
- The structural difference is that OiiO splits the work into 7 roles and channels. A persona can help write the story; a script can help build the character; users can enter at any production stage and complete the workflow. “Seedance is only the final generation step.” It suits production teams that already know exactly what they want in every shot. The industry’s “draw-card operator” role is, in rough terms, about prompt skill. “We solve 60% of it. A piece that others need 6 or 7 draws to produce, we solve in 2.”
23. Consumer Is a Temporary Necessity: Inference Costs, Doubao Charging, and “Unlimited Cost”
- She disagrees with the thesis that foundation models should abandon To C and focus only on Pro C and B, but agrees with the current reality. “Would ByteDance not want to serve consumers? Of course it does. But with Seedance this expensive, is it really making money? Not necessarily.” Once costs rise, consumer usage cannot circulate economically and inevitably shifts toward somewhat more professional users.
- The strongest evidence is Doubao starting to charge. “A company that has created so many To C breakout products charging for a To C product is unimaginable. That shows that even a company as large as ByteDance cannot completely ignore this. Free is simply too expensive.”
- Video makes the cost structure even more extreme. In the past, a self-media creator might spend several thousand yuan on a DSLR or Insta360—a “finite cost” that leaves room to make mistakes and improve. “Generation is an infinite cost. Every attempt is another attempt. It is a bottomless pit.” That is why people post their AI spending bills: “You cannot guarantee that you will produce a viral hit. You are genuinely gambling.” OiiO’s role is to improve the odds and raise certainty.
24. June 10’s 2.0: Shot-by-Shot Recreation, Skills, and What the Black Box Actually Hides
- The release has 2 new features. Shot-by-shot recreation reflects the fact that “all creation begins with paying tribute to and imitating excellent work”; it raises both the probability of producing something good and the user’s motivation, because there is at least feedback. Skills are generated by feeding in around 5 videos the user considers particularly good, automatically creating a skill such as “interview.” “It is somewhat like a template, but less fixed than a template. It has the feel of the source.”
- Why can’t competitors simply copy the layer? The shot-by-shot process is itself relatively black-box. “You cannot really go through 24 frames per second one by one—if you truly did that, there would be no expression, because it would be static. You have to go shot by shot. But how the shots are cut is determined by people, and how the content is understood after the cut is also determined by people.”
- Her experience building Douyin effects is evidence. A single cartoon transformation effect took 3 months of training, including extremely fine-grained human adjustments to facial structure and definitions of what looks beautiful or not. Making an effect go viral also depends on layers beneath the algorithm: users interested in celebrities receive celebrity effects, while parents may receive them through the social graph of nearby friends. “There are many underlying things in the middle beyond the algorithm.”
25. Sector Call: De-Staffing Short Animation Dramas; Self-Media Is the Bigger Pool
- Her case for ACG—animation, short animation dramas, and games—is that large companies will spend heavily training models and push urgently toward commercialization. “The low-hanging fruit will definitely be picked by large companies first.” Startups have room in the uncertain middle layer where new value is still being formed. She is blunt about animation: “It does not make much money. The industry itself is quite difficult. Productivity has been released, but production costs are still high for ordinary people.”
- Her core organizational call is that short animation dramas have remained labor-intensive over the past 6-12 months, but “we believe they will definitely become non-labor-intensive. Labor intensity is only a temporary cost advantage; it still cannot turn content into premium content.” Many industry workers who leave will form small units, and OiiO serves this small-B segment. “That is also the trend now.”
- The larger opportunity is “self-media.” Hongguo follows a drama-industry-chain model with obvious concentration: “Maybe only the top 5% of people are making money.” The remaining 95% will flow to larger platforms such as Douyin, Xiaohongshu, and Bilibili, adopting content formats that fit them and finding a way to survive. That is her definition of self-media. She agrees with 卫诗婕’s summary: as production formats evolve, the individuals pushed out of the old system will have strong incentives to use her tool.
26. The Relationship with Big Tech: “I’ve Always Thought Everyone Was Sitting in the Car”
- How does a startup avoid being crushed under a large company’s tire tracks? “I have never thought we were under the tire tracks. I have always thought everyone was sitting in the car. It is not a competitive relationship. People who have never worked at ByteDance are more likely to make the other judgment.” Her ByteDance experience provides the mechanism: parallel exploration across many directions reflects a “try-it” mindset; the people selected to run a project are not always deeply committed, resources are spread across projects, and the setup can even be a disadvantage. She is focused on a single direction that can support itself. 卫诗婕’s metaphor also works: Agents are like nomads, living off whatever mountain or grassland they find.
- The mismatch in value scales protects startups. “ByteDance has many products with tens of millions of DAU that are considered to have no value inside ByteDance. For a startup, tens of millions of DAU definitely have value.” The bottom line is independence: “Even if I have only tens of thousands of users, if I retain my independence and can run positively, this startup is not bad. Once a startup loses its independence, it becomes very easy to eliminate.”
- She has no commercial allegiance when selecting models: “Choose objectively the model best suited to the situation, unless it refuses to serve China.” The method is to run 100 cases 100 times in each scenario and use whichever executes most accurately. The main domestic models currently routed are Kling, Seedance, and Vidu; Hailuo is used less now.
27. Price Hikes and the End State: Partners, At Least Not Rivals; 70-80% Chance the Agent Layer Survives
- What if model providers keep raising prices? She starts with the incentive: “If raising prices means more people cannot make money, it is unsustainable and they cannot keep raising prices. My judgment is that they cannot.” Short animation dramas, despite being a strong segment, have low margins; ByteDance cares more about a flourishing AI-content ecosystem on Douyin and needs to release UGC creative capacity. “Maybe cutting prices actually maximizes value.” She says you have to trust that ByteDance is a highly rational company: if it can raise prices without harming most stakeholders, it will; if it does not, the current price may already be optimal. She allows an exception: Anthropic can charge aggressively while it holds a lead in coding, until a competitor such as Codex changes the landscape.
- Her relationship with foundation models is partnership, not rivalry: “We are partners, grouping together to fight. Everyone needs building blocks, and the better the blocks, the better it is for Agents.” But she assigns only a 70-80% probability that the Agent layer will not be absorbed. When 卫诗婕 asked why she would not say 100%, she replied: “This is not conservatism. It is objectivity. If something is 100%, believing it is so is stubbornness. I am not stubborn, and I am not betting. I am just reasoning it through.” The logic is cold but coherent: “If it can eat every Agent, then we will stop doing this. There is no point. We will not create value just to prove that we have value.”
- Asked about the market’s argument that founders have to claim 120 points to create positive incentives, she offered one of the episode’s lines to remember: “People speak with far too much certainty. ‘The wolf is coming’ has been cried too many times, and eventually no one may believe it. What is most scarce at that point is saying something objective and relatively true.”
28. Timing, Seedance 2.0’s Lesson, and the End State: Humility as Courage
- Was it too early to build an Agent when models were usable and approaching good but not actually good? “To be honest, it was a little early, but I no longer think being early is bad.” You cannot always catch the perfect entry point; entering early lets a team accumulate knowledge about what works, what does not, and where the gaps are. “The first person to appear will always be remembered by more people.” Was the timing perfect? “I think we should have entered then.” Her way of reading model companies is to identify creator pain points—“you can be almost certain that they are directions the model must pursue”—and then talk to former classmates.
- What Seedance 2.0 truly taught founders was to “make the product detailed.” Everyone knows data matters, but Seedance showed exactly where and how to execute that insight. “The execution is fragmented and boring. It is people stacking up extremely dirty, exhausting work. It is not flashy at all, not sci-fi at all—just honestly and solidly doing what needed to be done.” AI startups are particularly short of that investment in details and dirty work. 卫诗婕 pushed back that startups cannot beat big tech on detail-heavy execution. 闹闹’s answer is that startups have their own hard labor: algorithmic and engineering details in the layer above the model.
- Her ideal endgame for a top-tier AI product jumps from practical execution to science fiction. It is “not a phone and not glasses,” but content interaction in which audio-visual language creates stimulation signals inside the brain through neural transmission. “Being able to interfere with those signals is more fundamental.” 卫诗婕 framed that as the true metaverse; 闹闹 answered only, “Maybe.” Her profile of the top AI-era product manager includes curiosity and the ability to rationalize aesthetics, but she ultimately lands on a rarer quality: “Humility means you can learn. Being willing to admit your shortcomings is not low self-esteem; it is actually a sign of great confidence. It is a kind of courage.”