OiiOii Founder 闹闹 on 2026, the Year of AI Video
Summary
- OiiOii is still only about 4 months old, with 1 month of beta testing, 18–19 full-time employees, and zero revenue or profit—but its product bet is already highly concentrated. The team is raising a Pre-A round that is now nearing completion; after Sora 2 emerged, it paused the original first-and-last-frame pipeline and decided to “stop building the old pipeline for now and switch everything to Sora 2.” It is a deliberate choice to trade consistency risk for stronger finished-video performance.
- 闹闹 does not believe future models such as Sora 4 and Sora 5 will simply swallow video Agents, because foundation models and vertical applications are more like a “market” and a “restaurant.” Models provide ingredients with different strengths, while Agents build knowledge bases, visual language and orchestration strategies around specific audiences; roughly “60–70% of the work” at OiiOii goes into seasoning and controlling the heat, not user-facing features.
- OiiOii’s initial wedge is neither the heavily distribution-driven manhua-drama market nor a UGC community that requires massive users and inference costs, but individuals and small self-media, animation and science-education studios. 闹闹 says these 2–3-person teams used to release roughly 1 episode a week; with OiiOii, they could theoretically release “10 episodes a day,” or 1–2 carefully selected episodes a day. The first value to validate is a step-change in production capacity.
- The beta revealed demand broader than the original target market, but the team plans to take verticals one at a time rather than serve everyone at once. Manhua-drama teams download only the efficiently generated storyboards and edit them themselves; non-professional users use animation to express mental states they do not want to show on camera; parents, couples, pet owners and adult children also use it for private relationships. The strategy is to turn the structure of each category’s best videos into a knowledge base and “攻 one by one.”
- One central challenge for Agent products is the structural conflict between freedom and reliability. OiiOii has rebuilt its architecture 4 times in 2 months: from free but disobedient System Prompt-based multi-Agent workflows, to stable but uneditable Workflows, and then to a system where Agents could jump out, modify the process and return. It still sometimes “can’t get out,” “stops moving again,” or forgets to return after jumping out.
- Sora 2 has changed the value distribution across editing tools: complex editing may be replaced by models first, while the simplest cutting, sequencing and TTS assembly remain. 闹闹 once thought transitions, effects and other heavy editing would be difficult to replace; she now believes models absorbed those techniques while training on video. The future is more likely to compress a 3-hour CapCut job into 30 minutes than to make CapCut disappear.
- 闹闹’s view of 2026 remains restrained: quality, editability and real-time performance will probably continue improving, but the technology leap that truly expands the audience may not be the grandest one. Stronger editing could push products toward professional users, while real-time interaction may be constrained by passive consumption habits shaped by “swiping up and down”; Sora 2 has already shown that simply making cuts more natural can create a much larger audience.
Deep dive
1. OiiOii Is Still in a Zero-Revenue Beta, but Its Bet Is Already Concentrated
OiiOii is still at a very early stage: the company was founded about 4 months ago, the product launched in beta 1 month ago and still requires an invitation code. The full-time team has 18–19 people, a Pre-A round is underway and nearing completion, and the company currently has “no revenue and no profit.”
The product’s one-line definition is “using AI to make animation.” 闹闹 previously spent years in video content and creative tools, from Tencent QQ Mail’s mobile app to an extreme-sports content startup, then ByteDance, where she oversaw CapCut and effects for Douyin and TikTok, before running animation at Bilibili.
Actual use cases have exceeded the team’s expectations. Koji used it to make a Christmas-song music video for his daughter, parents make daily shorts for their children, and teachers create animated explainers. 闹闹 acknowledged that these creators are “quite different from what we had imagined.”
2. A More Than Decade-Long Dream of Animation Found Its Entry Point in Multimodal Models
闹闹 had wanted to make animation since college. During the 6 months after leaving Tencent, she studied character design and Maya, but backed away after finding the software “too difficult to use,” while the industry offered low pay and often relied on people “working for love.” At a dim, air-conditioner-less effects company, a modeling director had worked there for 5 years and earned RMB10,000 a month, yet still had a spark in his eyes—a scene that stayed with her.
The true technological trigger came with DALL·E 2 in 2022. 闹闹’s first reaction was not about images but: “This could make animation.” She then learned about the industry on Bilibili and kept probing for an entry point while working on “离谱.”
3. The First Version Used First-and-Last Frames and Model Routing to Stabilize Individual Shots
The team planned the product in July and began development in August, initially following the first-and-last-frame route common among video Agents at the time. If a finished video contained 8 shots, the system generated 9 images, then created 8 video segments from adjacent first-and-last frames and stitched them together, improving character and image stability within each shot.
The first version’s real differentiator was the Task Agent. It selected models based on the storyboard task: an action-specialized model for a fight scene and a more nuanced model for an emotional close-up. 闹闹 later gave the example that a fight sequence might call Hailuo, which performed relatively well in combat scenes.
After 1 month of development, the pipeline was already producing good results, and the original plan was to add TTS, background music and sound effects. The core assumption was that video models would struggle to become a single universal model: differences in data, labeling and training direction would ultimately produce different strengths.
4. Sora Revealed the Potential of Finished Videos; Sora 2 Made the Team Shut Down the Old Pipeline
After Sora emerged, 闹闹 saw an animation made with it on Twitter and “couldn’t see any trace of AI at all.” After integrating it, the team generated a complete short film of a little crab and a little star playing basketball. 闹闹’s first reaction was: “Could we really have pulled this off?”
闹闹 saw Sora’s standout capability as natural cutting, with montage-like visual language appearing in the finished work. When Sora 2 later arrived, she felt it must have been trained on a great deal of film material for animation; even its remixes carried unmistakable cinematic traces.
In the reference-image-plus-text approach, text carries substantial weight, while the reference image is not a hard constraint in the way a first-and-last-frame setup is, so consistency is less stable. Even so, after Sora 2 appeared, the team moved from running 2 pipelines in parallel to “stopping the old pipeline for now and switching everything to Sora 2.”
5. Foundation Models Are the Market; the Agent’s Value Lies in the Cooking
Koji asked whether stronger models such as Sora 4 and Sora 5 might eventually swallow OiiOii end to end. 闹闹’s answer was “no,” because models are difficult to unify; their relationship is more like a large market and a restaurant. Users can buy ingredients and cook for themselves, or go to a vertical restaurant for a finished dish.
If OiiOii is a Sichuan restaurant, it must select ingredients from different markets that suit people who like spicy food. The real differentiation is not whether it calls Sora 2, but how the chef controls flavor, heat and combinations. 闹闹 estimates that “60–70% of the work” at the company goes into these invisible details.
闹闹 acknowledges competition among video Agents, but believes video content is rich enough to support many specialists: Sichuan cuisine, Cantonese cuisine, Hunan cuisine and hotpot can each have their own workflows. The market may be more like collectively expanding a “snack street” than a universal model exhausting every professional need.
6. Knowledge Bases and “Just-Right Incoherence” Form the Invisible Product Layer
Take an anime music video. 闹闹 believes that giving Sora 2 the same Prompt and images still would not produce OiiOii’s results. The team collects high-quality anime MVs, has the model learn their content and builds a knowledge base. When a user writes, “Have these characters dance to a K-pop song and make an MV,” the system expands the Prompt and supplies visual language and sound effects.
Coherence is not always better when maximized. If every storyboard beat connects perfectly, the plot becomes “very flat,” weakening sudden turns and emotional rises and falls. The team intentionally preserves some breaks, then keeps tuning toward the right balance between dramatic tension and viewing smoothness.
Experiments with scene consistency showed that the team could overdo the “ingredients.” Adding a scene reference image did unify the background, but it also suppressed the model’s imagination, making the characters look pasted onto the setting. 闹闹’s analogy: “The user says this Sichuan restaurant isn’t spicy,” so the team adds chili peppers aggressively—and ends up making it too spicy.
7. The Team Deliberately Avoided Manhua Dramas and UGC to Start with Small Self-Media Studios
OiiOii initially considered manhua dramas, but ultimately decided it was not fully ready for them. The category places heavy demands on scripts, which are not the team’s strength; commercially, it depends on paid distribution; operationally, production remains an efficiency optimization within a mature workflow and still relies on labor, contrary to the team’s goal of reducing dependence on manual work.
UGC content communities were also set aside for now. 闹闹 believes users already consume information at high density, while UGC has short consumption cycles and insufficient information density to support a large app. Communities also require huge user numbers and carry high product and inference costs.
The team therefore chose the middle ground: individual creators and small 2–3-person self-media studios, including animation authors who continuously operate IP, ACG MV creators, and people who previously made history or science content but avoided animation because production costs were too high.
闹闹 cited research showing that these small studios typically release 1 episode a week. With OiiOii, they could “basically release 10 episodes a day”; even with careful selection, 1–2 episodes a day would be manageable. The first value to prove is a compressed production cycle.
8. Beta Users Crossed the Original Boundaries, with Private Expression Emerging as Unexpected Incremental Demand
Manhua-drama teams came looking for the product themselves. They did not need OiiOii to deliver a complete finished video; as long as the storyboards lined up, they could download the full set and edit it themselves. The current product has no manhua-drama-specific capabilities such as script uploads, but it has already solved one time-intensive part of the process.
A second unexpected user group had never made videos and disliked appearing on camera, but wanted to use animation to express their own state of mind. 闹闹 believes animation is naturally suited to expressing an inner world. These are consumer users, but not the professional creators the team originally envisioned.
Some content is not made for public consumption at all, but circulated within relationships: parents making videos for children, students for teachers, couples sharing with each other, people making videos of their pets for private viewing, and adult children making videos for their parents. This private-expression demand only became visible after the beta.
9. Attack Verticals One by One; Keep the Agent Interface from Becoming a Feature Pile
Faced with multiple user groups, 闹闹 does not plan to expand into all of them simultaneously. The team will borrow Douyin’s approach to verticals: find high-quality creators in areas such as science education, break down the structure of their animated videos, turn it into a knowledge base and “攻 one by one” according to priority.
闹闹 believes Agents can break the cycle in which creative tools evolve from simple to bloated. Photoshop, Sketch and Figma, as well as Premiere, Final Cut and CapCut, have all essentially continued piling features onto a GUI. Agents can hide complex capabilities internally and let users discover uses through conversation that “even we didn’t know about.”
10. Four Architecture Rebuilds Have Yet to Resolve the Tension Between Freedom and Stability
Koji pointed out that while Agents are flexible, they are less controllable than CapCut or a fixed Workflow. OiiOii must keep the production line stable while allowing users to converse freely with every Agent along it, making the underlying architecture the product’s central challenge.
The team rebuilt the architecture 4 times in 2 months. The first version used System Prompts to define Agents and let the model decide the Workflow itself; it offered high freedom but was “disobedient.” The second switched to a strict Workflow, which stabilized the process but removed room for modification.
The third version allowed an Agent to jump out of the Workflow after receiving a signal, make changes and return. The fourth further allowed it to move between processes. The direction is more reasonable, but failures remain common: it “can’t get out,” “stops moving again,” forgets to return, or never jumps out at all.
11. Sora 2 Removed Heavy Editing, While Light Editing Is Harder to Eliminate
闹闹 sees OiiOii as additive rather than directly cannibalizing CapCut. CapCut offers general-purpose tracks and is connected to Douyin’s templates and content ecosystem; OiiOii delivers specific MV, science-education or animation content, and users typically return to CapCut for post-production after generation.
闹闹’s view of a “Cursor for video editing” changed because of Sora 2. Previously, single-shot models could not cut naturally, so transitions, animation and complex effects still had to be handled by editors. Now the model may have learned those techniques along with the training videos: “That layer has been removed from the original judgment.”
The operations least likely to disappear are cutting the final 0.1 seconds, changing the order of clips and assembling TTS at the end, because dragging is more direct than writing a sentence. The result may not be CapCut disappearing, but a single editing job shrinking from 3 hours to 30 minutes. “The optimal solution is the combination of the two.”
12. Seven Character Agents Are Only the Front End; Emotional Language Is the Key to the Finished Video
OiiOii designs creation as a team serving a director: the script Agent summons the character-design Agent, like inviting a colleague into a group chat. Users see roughly 7 characters on the surface, while invisible assistant Agents operate underneath, forming a two-layer system of collaboration, context and memory.
闹闹 sees a close parallel between product managers and directors: both must orchestrate specialists with different capabilities to complete a work. Having an Agent join the group chat is both an engaging character experience and a visual representation of Workflow-state changes.
The “feeling” of the finished video comes from translating emotion into cinematic language. When a user says only “sad,” “joyful,” “healing” or “lonely,” the system must expand that into a long corridor, gray-white tones, composition and visual elements—“using rational elements to express something emotional.” But that is not the same as an artist’s flash of inspiration. 闹闹 acknowledges that advanced creators can still produce work she finds unbelievable, while the product’s freedom remains insufficient.
13. WeChat Trained Human Intuition; ByteDance Trained Data Strategy
Tencent left 闹闹 with a product philosophy. Teams were immersed every day in huge volumes of user feedback, listening not only to what users said but also watching what they actually did. 闹闹 describes user experience as an intuition strengthened through long-term training, “a bit like a large model.”
The power of the WeChat system often lies in details. 闹闹 remembers the version copy “Everything I say is wrong” paired with an image of Michael Jackson, and also points to the inspired touches in Mini Games. Zhang Xiaolong-style decision-making was not about rejecting opinions, but about listening openly before converging more deeply.
When she first joined ByteDance, 闹闹 resisted a purely data-driven approach: enlarging a button might improve a metric without improving the experience. Later, while working on the effects submission rate, she came to understand strategy products—using real behavior such as browsing, saving and carousel clicks to push a trend to the people most likely to use an effect.
闹闹 ultimately summarizes the 2 systems as “one right brain and one left brain”: WeChat excels at understanding users and human nature, while ByteDance excels at data science, recommendation and growth. Their shared principle is to “push their strengths to the extreme,” and AI animation happens to require both artistic intuition and model strategy.
14. A Good Product Manager Needs Empathy, Self-Reflection and Technical Sensitivity
闹闹’s first capability is empathy: even with extensive experience, she must be able to switch instantly into the mindset of a complete beginner. Her method is to “frequently pull myself out of ‘myself,’” observing both herself and users so that experience does not become a blind spot that hides unmet needs.
The second is “50% confidence and 50% self-reflection.” Confidence alone can become arrogance; recognizing shortcomings should not become low self-esteem. 闹闹 wants self-reflection to reinforce confidence while reducing an overdeveloped sense of self.
The third is technical sensitivity. A product manager does not necessarily need to write code—闹闹 herself does not—but must understand what technology can make possible. That matters even more in the AI technology revolution and remains a basic skill. Her long-standing interest in visual, auditory and physical rules happens to converge with multimodal audio-visual technology.
15. Conflict and Calm Are Two Sides of the Same Force; Entrepreneurship Filters for Long-Term Conviction
闹闹 believes “conflict is the force that gets things done.” Healthy competition is like facing an opponent in basketball: it unlocks potential. When multiple ByteDance teams competed for the effects business, fighting together through the conflict turned opponents into teammates because everyone ultimately put “the work first.”
During her first startup, the first person 闹闹 fired was a close friend. The friend cursed her out, then poached several team members to do the same thing. A year later, the friend added her back on WeChat and said he finally understood 闹闹. She came to believe that as long as the starting point is not to harm anyone, “time will ultimately prove everything.”
The rock music, extreme sports and rebellion of her youth were attempts to find maximum freedom in the outside world. Later, 闹闹 realized that this forceful outward expression could also become a cage. Today’s Peace is not a lack of force, but a broader, steady force that flows over time. To 闹闹, noise and stillness are simply two sides of the same search for freedom.
16. 2026 May Not Be Won by a Grand Breakthrough; OiiOii Will First Be a “Container for Expression”
For 2026, 闹闹 expects video-model quality and editability to keep improving, while real-time performance may also strengthen and enable interaction. But she does not equate those developments with mass-market breakthroughs: greater editing freedom may push audiences toward professionals, while each additional step in interaction may filter out users accustomed to passive swiping.
闹闹 also refuses to place hard limits on the trend. Sora 2’s core change may look like nothing more than more natural cutting, yet it significantly improved the consumption experience. The market may therefore expand through “small changes built on the existing medium,” rather than through a grand technological revolution.
OiiOii may eventually extend from a tool into more possibilities, but 闹闹 does not want to predefine it as a content platform or “Pixar.” The immediate goal remains helping everyone who wants to make animation actually make one. Defining it too rigidly would turn directions that are still open to exploration into fixed objects.
闹闹 ultimately defines both herself and OiiOii as “containers”: not expressing things for other people, but serving as the channel between animation and creators. Faced with a $3M angel investment allocation, she still chose to invest in herself—not by claiming she is the best, but because “the person I know best is this one.”