Pioneers Insight Method Research Author
A Conversation with 晨然 on Medeo, AI Video, and Young People
Back to Episodes

A Conversation with 晨然 on Medeo, AI Video, and Young People

Summary

  • Medeo’s core bet is a “dimensional upgrade in modality.” One Two X believes video will become an information format as fundamental as text, turning high-density written information into video: “Its unit price would then rise tenfold—I’m just guessing here; it’s not something that can be verified.” The team ran an MVP through its own AI news account: articles with only a few hundred reads drew strong viewer interest once converted into video, at a manageable production cost. That was the team’s first instinctive validation of the sector.
  • Chen Ran does not shy away from the comparison with what InVideo could already do a year and a half ago. Editing products already have an SOP for solving the problem, and “at the very beginning, they will definitely look extremely similar.” Differentiation comes from chasing the 3% innovation—“fewer tokens, stronger generalization, and higher-quality information products.” He does not believe this is a winner-take-all market: Jianying is like an IDE that can do everything, and “bloat is a disadvantage.” Go vertical, go deep, and keep polishing—“people will pay for that.”
  • His clear architectural stance is that video should be “80% workflow and 20% agentic.” Editing has a fixed human SOP—organize footage, write the script, rough cut, fine cut, then package it—and agents may deliver less reliably than workflows in highly procedural settings. Agents are better suited to open-ended problems: “Open problems require open-ended solutions.”
  • Two pieces of counter-consensus are worth recording. First, he is “extremely cautious” about MCP: it is a protocol for ecosystem portability, not a technological breakthrough, and “at the end of the day it is still a way of managing prompts.” Integrating it into a vertical product adds engineering complexity and may not solve a real technical problem. Second, do not worship the best model: switching Sonnet from 3.5 to 3.7 “actually made the results worse.” Each vendor makes different alignment trade-offs, and this kind of knowledge “generally doesn’t circulate publicly”; only people tuning prompts on the front line see it.
  • Reliable delivery is the hardest and most underestimated opportunity in AI products: “AI products are looking for certainty within randomness… You have to do work to overcome entropy and get a stable result.” He sees this as a product problem, not something to wait for in the next generation of technology. His analogy is Zelda’s open world: freedom is created by carefully designed rewards, from towers to Koroks. If an AI product can design the positive feedback from randomness well, it may have its own Breath of the Wild moment.
  • On opportunities for young people, Chen Ran makes an explicit conflict-of-interest disclosure: AI is so new that experience can become a constraint, and “in this new era of AI, I think everyone is equal.” 融汇 also cites former guest 童超’s view: the race is about iteration speed. At 25, with no business experience, Chen Ran’s experience at a major tech company was that “because nobody understood it, I became the person who understood it best.”
  • One footnote of the era: In a downturn, “just keeping your emotions stable is already quite an achievement.” When he talked about dreams with a close friend, the response was, “That’s great. What does it have to do with me?” For this generation of founders, preserving the urge to create has become one of the hardest things. His defense is to physically separate himself from his phone, read for input, and admit that “most of the source of my creative drive comes from pain.”

Deep dive

1. No launch campaign: Tool products need a year of polishing, so “develop quietly”

  • Medeo’s servers crashed after launch, and the team “probably pulled an all-nighter.” The product briefly flooded people’s feeds, but Chen Ran says it did not feel especially real to him. The launch was defined as “a very embryonic attempt,” with no distribution targets set from the start. Seeing public accounts and KOLs share it spontaneously was “honestly, still somewhat beyond expectations.”
  • The logic behind the low-key launch is that tool products need a year to polish the details; it is too early to talk about a product with a high level of completion. The company’s DNA is “not that of a team trying to chase traffic.” It would rather “develop quietly” and build word of mouth, and at this stage it is explicitly not focused on metrics such as DAU.
  • The only feedback that genuinely reassured him was users saying the functional areas were clearly divided: “I feel like those several months of major restructuring weren’t wasted after all.”

2. Why video: A dimensional upgrade in modality is “an overwhelming economic upgrade”

  • One Two X was founded by former Kimi product leads 王冠 and 姚行, together with a technical co-founder. Chen Ran joined because of the intersection of 2 traits: he is a programmer and a content creator who has worked as a director. “AI video happens to be where I can make the most of both traits.”
  • His timing call was that, last year, the cost and reliability of fully automated video generation were still insufficient. The team pivoted along the way, then saw more and more generation workflows emerge and concluded, “The pace may already have become fast enough.” The positioning is end-to-end generation of a deliverable video that can still be edited in the project file. Whether the underlying process involves generation, editing, or retrieval “is something the AI behind Medeo handles, not something the user needs to care about.”
  • The initial economic validation came from running an AI news account: articles with a few hundred reads attracted viewers when converted into video, and the cost was not high. The core belief is that “a dimensional upgrade in modality is an overwhelming economic upgrade”—turning a piece of text into video “would raise its unit price tenfold. I’m just guessing; this isn’t something that can be verified.”

3. The first lesson from turning articles into video: Video has its own grammar, and viewers leave when you read out data

  • 融汇’s sharp question—she has written scripts herself—was confirmed by the answer: a transcript script and a video-language script “do not need to match the article itself very closely.” How to write the hook, land the ending, and deliver emotional value in the middle all belong to video’s own language system. The most valuable information density in an article—tables and data—is precisely the hardest thing to use in video: “Once you read the data out loud, viewers don’t want to watch.” More important than hard information is the “feeling of getting value”—making viewers feel they learned something and enjoyed watching it.
  • The team has not built detailed script templates by category; the system is “still fairly makeshift.” Half the reason is limited bandwidth, and the other half is a deliberate belief in less structure: as models improve, writing tasks do not require people to provide so much structural guidance. “Giving the model a certain amount of authority may work better. I’ve really felt this firsthand—just let it take control.”

4. One person is a whole production unit: Watching videos is work, and writing code is too

  • Externally, he is the product lead; in practice, his main job is improving video quality: researching new video possibilities, defining categories, manually cutting demos, and doing prompt engineering. “I built the entire video-generation algorithm for Medeo.” His daily work combines intuition and analysis: “One part is making content, and one part is writing code.”
  • He spends a huge amount of time immersed in Douyin, Xiaohongshu, and YouTube Shorts: “Only by staying immersed in the vibe of these videos can you possibly get how the emotion is actually being conveyed.” He then manually imitates the edits, analyzes the steps, and turns them into runnable code.
  • That immersion has sharpened his judgment. Videos that survive on Douyin “all have an extremely similar sense of rhythm, almost exactly the same,” allowing him to tell at a glance whether a video will survive in the feed. YouTube Shorts explainers “basically have one style”; since it is a style, “I believe it can be structured.” His final conclusion after all that watching is “video is emotional value.” How much knowledge it transmits is secondary to whether it feels good to watch.

5. Fully remote, no KPIs: Solve management problems at the hiring stage

  • One Two X is fully remote. 王冠 and 姚行 are both in Beijing but do not work from the same office every day. Recurring meetings are deliberately kept under 30 minutes, and all-hands meetings under an hour. “After you reduce meeting time, most of your time is spent alone… That places extremely high demands on people’s initiative.”
  • The team sets no KPIs. Chen Ran’s explanation is that the job should be done through hiring: “If you hire truly high-quality people and give them an environment that is well suited to innovation, they will proactively produce innovative things.” His coordination with the 2 co-founders is almost entirely strategic—when to do what, and which paper validated which technical hypothesis. “The dirty, difficult work… I feel like I can handle it alone.”

6. Bets and validation: DeepSeek validated reinforcement learning, while Google’s unified multimodality remains a major bet

  • The team had been studying reinforcement learning throughout last year. “We believed this path would have a very big breakthrough.” Then DeepSeek appeared: “We were incredibly excited that day. This path was right,” because the product assumptions had already been built on that technical hypothesis.
  • 姚行’s paper-review discipline is worth noting: every week, he uses O3 to go through every paper in the ecosystem, then selects several papers on important technical paths to share at the weekly meeting. Another long-term bet is Google’s unified multimodality: “We were betting on this path last year, which is why One Two X exists.” Last week, Google “dropped another thing.”

7. Facing the sameness critique head-on: The solution has an SOP; everything looks similar at first, and the difference is 3% innovation

  • The host relayed the market’s negative reaction: Medeo launched in May 2025, but “what it can do is what InVideo could already do a year and a half ago.” Chen Ran acknowledges that “this kind of criticism was inevitable.” The team saw the leading competitors from the start: AI editing products cannot escape the functions required by video expression, and one-click generation plus the editor already have “an SOP for solving the problem.” Convergence is structural.
  • His response is to “look for that 3% innovation”: “We want fewer tokens, stronger generalization, and higher-quality information products as the output.” “Making another Jianying or another InVideo doesn’t seem to have much meaning.” It may look perfectly normal at this stage, but “future iterations will look very different.”

8. The mission: One piece of information, multiple forms of expression; video is only the first

  • The underlying insight is that people are naturally trained in language, but “nobody is naturally trained in the grammar of video.” How to tell a story in a vlog and how to order an edit are bigger barriers than learning the tools. Medeo wants both untrained beginners and professionals to express themselves quickly through video.
  • The company’s strongest conviction is “one piece of information, multiple forms of expression.” The same source information should be expressed differently for older people, children, and young adults. With AI, users no longer need to care whether the middle layer is editing or generation. “What we want to do is minimize as much as possible the pain between one piece of information and multiple forms of expression. Video is simply the first form of expression we’ve chosen.” He also admits that Medeo has not yet clearly defined which video categories it serves; that will depend heavily on market feedback and the direction the technology takes.

9. 融汇’s follow-up: If others agree with the mission too, why won’t the market be winner-take-all?

  • 融汇’s pushback was direct: “I think Veed (phonetic) and InVideo may also agree with this mission… Could their mission be the same, while they were already doing what you’re doing now a year ago?” Chen Ran says it is “a very reasonable question.” His answer rests on the breadth of the video market: marketing videos rely on emotion, sound effects, and transitions, while podcast videos are driven by voice and text. “The expressive elements they depend on are different,” and the category determines the shape of the tool. Medeo has tentatively chosen news, explainers, knowledge, and stories, rather than talking-head videos with captions.
  • He directly rejects the idea that Veed has not yet solidified its product: “I believe Veed solidified its product a long time ago.” Its marketing shows that the core selling points are largely pain points around captions and talking-head videos, and it has built recognition through flashy caption effects.
  • His argument against winner-take-all is that Jianying “is like an IDE that can complete all kinds of code.” Even if it does everything, “its bloat is a disadvantage.” Some users will want a simple tool that only edits podcasts or only adds captions. “And if you go sufficiently vertical, you can keep optimizing over and over. You might even train a model… People will pay for that.”
  • On the threat from general-purpose agent companies adding video generation—Lovart (phonetic), for example, can already generate video directly as a design agent—he says generated output is “very hard to modify.” If you need to change captions, the GUI editor is still the better destination. “The final products may all end up looking similar, but the paths will be very different.”

10. 80% workflow + 20% agentic: Editing already has a human SOP

  • The team has repeatedly discussed the boundary between agents and workflows. An editor’s work is already a fixed process: organize the footage, understand it, conceive the script, rough cut, then fine cut—tighten pauses, add captions, and package the video. “Since human brains, or editors, already work this way… if something is already highly procedural, I don’t think an agent is necessarily the best fit.”
  • He prefers the word agentive to agent. An agent sounds like a worker entity; agentive describes a technical approach suited to open-ended problems where “you don’t know what the user will do.” “Open problems require open-ended solutions.” His conclusion is clear: “The video space should be 80% workflow and 20% agentic. That is how you can deliver a result more reliably and consistently.”

11. Counter-consensus No. 1: MCP is an ecosystem protocol, not a technological breakthrough

  • Asked which emerging consensus he agrees with least, his first answer is MCP: “I have a very cautious attitude toward MCP.” It is a protocol that solves ecosystem and portability problems, not technical problems. “It has not brought any epoch-making technology… at the end of the day, it is still a way of managing prompts.” Context management, structured output, and cost reduction already have many other solutions.
  • The market discussion is intense, and the term appears in every marketing hook, “but in actual production deployment, it has not really solved a major problem.” He distinguishes the beneficiaries: platform products such as Doubao naturally benefit from a pluggable plugin ecosystem; but vertical products do not need to talk about ecosystem expansion in the first place. Adding MCP brings engineering complexity without solving the product’s actual technical problem.

12. Counter-consensus No. 2: Solving a user need does not necessarily require the best model

  • In a system combining multiple models, not every node needs the strongest model. He found in testing that switching Sonnet from 3.5 to 3.7 “actually made the results worse.” Version 3.7 had a small bug in instruction following. The reason is that vendors make different alignment trade-offs: Claude prioritizes coding ability, “so it may incur small losses on other tasks.” You only see that by testing the model in your own business context.
  • The meta-insight is more valuable than the conclusion: “You only discover this kind of thing by being on the front line and actually tuning prompts. These insights generally don’t circulate publicly.” They may appear in one obscure corner of one forum. The lesson is not to blindly assume that the best technology is always the right choice.

13. Models have personalities: Claude 3.7 is an “ADHD child”

  • He uses Claude 3.7 thinking and Gemini 2.5 Pro Preview most often for coding. Gemini’s “aesthetic judgment is indeed a little better than Claude’s,” but its output is less complex. Claude 3.7 is “an ADHD child”: even after completing the task, it will “deliberately refactor it, optimize it, and add a few more files.” Its instruction following is worse than 3.5’s.
  • He uses OpenAI models less and less, relying on them only for image models. Their output is “serious, extremely mathematical,” and they love tables. This leads to an amusing personality split: 姚行 is a J, using OpenAI intensively because it is “extremely substantive, not chatty, all tables.” Chen Ran calls himself a P who “really needs emotional value”; “tables make me dizzy.” 融汇 sides with 姚行: “When even this kind of chaotic information can be organized into a table, the world suddenly becomes clear.”

14. What Cursor taught him: AI products have 3 ends and 3 edges, and you must first become a creator yourself

  • The framework he thought about most while building Medeo came from Cursor. Traditional products have 2 ends and 1 edge—product and person—and the user journey can be enumerated in advance. With AI, it becomes 3 vertices and 3 edges: person, product, and AI. Cursor clearly understood the progressive allocation of AI permissions: inline tab → edit → chat → agent. “I think it has thought through this model extremely, extremely clearly.”
  • The deeper change is that “when you build an AI product now, you basically have to become a creator yourself.” You pre-build part of someone’s creative process into the tool, “and then the branch conditions can no longer be held in place, because there are infinite possibilities.”
  • He has a clear-eyed view of AI’s creative abilities: “AI actually finds it very difficult to come up with ideas that surpass humans… After testing enough of its scripts, I can predict what kind of script it will write.” AI is best suited not to replacing creativity, but to serving as “an execution role with some intelligence”—writing prompts in batches, drawing cards in batches, writing storyboards with an SOP, and running automatically 24 hours a day. “If you want it to give you a truly aha-moment idea, it’s really difficult… Pushing hard in that direction makes it easy to hit a wall.”

15. Reliable delivery: Finding certainty in randomness; the answer is in Zelda

  • Cursor is the best AI product in his view because it “really delivers a result reliably in production.” That is also the counterintuitive difficulty: “AI products are looking for certainty within randomness… You have to do work to overcome entropy and get a stable result.” After the first experience amazes a user, the expectation becomes, “I hope it’s still this good next time.” But anyone who has built these systems knows that “the probability it will still be this good next time is very low.”
  • His key judgment is that reliable delivery is a product-level problem; “it is not something that requires waiting for a new technology to solve.” Products need time, not another generation of models.
  • His favorite analogy is Zelda: Breath of the Wild and Tears of the Kingdom. An open world does not mean abandoning players to total freedom. “Freedom only becomes meaningful when discussed within certain constraints.” Towers are the major goals, Koroks are the small rewards, and every detour is a deliberately designed reward point. The hard part of an AI product is also designing the positive feedback created by randomness. One day, an AI product will create the same aha moment as the launch of Breath of the Wild. “It may be something we build, or something someone else builds.”

16. Opportunities for young people: Everyone is equal; the race is to iterate faster

  • His circle is full of super-individuals—海星, 阿文 (phonetic), and 倪豪, who recently won Xiaohongshu’s independent developer award. “They all took their own wild paths. It’s absolutely not about something they learned in school.” AI has broken the traditional path of “first learn front-end or whatever,” and “everyone’s route may now be different.”
  • His judgment comes with a conflict-of-interest disclosure—“I’m young, so of course I would say this”: experienced people can be constrained by old frameworks, while “young people are naturally more daring; with skill and courage, they can still build many fun things.” His experience at a major tech company confirmed it: “Because nobody understood it, I became the person who understood it best, even though I was clearly the least experienced… In this new era of AI, I think everyone is equal.” 融汇 added former guest 童超’s point: nobody’s understanding is necessarily further ahead in the AI era; “the competition is about who can iterate faster.”
  • The concrete form of fast iteration is demo culture: improvisational, aimless, and mostly “not very valuable,” but the experience suddenly becomes useful in later projects. He stayed up to watch the Claude 4.0 livestream and tested the API as soon as it launched—“as expected, it was amazing.” When GPT-4o’s image generation first achieved image-to-image consistency, Lovart (phonetic) “took off because of it.” “When a technology appears, you have to test it extremely quickly to be among the first to encounter that insight.” For him, making demos “is a bit like building with LEGO… it’s a way of life.” AI founders also have 海星’s Demo In, where they show each other their work.

17. Creative drive and the downturn: “Healthy crying” and “What does it have to do with me?”

  • His starting point with ChatGPT was 2021, the year he graduated from university at the peak of the mobile internet. “Only after it passed did I realize that was the peak.” In a mature industry with a rigid hierarchy, veterans had the advantage, and computer science was “just a tool for making money” to him. He stayed up playing with ChatGPT on its first night—debates, turtle soup puzzles, emotional companionship—and only later realized that the product itself was a sector. His product insight was: “If the creator finds the process fun… then what they create will also let viewers and users feel that.” Koji contrasted this with his own “Wudaokou era” in 2010, when he was excited about Web 2.0 but felt out of place as classmates joined Microsoft Research Asia and IBM. That experience motivated him to build an AI Hacker House: a place where the next generation of founders can find their peers.
  • His admission about creative drive is the most moving passage of the episode. After graduating, his desire to create disappeared for 3 years: “I knew what I wanted to express… but it was always just a little short.” His defenses include physically separating himself from his phone, living in Pudong for quiet, and reading to replenish his inputs. “I injured myself by watching too many videos before… Douyin is genuinely toxic.” Most of his creative drive “comes from pain.” The day before yesterday, he suddenly cried for the entire high-speed rail journey: “I didn’t know why I was crying, but I felt great. It had finally come back”—a kind of “healthy crying.”
  • He does not avoid the backdrop of the era: “The downturn began with our generation.” Older students found jobs easily; his cohort began struggling to find work. “Just maintaining emotional stability is already quite an achievement.” When he talked about dreams with a close friend, the answer was, “That’s great. What does it have to do with me?” That is why “preserving this passion, preserving this vitality, preserving this creative drive is extremely, extremely, extremely important.” Koji closed with Paul Graham’s How to Do Great Work: find the thing that feels effortless to you but difficult to others, and work at it seriously. The returns will be very good.