Pioneers Insight Method Research Author
123: 5Y Capital’s Meng Xing and Two Challengers: 72 Hours with Only AI
Back to Episodes

123: 5Y Capital’s Meng Xing and Two Challengers: 72 Hours with Only AI

Summary

  • The clearest takeaway from the 72-hour experiment is that model capabilities have outpaced the digital infrastructure built for Agents. Under conditions without smartphones, ready-made browsers, or complete development environments, the 7 participants ultimately used AI to get through CAPTCHAs, payments, deliveries, and food orders—but only through a maze of scripts, environment variables, and semi-manual workarounds. Meng Xing put it most sharply: “An Agent has to kneel before the CAPTCHA and pretend to be a person.” Authentication, payments, authorization, and machine-native interfaces are opportunities for Agent infrastructure, but also battlegrounds involving security, regulation, and platform interests.
  • AI coding has pushed the cost of starting a personal product low enough to generate real cash flow, but it has not eliminated the friction between a demo and a reliable product. Chen Zhiyue, who could not independently complete a computer-science assignment, built Adventure Snap in 3 months, generating more than €1,000 a month in subscription revenue from European users; but he had expected to complete 70-80% of his virtual livestreaming product in 72 hours and ultimately managed only about 30%, with a single WebSocket problem consuming 7-8 hours. “AI is still solving the productivity problem, not really the competition problem.”
  • A one-person team may start more easily, but that does not mean the winners will remain one-person companies forever. AI lets non-programmers such as product managers and investment managers “get in there and build something themselves,” but once a product is validated, ASO, promotion, engineering depth, and competitive follow-through still push teams to expand. Meng Xing believes the long-term equilibrium “may not be that small”; small, beautiful businesses will exist, but low development costs cannot be extrapolated into eternal one-person unicorns.
  • AI imaging is reversing the creative workflow, but the director’s authority has not been taken over by the model. Lijianlei has shifted from “story first, visuals second” to having AI generate images in bulk and then assembling them into a story; meanwhile, a 5-second “masterpiece-level” Kling shot costs about RMB10 and may be entirely unusable, while voice acting, emotional intensity, and precise composition remain constrained by human performers and traditional post-production tools. His conclusion is that AI cannot create “something from nothing”: the director must provide the motif, style, and boundaries, and steer between ambiguous emotion and the one precise image.
  • Extreme isolation showed that demand for AI companionship does not belong only to the imagined category of “lonely users.” The extra request participants made most often was not more tools, but the chance to talk with the other contestants; a warm sausage, mutual help on the message board, and the urge to hug people after leaving were more powerful than any feature. Chen Zhiyue also went from never discussing emotions with AI to interacting continuously with his virtual-livestream demo for half an hour; with model commentary costing about RMB0.02 a minute, he described the experience as “not livestreaming as an internet celebrity, but livestreaming as a star.”
  • In the AI era, what may be scarce is not executors but domain experts who can ask questions beyond the boundary. Meng Xing initially placed more value on specialist researchers and fast young people, but after the experiment concluded that people without computer-science backgrounds may be the ones who push the boundary of knowledge outward, while machines are better at filling gaps between existing points. “The most important thing people do is ask the right questions, not answer them.” Artists should not merely give engineers feedback; they should define tools and works together with engineers.
  • Meng Xing sees multimodality and world models as variables worth tracking in the next phase, but not yet productized or settled. Odyssey has demonstrated real-time generation of realistic scenes that let people walk through and view generated spaces, but meaningful interaction remains limited; the more distant goal is to turn a film, or even the entire world, into an enterable parallel space. The real challenge is not image quality alone, but unifying understanding, generation, and alignment—and encoding pixel and spatial information that humans have never structured in advance.

Deep dive

1. AI coding’s revenue curve surged first, while product experience remained one step short

  • Meng Xing dates the experiment’s starting point to February, when AI coding products such as Cursor, Bolt, and Windsurf were gaining strong traction and rapidly growing revenue and user numbers. The media began speculating about “one to 10 people building a super-unicorn,” including stories of companies reaching several million dollars in ARR within a few months.

  • Actual product experience at the time was less smooth. Products such as Bolt and Rosebud, focused on front-end or game generation, were still “one step short” for even simple personal projects; just over 4 months earlier, Cursor was still commonly described as a tool for professional programmers to vibe-code a demo in 2-3 days, not for large-scale production deployment.

  • The picture was changing fast enough to invalidate old judgments almost immediately. By the time of the recording, Meng Xing already saw Cursor as an important force in enterprise development, making a real-world experiment more valuable than asking founders to describe the limits of their capabilities.

2. What 5Y Capital really wanted to test was who could become a creator

  • In conversations with teams such as Augment Code and Windsurf, Meng Xing heard different versions of the future. One path was to plug into enterprise systems and expand the productivity of professional developers; another, like Replit, was to move from “tools for you to play with,” such as H5 games, toward genuinely complex works—the extreme case being an individual building a GTA-level product.

  • He reduced the disagreement to 2 questions: What tasks can AI help people complete today, and what will it be able to complete in the future? The more important investment question was whether AI was merely a coding-assistance tool or a general-purpose creative tool that could help people who knew code, and eventually people who did not, build influential products and works.

  • 5Y Capital also recognized that it was living inside a “bubble,” interacting daily with founders who understood AI best. The experiment was therefore designed to broaden the sample to ordinary people and understand how the technology actually changed their lives, rather than continuing to infer mass-market capability from industry insiders’ accounts.

3. The 1999 Internet survival challenge became a 3-layer AI experiment

  • Meng Xing first encountered the Macintosh at just over 3 years old and began using the Internet in China in 1995, making him unusually sensitive to how each generation of technology entered people’s lives. The 72-hour Internet survival challenge in 1999 left a deep impression on him at age 14: Taobao did not yet exist, while e-commerce, payments, food delivery, and courier services were all extremely immature.

  • That challenge tested whether weak Internet infrastructure could sustain physical survival while people still had computers and other tools. 26 years later, the infrastructure had matured, so the new question was whether the agent behind the computer could switch from a human to AI—or whether humans could use AI to complete actions.

  • The new experiment followed Maslow’s hierarchy and set 3 levels of objectives: Could participants survive using only AI? Could they obtain more resources than they started with and improve their living conditions? Could they complete an application, piece of content, or other form of self-actualization beyond survival?

  • Meng Xing believes that if AI 26 years later can still solve only the problem of physical survival, “then it may not have progressed much.” The real question is whether it can help with thought, spirit, and creation—even “ignite your soul.”

4. RMB100, a bottle of water, and a book reduced resources to the bare minimum

  • Each participant received only RMB100, a bottle of water, and one chance to choose a single item from options such as snacks and books. By comparison, the 1999 challenge provided roughly RMB1,000 in cash and RMB1,500 in virtual tokens—RMB2,000-3,000 in total—with much greater purchasing power at the time.

  • Lijianlei chose My First Programming Book because its cover promised to be suitable for beginners. He could not understand the second page and could advance only about 1 page every 5 minutes. He later admitted that learning programming from the book within 72 hours “was obviously never going to work”; the book ultimately became a mosquito swatter.

  • The 2 university students, both veterans of multiple hackathons, were attracted by the scarcity. Traditional hackathons arrange food, drinks, and sleep so people can focus on development; 5Y Capital “provided absolutely nothing,” which they considered “too cool.”

  • The rules were designed to minimize physical supplies and pre-existing tools while making the AI options sufficiently broad. The test was not who knew existing apps, but whether participants could use AI to rebuild the interfaces they needed.

5. Only 2 months of technical rehearsals proved the challenge would not become simple starvation

  • From March to May, the team spent most of its time testing IT feasibility. They initially banned phones entirely, but quickly realized that not receiving verification codes meant being cut off from the infrastructure of real society; the final compromise was a feature phone that could receive codes but could not run apps.

  • Payments and ordering also required fallback routes. Participants could ask AI to write a browser and use browser-use to build a small Agent that clicked through websites, or open a browser from a coding tool and complete the process semi-manually. The prerequisite was always that the target website allowed the chain to work.

  • Manqi summarized the rules precisely: Without a smart mobile device or a ready-made browser, how could AI be used to reconnect to Internet services? Meng Xing added that the experiment removed every “tool someone else had already programmed” and asked one person to rewrite, within hours, a capability built by a company of 10,000 people.

6. This was not project selection disguised as an experiment

  • Manqi’s first instinct was: “Why not spend this time looking at more projects?” Meng Xing’s answer was that nothing in the world is inherently something one “has to do.” Entrepreneurship, investing, and experiments can all be “a way of asking the world a certain kind of question”; execution is simply how one answers it.

  • He believes that as AI becomes more powerful, people’s core role moves closer to asking the right questions rather than answering them personally. 5Y Capital wanted to know how far AI’s impact on ordinary people had reached; if designed well, the experiment could also help more people understand that “it has reached this point” or “it has not reached that point yet.”

  • 5Y Capital’s founding partners supported the idea. The firm had previously taken portfolio-company founders to an island for a physical survival challenge, but the one objective deliberately excluded this time was using the event to screen projects, find founders, or source deals: “We wanted to do something nobody else was doing.”

7. The 2 challengers arrived with completely different questions

  • Born in 2000, Chen Zhiyue had no memory of the 1999 challenge and had never systematically considered how society should meet new technology. He already used AI naturally for work, daily life, papers, and products. He had even noticed his mother using Doubao earbuds every day to ask about the stock market and stock recommendations, but had never elevated those fragments into a question.

  • When he applied, he did not even understand what the challenge was supposed to involve. He mainly wanted to see what answers other people would produce. He had long had the idea for virtual livestreaming but had never written the code; the application and subsequent interviews became a trigger, forcing him to think through the product form, show format, and technical route.

  • Lijianlei grasped “72 hours” and “a closed space,” but not the word “challenge.” His initial response was, “I’m here to take a break,” hoping to force himself to pause. During the film and television downturn, he had considered changing careers, but an AI work shortlisted for the Golden Rooster Awards “pulled him back into AI,” so he also saw the experience as “coming to pay homage to AI.”

  • Lijianlei had already monetized AI-image advertising and training, but every previous exploration had a known premise. This time he knew only that he wanted to make a film, without setting the content or theme in advance, hoping to “dance in shackles” under unknown conditions.

8. More than 300 applications showed that curiosity belongs to no single age group

  • The organizers initially worried that nobody would want to participate, but ultimately received more than 300 applications, spanning people in their teens to people in their 60s. The number of applicants, their diversity, and their motivations all exceeded expectations.

  • A lawyer in her early 50s had no programming or AI skills, but was prepared to “go hungry for 72 hours.” She was simply curious about what young people were doing and wanted to see for herself “how you live.” Another applicant hoped to make something with her child, and the team even discussed a parent-child challenge.

  • The 2 university students applied separately. During interviews, staff discovered that they knew each other and were already starting a company together, so the organizers formed them into one team at the last minute. Recruitment itself showed that the same “AI survival” title could trigger entirely different forms of self-testing.

9. On the first night, one person slept while another coded past 3 a.m.

  • The challenge began at 6 p.m. While everyone else quickly started writing code, Lijianlei meditated and then fell asleep. He wanted to create a buffer between the noisy real world and the closed space: “I can be very far from this world, but I hope to be closer to myself.”

  • He did not know that orders had to be placed before 11 p.m. When he woke at 10:30, only 30 minutes remained. Lijianlei admitted that if he had known the deadline in advance, “I definitely wouldn’t have let myself sleep.” The nap was not a strategy; it was the real consequence of misunderstanding the rules.

  • A doctoral student returning from the Middle East woke at about 3 a.m. because of jet lag, quickly began coding, and posted the result on the message board. Meng Xing liked the contrast: If the experiment became a race in which everyone scrambled toward a target, it would be nothing more than a utilitarian hackathon.

  • Participants’ sense of time was also split by their survival state. Before hunger and ordering problems were solved, “every second felt like a year,” as if they had entered another dimension; once food was secured, attention returned to works and products.

10. The first survival route was not food delivery but next-day courier delivery

  • Chen Zhiyue’s first project was a next-day courier tool that allowed orders to be placed from a webpage. He assumed payment would be the hardest part, but linking the bank card to an existing account and entering the verification code received on the feature phone was enough to complete the payment.

  • He coded through the night from 6 p.m. on the first day and did not see the payment confirmation and delivery notice until around 5 a.m. the next morning. The excitement of “I actually made it work” came half from closing the technical loop and half from the certainty that he would not go hungry.

  • He ordered a case of bottled water, a case of toast, a case of chicken breasts, and several tea eggs. He had also planned to order salad and vitamins, but the actual selection was limited: “Getting anything through successfully was already good enough.” Real-time food delivery was not unlocked until the afternoon of the second day.

11. 24 hours of hunger turned code into Lijianlei’s antagonist

  • Lijianlei downloaded scripts shared on the message board but did not know which window to paste them into. When the system said Python could not be found, he could not tell whether the problem was the version, the path, or a dependency. He handed the code to DeepSeek, copied each terminal error back into the chat, and did nothing more than press Enter again.

  • The loop of “error—copy—ask—another error” continued until he had gone more than 24 hours without eating. He eventually used his one technical-support request to have the organizers configure the environment variables, which finally allowed him to place an order.

  • When Manqi asked how he saw AI at that moment, the answer had none of the usual romanticism: “It’s just so annoying.” He knew every English word, but putting them together was maddening. The raw emotion of wrestling with code became the first motif of his short film: “I want to fight the algorithm.”

12. A bowl of instant noodles and a sausage restored warmth to the world

  • Lijianlei’s first meal was instant noodles and a piece of chicken breast. He was too hungry to work, and for the first time in front of the camera felt that instant noodles tasted extraordinarily good. The most profound moment had almost nothing to do with AI; it was the relief of physical deprivation.

  • On the afternoon of the second day, one participant first ordered milk tea for everyone and then shared a food-delivery tutorial. Chen Zhiyue used it to order one sausage for each person. Even though the sausages had cooled somewhat by the time staff distributed them, they were still warm inside: “That bite was so warm—it had the warmth of people and of this world.”

  • He remembered the crisp casing, soft interior, and evenly distributed seasoning, and developed a strong liking for the delivery brand he had ordered for the first time. An extreme environment magnified sensory and emotional value that would rarely enter ordinary product analysis.

13. Separate rooms did not eliminate social interaction; they made connection scarce

  • Before the challenge began, staff were still in the rooms setting up equipment. At 6 p.m., everyone left and the doors closed, leaving Chen Zhiyue suddenly alone. He kept asking on the message board, “Is anyone there?” and “Does anyone want to chat?” He struggled to adapt to a state with no real person speaking.

  • Participants discussed rolling up a blanket or towel and lowering food from upstairs to people below. The organizers confirmed that this violated the rules, so the plan was not implemented. What mattered to Lijianlei was not the outcome, but the “flow of love” that appeared in the discussion; it became the second motif of his short film.

  • In the final 10 minutes, Chen Zhiyue had to shoot a vlog, give interviews, record the countdown, and clean his room, leaving him exhilarated and untethered. When he rushed downstairs and saw the crowd, he wanted to talk to and hug everyone: “It felt like returning to this world.”

  • Manqi’s summary received his immediate confirmation: people need to be with real people. Virtual feedback can supplement emotion, but it did not make real connection less valuable under extreme conditions.

14. What the organizers ultimately saw was “The Martian + Escape + collective creation”

  • Meng Xing had worried that some participants would fail to survive and drop out, or that everyone would pursue only their own goals and ignore the message board. In practice, the page kept refreshing, while mutual help, conversations, script sharing, and attempts to exploit loopholes supplied much of the experiment’s energy.

  • After the event, the organizers asked each person what other help they needed. The most common answer was not more compute or more tools, but the chance to meet and talk with the other participants. Meng Xing realized: “Although everyone was essentially challenging themselves and AI, connection turned out to be so important.”

  • He described the scene through 3 images: The Martian, with people saving themselves using limited resources; “Escape,” with participants passing things by bedsheets and searching for loopholes; and people spending long periods creating with AI in their rooms, resembling Chen Jingrun working alone to prove a conjecture. The experiment did not converge into a single genre.

15. AI capability changed scale several times between planning, execution, and review

  • In February, the discussion was still centered on AI coding; in March, Manus and a wave of AI Agents appeared; by May, general capabilities had improved enough that some of the original scaffolding could be relaxed. Meng Xing therefore repeatedly emphasized that this was not a static measurement capable of answering the question permanently.

  • Based on the results, surviving with AI was broadly feasible, but not naturally accessible. Participants used hacker-style routes to bypass environmental constraints; some e-commerce sites worked and others did not. The result could not be summarized as the Internet either supporting Agents or not supporting them.

  • The split in outside reactions also showed that people were operating at different cognitive scales. Some thought it was “incredible” that ordinary people could do these things in 72 hours; others thought opening a coding tool and having it call up a browser would take only a few dozen minutes, making the event “extremely boring.” Meng Xing believes the disagreement itself was part of the experiment’s value.

16. The Internet was designed around phones and people, making it difficult for Web-based AI to reach offline services naturally

  • Today’s payment, CAPTCHA, app connectivity, and account systems were built mainly around phones over the past 12-13 years; older PC-web structures remain underneath. In both generations of systems, the default actor was a person, not an Agent.

  • Manqi pointed out that the new wave of AI tools is concentrated on the Web and in the cloud. Because of model size, latency, and other constraints, they may not land on mobile first, while the offline services needed for survival—food delivery and retail—are heavily concentrated in mobile apps. A natural gap exists between the two.

  • The 2 university students also observed that existing systems depend too heavily on specific terminals, especially phones. The challenge did not really test whether a model could click; it tested whether model capability, Web entry points, mobile services, identity verification, and payments could be assembled into one continuous chain.

17. Agents can already work on behalf of people, but have “no dignity” in the digital world

  • Meng Xing described Agents as high-concurrency actors capable of replacing people, but said that “in their world of survival, they have no dignity.” Every CAPTCHA, ban, and identity check is a shackle: “An Agent has to kneel before the CAPTCHA and pretend to be a person.”

  • Most Agents today still use a human’s browser and GUI, imitating clicks, logins, and payments through computer use rather than operating in their own sandbox and machine-native environment. They are not people; at best they are authorized virtual people, and they may not even have completed the authorization process.

  • Genuinely independent Agents will need accounts, wallets, permissions, and cross-service identities, and one Agent may hold more accounts than one person. The capability exists, but the supporting identity system is not ready.

  • Platforms also have legitimate reasons to resist. If an aggressive Agent impersonates a real person, regulation and accountability become extremely difficult. Meng Xing said this could become a commercial battle; the program also noted that platforms in different competitive positions may make different choices about whether to open access.

18. From co-pilot to auto-pilot, platforms will shift from selling attention to selling outcomes

  • Manqi identified another conflict: many Internet services monetize human attention and advertising revenue. If users stop browsing personally and simply send machines to retrieve results, their existing business models face a direct shock.

  • Meng Xing distinguishes between co-pilot and auto-pilot. The former requires people and AI to collaborate continuously in discovery and decision-making, leaving human attention in the loop and potentially supporting both eyeball-based and outcome-based billing; the latter completes tasks automatically, in which case “the eyeballs are gone” in theory and pricing is more likely to shift toward outcomes.

  • Agent infrastructure is therefore not merely about adding an API or CAPTCHA component. It also touches platform traffic allocation, ad display, service pricing, risk liability, and who has the right to spend on a user’s behalf.

19. Machine-native interfaces and human understanding are both lagging behind model capability

  • Browser graphics were designed around the fact that humans have eyes, but machines do not need the same pages. Meng Xing noted that people reach their limit at roughly 2.5x or 3x video speed, while an Agent might process content at 300x; existing players were never designed for that mode of consumption.

  • He breaks the evolution into 3 lines moving at different speeds: technical capability rises exponentially; infrastructure supporting the technology moves more slowly; human understanding of what the technology can do moves more slowly still. “Each one is slower than the last,” and every layer loses something between them.

  • The activity tested the second line by identifying where infrastructure remained deficient, while also helping the third line catch up—showing people where AI “has already reached” and preventing capabilities that do not yet exist from being mistaken for reality.

  • Meng Xing rejects a binary story about Agent infrastructure in which “without it, nothing works; with it, everything works.” The reality is that legacy systems are compatible in some places and fail in others. Only running a complete task reveals the middle ground.

20. AI did not turn individuals into islands; it generated new collaboration networks

  • The purpose of placing participants in separate rooms was to create a tighter connection between each person and AI. That connection did form, but participants also crossed the physical barriers through the message board, scripts, tutorials, and food plans, creating a temporary network.

  • Meng Xing therefore imagines that AI may not lead to personal isolation. As individual productivity explodes, the ability to connect and combine with others may strengthen at the same time; independent development, joint development, and collective problem-solving can coexist.

  • He sees protocols such as A2A and MCP as early prototypes of this model, and thinks back to the Internet protocols of the Tim Berners-Lee era: people create independently, then connect their output through shared rules. The eventual organizational form is uncertain, but the “networked super-individual” is closer to what happened in the experiment than the purely isolated individual.

21. 7:41 grew out of fighting the algorithm, the flow of love, and the sound of an airplane

  • Lijianlei deliberately reserved the first 48 hours for experience, hunger, meditation, frustration, and emotional accumulation, leaving roughly the final 24 hours for production. Wrestling with code produced “the algorithm’s coldness,” while mutual help on the message board produced “the preciousness of human emotion.” The 2 motifs emerged one after another.

  • The guesthouse was near Shanghai Pudong Airport, with planes repeatedly passing overhead. Their sound reminded him of the disappearance of MH370, completing the third motif. The story shifted toward a person living with AI while being unable to make AI understand the memories of a missing loved one.

  • In the film, AI considers the protagonist’s years of voice memos from his wife and the sound wave from 2 seconds before the plane crashed to be meaningless white noise that should be deleted. But the alarm set for 7:41 every morning keeps ringing stubbornly on time, giving the film its title, 7:41.

  • DeepSeek helped expand textual details, such as coffee spilled in the shape of the Strait of Malacca. Such associations went beyond Lijianlei’s expectations and made him acknowledge how much AI helped complete the writing.

22. AI can expand a story, but it cannot complete the director’s leap from emotion to one unique image

  • Once the text entered the visual stage, Lijianlei immediately hit a ceiling. Precisely turning liquid into the shape of the Strait of Malacca required Photoshop, image-editing software, or After Effects, but the rules removed those traditional tools. Directorial authority was therefore forced to cede ground to the generative model.

  • Voice acting exposed a similar loss of emotion. Text-to-speech could not reliably handle Chinese heteronyms, questions, affirmations, or strong exclamations. He ultimately used his own voice to express the emotion first, then applied voice-to-voice conversion, which brought the result somewhat closer to the target.

  • In his normal work, he still prefers human voice actors and relies on sound libraries and music to adjust emotional intensity. The 72-hour experiment, without those people and Internet resources, showed that a director cannot become a fully self-sufficient “super-individual” using generation alone.

  • The boundary Lijianlei ultimately drew was that AI cannot reliably create “something from nothing,” nor automatically grasp human emotion. The director must first provide the motif, framework, and style, then let AI execute within those boundaries—not simply say, “Write a story about a person coexisting with AI.”

23. Free compute enabled creative freedom while exposing the real cost of generative imaging

  • The event provided image-generation tools such as Jimeng and Kling at no cost to participants. Lijianlei called the situation “the creative freedom that algorithmic freedom gives me”: normally he would have to distinguish between standard and master models, but here every shot could use the master-level model.

  • His estimate of the everyday cost was about RMB10 for a 5-second master-level Kling shot, with any single generation very likely to be unusable. That is already far cheaper than live action, special effects, set construction, and actors, but ordinary creators may still be unable to absorb the sunk cost of repeated “draws.”

  • When Manqi asked about industry practice, Lijianlei cited the work that had been shortlisted for the Golden Rooster Awards: shoot live action first, then use AI for frame-by-frame transformation to preserve consistency in the characters, space, and image structure. Pure generation makes it difficult to ensure that every element in the frame was designed by the director. “Generation right now is basically drawing lots.”

24. AI inverted “story first, visuals second” into “look at the visuals first, then find the story”

  • Lijianlei now gives himself 2 hours to have AI generate a huge number of images without first conceiving a plot. He then selects and assembles them, observing what story the images can form. “Before, there was a story first and then the visuals. Now I have AI generate the visuals first, and then I have the story.”

  • He plans to continue developing 7:41, expanding the 72-hour version into a medium- or feature-length film. It will continue to explore people and AI, folded time, and streams of consciousness, perhaps release by the end of this year, and continue to be made with AI.

  • AI works once pulled him back from the edge of changing careers in film and television, and this challenge again suggested a chance to “overtake on the bend.” He wants to explore the distinctive style of a “first-generation AI director,” or China’s “seventh-generation director,” but did not package an incomplete work as a validated success path.

25. Chen Zhiyue did not learn traditional programming; he learned to use projects to drive AI to build products

  • Around June 2024, Chen Zhiyue, then working as a product manager, often faced engineering teams with 10-15 years of experience. Their response was: “You don’t understand code, so stop giving stupid orders.” When he asked why a requirement could not be built, they would end the discussion with technical details.

  • He initially followed an Indian programmer on YouTube, typing code step by step from Hello World. After 2 days, he decided the path was too stupid: “I’ll never learn to code in my life.” The turning point was not mastering syntax, but switching to AI coding and first defining the product he actually wanted to complete.

  • He chose a full project spanning the UI, front end, back end, deployment, App Store review, and overseas operations, rather than a to-do list or clock. About 3 months later, Adventure Snap went live, using a VLM to identify buildings in front of travelers and having a model provide historical and cultural explanations.

  • He has always used Cursor, even though early models and context capabilities were far weaker than today. He makes no attempt to hide the dependency: “Although I’ve developed several products, without AI coding I have absolutely no way to complete any computer-science course assignment.”

26. What people will pay for in Adventure Snap is not recognition, but explanations of details in front of them

  • The product initially answered only “What is this?” and provided basic background. After iterating, the team realized that users standing in front of attractions such as the Louvre often already knew the general facts; the paid value lay in pointing out specific details they could see but did not know how to interpret.

  • Chen Zhiyue used Hagia Sophia as an example. Users might know that Christian and Islamic histories overlap there, but the product should identify the exact symbols on the door, where they appear in the photo, and why they prove that history exists.

  • The team also connected Apple’s latitude and longitude data to help determine the user’s location, while distinguishing encyclopedia-style knowledge from anecdotes and curiosities. For Bangkok’s Grand Palace, the point was not simply to state its name, but to explain its “heart” position in local feng shui narratives about Bangkok and the palace.

  • He describes the goal in everyday terms: the information should give users something to “brag about” after they return from a trip. The 2 rounds of iteration were not about producing more words, but about matching content to the user’s current field of view, location, and social needs.

27. Someone unable to complete a homework assignment independently built more than €1,000 in monthly subscription revenue

  • Adventure Snap is now live on overseas App Stores. Chen Zhiyue posted only 2-3 introductions on Reddit, yet attracted backpackers from France, Germany, and elsewhere, and continues to charge through a membership model.

  • He disclosed that monthly revenue was “roughly above €1,000.” The product offers a 1-week free trial, and users who stay on as paying customers form a stable membership base. Apple exposure, old posts, and social-media recommendations continue to bring in new users, while normal churn remains; overall growth is slow but steady.

  • Faced with a US competitor called ChatCici, Chen Zhiyue does not think the 2 products target exactly the same user. ChatCici aims to cover travel, daily life, education, and other scenarios, approaching a search engine on the app; Adventure Snap serves only people willing to understand a destination in depth.

  • He acknowledges that many travel lovers do not need history or details and only want to check in, take photos, and read a basic introduction. The product is not trying to serve everyone; it wants to serve a tiny group of deep travelers well, then iterate further when it has the capacity.

28. Chen Zhiyue expected to complete 70-80% of the virtual livestreaming product, but finished about 30%

  • Chen Zhiyue did not want to build a virtual streamer. He wanted a real person to livestream to a group of AI viewers, allowing ordinary people without followers to share their feelings and show their lives while receiving timely feedback in a friendly, safe comment environment.

  • The basic chain seemed simple: ASR would convert audio into text, then an LLM would generate comments. Images could be added selectively, while full video understanding would be deferred because multimodal processing would bring higher latency and cost, and the first version did not need every piece of information.

  • The real difficulty was audio and video transmission. He had no previous experience in the area, and different components and cloud services required different file formats. Unable to check documentation, communities, or the Internet, he spent 7-8 hours on a WebSocket connection problem and could not debug it by sleeping and trying again the next day.

  • He had initially expected to complete 70-80%, but ultimately rated himself at only 30%. The demo ran on a computer without camera permission, leaving the screen black and allowing it to receive audio only; the AI could respond, but he had also hard-coded a large number of canned banter comments.

29. AI comments costing RMB0.02 a minute were cheaper and more alive than prewritten banter

  • When he continued development after the challenge, Chen Zhiyue changed his assessment. Hard-coded comments seemed to save money but were actually noise because they could not respond to what the streamer had just said or naturally enter the next round of dialogue.

  • He measured the model-comment cost at roughly RMB0.02 per minute of livestreaming. The interface expanded from the usual 4 comments to 6, updating continuously as they scrolled; likes and the number of people entering the room were also used to create the sense of an audience, with a possible expansion to 7-8 comments later.

  • When he recorded a 5-minute demo himself, he did not need to prepare topics in advance. He read one comment, answered it, let the new content re-enter the model, and received the next response; the loop formed naturally, allowing even someone without livestreaming experience to imitate the state of a normal livestream.

  • Everyone asked questions and offered encouragement around the streamer’s small emotions and movements, pushing the streamer’s agency to the foreground. Chen Zhiyue described it as: “Not livestreaming as an internet celebrity, but the feeling of livestreaming as a star.”

30. Extreme conditions turned a product creator into the user he had once looked down on

  • Chen Zhiyue reflected that he initially had a “very arrogant” product mindset. He assumed the product served people unable to obtain timely emotional feedback from real life, while he could always find friends to talk to and therefore did not belong to the target user group.

  • Once the 72 hours pushed his conditions to the limit, he discovered that he also needed AI feedback and could hand part of his emotional needs over to a machine. After returning to ordinary life, he interacted continuously with the demo for half an hour and believed for the first time that he could keep using it over the long term.

  • He had previously downloaded Character AI, Talkie, Xingye, and Zhumengdao only for research, never used them frequently, and did not understand why users needed them. After becoming a demand-side user himself, he finally saw how friendliness, emotional modes, and interaction pacing should be adjusted.

  • The product roadmap therefore added emo and non-emo modes. When users are emotionally low, the system provides almost entirely positive feedback; when they want to simulate a real livestream, it allows absurd, off-topic, and less-positive comments. The next step is a TestFlight beta; the full launch depends on Apple review and algorithm filing requirements and may take another 1-2 months.

31. AI gives more people products, but does not automatically build them a moat

  • After Chen Zhiyue shared his development process on Xiaohongshu, people from Hong Kong investment banks, major tech companies, and investment firms asked about his scaffolding and development approach. His observation is not that everyone wants to become a professional indie developer, but that “everyone wants to build their own product” to test a need they have glimpsed.

  • He maintains multiple products at once. An independent product can generate some cash flow to cover living and development costs; there is no need to raise money first, and an idea can quickly become a usable version before the creator decides whether to invest further.

  • Meng Xing agrees that AI has dramatically reduced the size of the starting team. Products that once required several people can now be validated by 1 or 2. But once a direction works, competitors enter, requiring stronger ASO, promotion, engineering, and more features; marginal iteration still demands more people and capital.

  • His long-term view is that teams at a stable equilibrium “may not be that small.” Small, beautiful businesses will exist in the long term, but most successful tools will enter a red ocean. “AI is still solving the productivity problem, not really the competition problem.”

32. The next organizational form remains unsettled because people keep raising their demands after AI meets them

  • Meng Xing admits that 1-2 years ago he was more convinced that AI would create extremely small new organizations, and more worried about wealth concentration: people with the most capital might command nearly unlimited Agent intelligence, while ordinary people might not have equal access.

  • The program raised one possibility: If Agents became powerful enough, the conversion loss between capital and productivity could fall, perhaps even allowing Agents to start companies and run projects. Meng Xing’s rebuttal was that this assumes human demand remains static. Once AI satisfies existing demands, demand will continue to rise, and new demands are often precisely what AI cannot satisfy yet.

  • Autonomous driving will not end demand for transportation. Users will still want different driving styles and car-to-car connections, eventually even pursuing “teleportation.” Zoom is a substitute for travel, but it is not teleportation. On whether AI will reshape company organization, he kept his answer open: “I think we’re watching. I don’t know the answer.”

33. The next major question may come from world models that can be entered, watched, and eventually interacted with

  • Meng Xing does not plan to simply replicate the same challenge. The current work still includes organizing the 72 hours of footage from the 7 participants and producing individual short films; if the experiment returns, its format and question must follow the new technological moment rather than turning the first edition into a fixed show.

  • He is currently focused on how far multimodality and world models can go. Odyssey has demonstrated real-time video generation of realistic scenes, forming a live generated space from a photograph and a real environment that people can walk through and view. Compared with World Labs and Decart’s Oasis, the real-time quality, real-world setting, and clarity of rendering are particularly notable, though interaction remains limited.

  • The more distant vision is to turn a film into an enterable parallel world, similar to the reconstruction of a film in the third episode of Season 7 of Black Mirror. If the context expands from films to the real world, the range of experiences available to users would grow dramatically.

  • The challenge is to unify understanding, generation, and alignment. Humans have already structured text through language, but pixels, video, and spatial information lack an equally mature tokenizer. Existing video companies focus more on realism, consistency, and impressive effects; physical plausibility matters only in some settings, and the frontier has not converged.

34. All 3 participants ultimately redrew the boundary between human and AI roles

  • Lijianlei concluded that directors must control the invariants: color, tone, style, camera parameters, and even color codes. What the characters do and how the subjects change can be delegated to AI. Prompts must be literal and unique: “black means black, 5 hands means 5 hands, looking up means looking up,” without relying on the model to understand vague lyricism.

  • Chen Zhiyue’s view of AI coding did not fundamentally change; what changed was his understanding of emotional products and users. He moved from believing that “AI only needs to help me solve problems efficiently” to acknowledging that he, too, could receive emotional support from a model.

  • Meng Xing saw that AI remains a long way from fully taking over online and offline life, but that the value of people without computer-science backgrounds is underestimated. Machines are good at filling gaps between points already inside the human knowledge sphere, but struggle to “expand the sphere another circle outward.” Artists and ordinary users may be the people who add new points beyond the boundary.

  • Artists therefore should not merely serve as a Value Network or environmental feedback for an engineer’s proposal. They should become a Policy Network alongside programmers, jointly defining the problem and the work. Lijianlei further predicts that AI will accelerate Chinese film and television’s shift from a director-centered system toward a producer-centered one, and that only the combination of programmers with traditional visual-media workers can produce a more standardized super-entertainment individual.