Pioneers Insight Method Research Author
VAST’s Song Yachen: Language Models Hit a Wall; 3D Models Get Started
Back to Episodes

VAST’s Song Yachen: Language Models Hit a Wall; 3D Models Get Started

Summary

  • VAST’s core view is that language models have already “hit the wall,” while 3D foundation models remain in a phase of rapid iteration, putting up a “new wall” every 3-5 months. Song immediately narrowed the claim: “It’s not that all models have hit the wall; I’m saying language has.” In fields where models keep leapfrogging, a pure-play application company may see its work erased by the next release just after it has “protected the old wall”; VAST works on both foundation models and workstations, giving it a view of which gaps the next version will close and which are worth productizing.
  • Tripo 3.0 crossed not the threshold for better demos, but the “pipeline ready” threshold at which it can be used directly in most scenarios. Song said users can now send generated models straight to a 3D printer without understanding structure, file formats or DCC software; after 1.5 years on the market, the product has reached roughly 3M-4M professional creators globally and more than 40,000 enterprise customers, including more than 700 major accounts. Only Standard has been released so far; Ultra, which will deliver better results at the cost of longer generation times, is still to come.
  • Tripo Studio shows that getting a model company closer to users does not necessarily dilute investment; it can instead close the loop between models, products and revenue. The AI-native workflow launched on May 31, after which company revenue “more than doubled”; the host then summarized that Studio already contributes more than half of revenue. It turns semantic segmentation, part completion, universal rigging, low-poly generation and Magic Brush into a single workflow, allowing “80-point assets” that once required professional refinement to be edited further at much lower cost.
  • VAST’s clearest current moat claim rests on roughly 40M high-quality native 3D assets, 50-60 Tsinghua PhDs, and a lead in funding and valuation within the sector. Song said competitors’ datasets are generally measured in the millions, while VAST’s reach “the level of Sun Wukong in Black Myth: Wukong”; the company has raised three rounds, each worth tens of millions of dollars, and has more than 110 employees. He still stressed: “A larger caravan does not necessarily find the oasis faster.” Beyond scale, route selection and luck matter.
  • The endgame is not selling a 3D tool, but building a zero-barrier, zero-cost, real-time 3D UGC network. Song points to Honor of Kings sustaining roughly 100M DAU for ten straight years and a global games market of roughly $260B as evidence of demand, and makes the bold call that future interactive 3D platforms could reach 2-3x the combined scale of today’s major social and content platforms; every game today would then be only “a small part of the gene pool” of the broader category.
  • At this stage, commercialization is fundamentally about the product, not stacking up salespeople, advertising and dinners. Song’s conditional view is that as long as professional users do not face a meaningful information gap and the product remains sufficiently differentiated, improving the product will be more effective than adding BD; once products commoditize or begin serving PUGC and UGC users who do not actively follow 3D models, growth and branding become necessary capabilities. Its CEO Program has interviewed roughly 1,000-2,000 real users, with demand expanding from games into industrial design, graduation projects, disability expression and XR applications.
  • The larger worldview is that the internet is gradually “decompressing” from text, images and video back into original 3D files, while human value shifts from physical production toward content and experience. In Song’s vision of a “fourth major industry,” the value created by one person is measured by the total time everyone spends in that person’s content and worlds; wealth ultimately takes the form of compute, recommendation capabilities and more creative agents. For Song himself, virtual worlds buffer the pressure of entrepreneurship: “When the physical world accounts for only 40-50% of your life, at least half of it doesn’t have to hurt.”

Deep dive

1. Song Yachen Went from “Building Worlds” to 3D Foundation Models—Not by Chasing a Trend

  • Song Yachen is 28. He earned his undergraduate degree from Johns Hopkins University, studied Hebrew and Arabic early on, then moved into AI; he worked on AI animation and games at SenseTime, helped found MiniMax in 2021, and started VAST in 2023.

  • According to Song, VAST has completed three financing rounds, each worth tens of millions of dollars, and has more than 110 people in total. Asked about revenue and profit, he gave no specific figures, saying only: “That’s a very pointed question. I don’t know if I’m allowed to say.”

  • His urge to create showed up early. In elementary school, he designed RPGs, leveling systems, equipment and adventure mechanics in a battered notebook, while classmates “recharged” with snacks such as mushroom beef and spicy strips. His takeaway: “Once you build a little world yourself, you become the god of that world. You have the final say on what it means.”

  • He is not good at writing and dislikes writing prompts, but he is good at conversation. That has directly shaped his view of future interaction: a keyboard should not pop up awkwardly inside 3D space; users should talk to a floating assistant and achieve “say it and it happens, think it and it appears.”

2. Generational Differences Are Moving Content from Text Toward AI and 3D

  • Song still instinctively opens Google or Baidu when he encounters a problem, while his younger brother naturally uses GPT and DeepSeek for searches and projects. Song is used to public WeChat accounts; his brother is more likely to look for answers in videos on YouTube and Bilibili.

  • He explains the gap through his own media history. As a child, he read 5M-word novels on an MP3 player that displayed only 10 characters per screen—“you had to press the button 500,000 times.” After buying an iPhone 4 in middle school, he began consuming images; he encountered video at scale in college, then became accustomed to short video.

  • A younger cohort of “born with 3D” users has already emerged. Song mentioned a junior of his brother who makes 3D works featuring Toilet Man and Camera Man on Bilibili—content Song cannot understand but that can still draw hundreds of thousands or even millions of views. Demand is not bounded by the previous generation’s taste.

3. Tripo 3.0 Moves AI 3D from Generatable to Directly Usable

  • Tripo launched around early 2024 and had been on the market for 1.5 years at the time of the interview. Song said it had roughly 3M-4M professional creators globally and more than 40,000 enterprise customers, including more than 700 “very, very large” accounts.

  • He defines Tripo 3.0, released around August 20, as an industrial threshold shift. Previously, AI 3D could perform only one step in a production pipeline and still required professionals to modify and refine the output; 3.0 is the first version usable directly across most industries and scenarios.

  • The clearest test is 3D printing. A user can generate an arbitrary model and send it directly to a household 3D printer; the finished result is already “very, very good.” Users do not need to understand internal structures, file formats or model repair, and do not need to master DCC modeling software first.

  • This was not a single-point breakthrough, but a systems effort combining more data, better algorithms and module-level improvements. Controllability, success rate, detail and performance all improved at once, with Song placing particular emphasis on geometric fidelity.

4. SparseFlex Is 3.0’s Key Representation, Not a One-Trick Solution

  • The host tried to trace the 3.0 architecture from Tripo 2.0’s composite DiT and U-Net architecture. Song did not directly confirm whether the architecture carried over, instead rejecting the idea that “it’s powerful because it uses MoE”: “That’s not how science works.”

  • He focused on SparseFlex, or SF, developed by VAST. The team open-sourced TripoSF in April of that year. It significantly reduces generation costs and improves speed, while simplifying processing by skipping the watertightness step.

  • SparseFlex supports spatial generation at resolutions above 1,000×1,000×1,000, enabling finer granularity and more geometric detail. Song compares it to a 3D token: the better the representation, the higher the compression, reconstruction and fidelity rates—allowing it to carry more training data while improving generation quality.

  • He sees mesh, NeRF, Gaussian representations and SparseFlex as different forms of 3D representation, not as one format responsible for all progress. Tripo 3.0 currently ships only Standard; the later Ultra version will produce better results but may take longer to generate.

5. Tripo Studio Turns “80-Point Assets” into a Full Workflow

  • Asked what Tripo 3.0 unlocked that Tripo 2.5 could not, Song declined to name a single scenario and instead distinguished the foundation model from Tripo Studio. The real addition, he said, is an entirely new editing and interaction paradigm for the AI era.

  • Tripo Studio launched on May 31, with the goal of replacing complex, redundant traditional 3D pipelines with an AI-native workflow. Song said revenue “more than doubled” after launch; the host then summarized that Studio now contributes more than half of company revenue.

  • The product is not designed to require the model to deliver a perfect result in one shot. It lets users generate something “80 points good” and continue editing it inside Studio, sharply reducing the cost, time and barriers of the traditional professional pipeline.

  • Song believes every vertical group with professional expertise will eventually have its own AI workstation. It must handle content creation end to end, expose sufficiently fine editing granularity and offer an interaction model unlike traditional software.

6. Segmentation, Completion, Rigging and Low-Poly Generation Put AI Assets into Production

  • Semantic segmentation and part completion solve the problem of generated content arriving as “one lump.” Song compares the system to Photoshop layers: AI images previously had no source files, and AI 3D assets could not be disassembled; now the system can identify objects, separate a hand from the water bottle it is holding and complete each into a full asset for replacement.

  • “Universal rigging” turns static sculptures into animation-ready assets. Beyond humans, cats, dogs, cows, snakes, fish, dragons, octopuses and spiders can all be automatically rigged, skinned and configured with motion, with detail reaching the level of individual fingers.

  • For real-time rendering pipelines in games, XR and the metaverse, VAST developed low-poly generation based on an autoregressive approach. Traditional assets can contain hundreds of thousands or even millions of polygons; the new models can naturally stay within the hundreds or thousands, significantly reducing local compute requirements.

  • Studio also offers Magic Brush, style extraction, symmetry, T-pose and Z-pose capabilities. These are not isolated features; together they connect generation, modification, structuring and deployment into a usable workflow.

7. Foundation Models Build the New Wall; Products Protect the Old One

  • Song believes both foundation models and agents or workstations will be built in the future. The relationship is not simply offense versus defense: “When you build a foundation model, you’re putting up a new wall. When you do engineering and product features, you’re protecting the old wall.”

  • The risk is that an application team may invest heavily in engineering around shortcomings in a previous-generation model’s faces, only for the next foundation model to solve 100 problems in one stroke and fix faces along the way. The old engineering then loses its value.

  • VAST controls both the foundation-model roadmap and user feedback, allowing it to judge which old walls are worth protecting and which gaps will disappear naturally in the next release. Studio brings the team closer to frontline users, whose needs then feed back into the foundation-model roadmap.

  • He uses Cursor as a hypothetical example: if GPT-5, GPT-6 and the rest advanced rapidly through GPT-9, with general models steadily absorbing vertical problems, a pure-play application could have “nowhere left to hide.” In his view, Cursor should eventually build its own foundation model.

8. “Language Has Hit the Wall” Is Both a Limited Claim and an Explanation for Application Growth

  • Song describes AI 1.0 as scientists manually tuning parameters and training large numbers of small models to solve long-tail problems one by one, such as garbage overflow and prison fights. AI 2.0 uses data-driven large models to generalize across whole groups of long-tail needs.

  • His sharp summary is that if general-purpose foundation models can no longer keep solving vertical problems rapidly, vertical engineering and products gain room to exist. But he immediately qualifies it: “It’s not that all models have hit the wall; I’m saying language has.”

  • As language-model progress slows relatively, vertical products and agents gain space. 3D, by contrast, is still moving at the pace of a major release every 3-5 months, making companies with no proprietary model and only an application layer more vulnerable to being covered by the next “new wall.”

  • The host cited DeepSeek as an example: Liang Wenfeng would rather reject commercialization than dilute investment in pushing the boundary of model intelligence. Song does not see building products simultaneously as a distraction; he believes doing only foundation models can become “academic self-amusement,” cutting the company off from real demand.

9. VAST Found the Real Prerequisite through Its Early “3D TikTok” Experiment

  • Early on, the company tried to build a “3D TikTok,” but quickly hit a wall: to create a UGC ecosystem, the world first needed 3D UGC. In reality, there was almost only PGC because ordinary people lacked usable creation tools.

  • Song compares AI 3D with the input method in the text era and the smartphone camera in the image and video eras. VAST therefore shifted toward foundation models, with the goal from day one of letting everyone create 3D content with zero barriers, zero cost and real-time interaction.

  • He repeatedly stresses that the company was not “holding a hammer and looking for a nail.” VAST did not first build a general-purpose foundation model and then search randomly for use cases; it first identified 3D UGC as the nail and then filled the missing condition—the creation tool. “You can’t make dumplings just because you have vinegar.”

  • The host questioned whether a company must build the product itself to understand users. Song’s answer was that pure foundation-model companies usually serve B2B customers and can only reach those customers and then ask about their customers. Since the endgame is serving creators, the team must work directly with creators to verify whether a problem has actually been solved.

10. Studio Serves PGC First; a True Mass-Market Product Will Deliberately Sacrifice Control

  • Tripo Studio currently targets Pro C or PGC users, not elementary-school students or Song’s grandmother. Professionals can understand editing, topology and workflows; mass-market users need an entirely different entry point.

  • The next step will gradually sacrifice some controllability and editing granularity in exchange for a large library of content paradigms and templates, allowing PUGC and UGC users to create without understanding 3D technology.

  • Song explains the route in terms of scores. Moving from 10 to 80 requires both UGC and PGC to first generate something that “looks like a thing.” The paths diverge from 90 to 98 or 99: UGC prioritizes speed and simple animation, while PGC prioritizes detail, topology and edge flow.

  • Studio is therefore not a detour from the endgame. It first serves users in the existing production relationship, pushes overall capability into the 80s or 90s, and then generalizes it to “native users” who previously lacked 3D capabilities and entered the production relationship only because of AI.

11. The Question “Why Would Everyone Make 3D?” Misses the Fact That Technology Shapes What Feels Natural

  • The host’s objection was that taking photos and shooting video feel natural, while 3D is a higher-dimensional form of artistic creation that people may not instinctively want to do. Song argues that this sense of naturalness is less than ten years old; it is a habit created after cameras, bandwidth and platforms lowered the barriers.

  • He asks in succession: before short video, how many films did people watch in a year? Before Xiaohongshu and Pinterest, how often did they visit a gallery? Before Weibo, Tieba and Twitter, how many books did they read in a lifetime? His point is that content platforms release consumption frequency previously suppressed by cost.

  • Games have already demonstrated demand for 3D. According to data cited by Song, Honor of Kings had roughly 100M DAU in its first year and still had roughly 100M DAU ten years later. Sustaining that scale in a country of roughly 1.3B people makes it a truly mass-market product.

  • He puts the global games market at roughly $260B, at least 2-3x the combined market for publishing, galleries and film. On that basis, 3D is not a niche medium; it simply lacks a corresponding mass-market UGC supply and distribution platform.

12. The Future of 3D Content Is Not More Games, but a New Category Yet to Be Named

  • Asked whether future platforms would be used for games or content like Jenny Three, Song admitted he could not name the category directly. Before short video emerged, people discussing UGC video could think only of film, not the vast majority of video formats that later appeared.

  • He uses Bilibili’s categories as an example. Film is now only a small slice of the video universe, and the short-drama market has already surpassed film. Likewise, match-three games, Genshin Impact, Honor of Kings and battle royale games will not exhaust the space of interactive content.

  • His strong view is that all games today will eventually become only “a small part of the gene pool” of the broader category of “three-dimensional interactive content.” VAST wants to serve not one type of game production, but a network capable of generating unknown content, gameplay and social relationships.

  • His more aggressive scale forecast is that future 3D platforms could reach 2-3x the combined scale of Twitter, Weibo, Xiaohongshu, Douyin, Kuaishou, TikTok, Snapchat and Instagram. This is Song’s end-state judgment, not a conclusion validated by commercial data.

13. The Three Bottlenecks for 3D Foundation Models Are “Feed, Trainers and Racecourse”

  • Song uses horse breeding as an analogy for the three elements of AI: data is feed, scientists and algorithms are trainers, and compute is the racecourse. All three are scarcer in 3D than in language because the early internet did not accumulate large-scale 3D content that could be scraped directly.

  • VAST reportedly has roughly 40M high-quality native 3D models, with quality at “the level of that monkey in Black Myth: Wukong.” Other major companies and competitors are generally in the millions. Asked how the data was acquired, Song replied only: “That’s our core secret.”

  • Talent is constrained by the age of the discipline. The team includes roughly 50-60 Tsinghua PhDs concentrated at the intersection of AI and graphics; many have worked with AI 3D for only 1-2 years. VAST’s advantage comes from identifying talent early enough, investing in it and getting the team to work together.

  • Song says VAST’s total funding “should be” the highest in the sector and its valuation “may also count as” the highest. Three financing rounds and the resulting capital reserves constitute the racecourse, but he retains uncertainty: “A larger caravan does not necessarily find the oasis faster.”

14. For the First Two Years, the Product Was Essentially Technology; Only Now Has the Productization Fork Appeared

  • Looking back at 2023-2024, VAST did one thing: push its technology to the global SOTA. The company went a long time without product managers, and the CTO personally wrote a large amount of code, because no peripheral feature could substitute for foundational quality.

  • Song’s analogy is that when competitors’ cameras reach 1080p, adding portrait, panorama and infrared functions to a 720p product is pointless. The priority is to catch up to 1080p and then 4K. “The essence of an early product is technology.”

  • By the end of last year, the team concluded that it needed a professional workstation, then launched Studio in roughly 6 months. From the end of last year into the current year, VAST hired a product and engineering team of more than 20 people to handle presets, stylization, poses, interactions and product and engineering issues.

  • The company’s long-term path has barely changed. At the all-hands sharing session each year in mid-March, Song puts screenshots of the old presentation into the new one and updates only the milestones that have been completed. The team has seen almost no attrition, and he hopes the fourth presentation will still be about the same thing.

15. One to Two Thousand “Chief Experience Officers” Expanded Demand from Games to the Entire Creative World

  • VAST runs a CEO Program, where CEO stands for Chief Experience Officer. It has interviewed roughly 1,000-2,000 real users so far. Song admits that many of the ways people use the product were completely outside the team’s initial assumptions about games, animation and 3D content.

  • The demand pool now contains hundreds of requests ranked P1, P2, P3 and P4; requests after P5 may not get done in time. Users asked to edit textures, so the team built Magic Brush; users wanted to alter geometry but could not operate professional tools, so the team began researching natural-language editing. Better topology, hard surfaces, corners, UV integrity and preserving brush strokes also came from interviews.

  • A single foundation-model iteration can solve a group of problems at once, such as facial detail, hard surfaces or texture quality. Other needs require dedicated AI algorithms and interaction design. The product team’s value lies in distinguishing problems the next “new wall” will take away from work that must be completed separately.

  • The host continued to ask whether building an entire product was far more expensive than user research. Song insisted that the company’s founding purpose was not AI 3D AGI itself, but a creator platform. Without putting the tools into the hands of real users, the team cannot know whom the model is actually serving.

16. The Most Valuable Adoption Came from Use Cases the Team Had Not Defined in Advance

  • Beyond games and animation, users now apply Tripo to 3D printing, industrial design and all kinds of design work. Graduating students at art schools use it for graduation projects in contemporary art, installation art, landscape art and new media art.

  • Song particularly highlights users with disabilities who use AI 3D to express themselves and create industrial designs and content. For people who previously struggled to enter professional 3D production, generative models are not simply productivity tools; they make participation possible.

  • XR is not only about hardware iteration. In interviews, Song saw an active software community creating 3D PowerPoint presentations, 3D AI picture books and small games. AI 3D may also create new gameplay; over the past decade, the only genuinely new game mechanic he can think of is auto chess.

  • In his view, the industry underestimates the speed of invention. Less than 3 years after the technology was invented, it is already being used by millions of people and more than 40,000 companies. Because so many inventions have appeared recently, people have not fully recognized that this technology is already mature enough for industrial adoption.

17. AI 3D Is Not Just Another Option; It Gives the Masses a Capability That Previously Did Not Exist

  • Song believes text-to-text, text-to-image and text-to-video are valuable, but people could already type, take photos and shoot videos. AI mainly adds a more convenient option. 3D generation is different: “You couldn’t do it before; now you can.”

  • He calls the sudden appearance of a usable object after pressing a button “Ma Liang’s magic brush,” “say it and it happens” or magic: “Now every single person can create anything. That is not a small thing.”

  • Menus provide the clearest commercial example. When 10 people order food, a two-dimensional shopping cart cannot intuitively show whether the dishes or portions will be sufficient. If each dish were placed as a 3D model on a virtual table, you could tell at a glance whether there was enough to eat.

  • Restaurants could never have spent RMB1M building a 3D ordering system. If generation brought the cost down to RMB150K per year, the decision would change. Beyond menus, billboards, business cards and all kinds of display media could be redesigned, because the cost of 3D modeling would no longer determine the form.

18. The Internet’s Endgame Is “Decompression”; Human Value Shifts Toward Creating Worlds People Want to Stay In

  • The host described VAST as adding dimensions back to a compressed world. Song refined that into “decompression”: people are not inherently drawn to low-poly graphics or the low resolution of Legend of Mir; networks simply could not carry more at the time. As compute improved, users naturally moved toward high-fidelity experiences like Black Myth: Wukong.

  • Humans first expressed themselves through statues, totems and cave paintings, and invented writing later because carrying and transmission were expensive. Text, images and video are not natural endpoints, but successive abstractions of the 3D world created by technical limits. Video selects a position, angle and moment; in 3D, “you’re inside it, and you stay there.”

  • Song sees content as the “fourth major industry” after agriculture, manufacturing and services. He imagines all physical value eventually being produced by robots, with human value centered on creativity. The measure will be the total time people spend in the content, experiences and worlds created by an individual.

  • This vision also explains his entrepreneurial choices and psychology. He would tell his self from 2 years ago: “Fuck, I got it right. I’m incredible. I was really brave. I really feel like I’m somebody.” When difficulties hit, he switches worlds through games and Dungeons & Dragons: “When the physical world accounts for only 40-50% of your life, at least half of it doesn’t have to hurt.”