Will Live-Action Films Survive? Director 陆川 on AI's Fear and Freedom
Summary
The inflection point for AI video is no longer whether it can generate; it is that production costs are collapsing first, forcing control and rights systems to catch up. 陆川 compressed traditional action-VFX previs from roughly 6 months to 48–72 hours and believes AI can cut post-production effects from several million yuan to the RMB100K-plus range; “Within 7–8 years, no director will say, ‘AI? I don’t think it works.’” But the systematic infringement accusations around Seedance 2.0 show that the faster costs fall, the harder it becomes to avoid questions around copyright, authorization and job redesign.
General-purpose models optimize for the majority’s taste, potentially sanding the authorship professional filmmaking needs most down to the mean. When 陆川 designed a male lead for a Ming-dynasty story, domestic models kept producing similar popular handsome stars regardless of the prompt; he believes consumer-side data from costume idol dramas not only trains models but may also “dilute the relatively high-quality knowledge base that was already there.” The real industry moat will be film-specific models capable of unifying characters, shots, editing and color grading through “industrial-grade delivery.”
AI-made work cannot command a film premium simply because it is AI; it will ultimately face the same comparison against a century of film language. 陆川 rejects cobbling together models and agents of uneven quality into a temporary workflow, comparing it to scattering the artillery, air force and navy across different countries and then barely assembling an army. “The important thing is not AI; it is film”—it must still be judged alongside The Godfather, Rashomon and Apocalypse Now; AI itself “is just a pen, just a camera, just an editing table.”
Job displacement is more likely to happen at the task level: AI can do about 90% of the work, so human budgets should go to irreproducible performances. 陆川 says AI can do 90% of it; repetitive clichés no longer need 100-plus people filming under blazing sun and torrential rain. Investors and creators must answer why a film needs human live action. He does not fear AI will kill live-action films, but believes human crews could become like horseback riding: moving from transportation to a costly pursuit of the elite, and an expensive, scarce mode of creation.
Voice shows what remains hardest for AI to replicate in human performance: not standard pronunciation, but context, collaboration and meaningful imperfections. 黄莺 cites pauses, stress and leaving emotional space for the protagonist in Jian Lai: the same line, “The sky is so blue today,” changes its emphasis and delivery depending on what came before; voice cracks, mic pops and tremors can also create an immediate shock. She still leaves the question open—“Who knows what the future holds”—while performances stripped of personal emotion and hardened into formulas will be easier for AI to replace.
Voice rights already have the same legal protection as image rights; what the industry lacks is enforceable provenance, evidentiary standards and rules for blended cloning. 黄莺 has negotiated separately since 2016 over whether her voice may be used for AI training, but unauthorized, professional-grade replication has surged over the past 2 years. If the defendant denies it, a voiceprint alone may not establish provenance; if a voice is synthesized 50% from each party, the rights boundary becomes even harder to define. Her view: voice rights and image rights enjoy equal protection, but the concepts remain blurry in enforcement.
AI expands content supply and individual creative freedom while lowering equipment and capital barriers; judgment, taste and professional training remain irreplaceable. KiSA used Midjourney in 2023 to rewrite the script while generating images and completed Overtime Night in about 15 hours, a film that would previously have been nearly impossible to shoot. 陆川 moved from fearing that works would be carried into his grave to believing every idea could be completed: “It makes me fearless.” He still advises young people who love film and can get in to attend film school, because the experience of working together offline is something “you can never learn online.”
Deep dive
1. Generation quality clears the gimmick hurdle; the copyright and jobs bills come due
The episode opens with ByteDance’s February release of Seedance 2.0. The model topped multiple benchmarks, while a clip of Brad Pitt and Tom Cruise fighting atop a skyscraper drew more than 7M views; the expressions, lip-sync and vocal timbres were nearly convincing enough to pass for real.
Disney, Paramount and the Motion Picture Association subsequently condemned the system for systematic infringement. The issue was not only protection of creative works, but the copyright regime underpinning millions of US entertainment jobs. 泓君 asks whether lower costs and higher efficiency will come at the price of creator unemployment and lost artistic value.
陆川’s timing call was unequivocal: “Within 7–8 years, no director will say, ‘AI? I don’t think it works.’” If effects costing several million yuan can fall into the RMB100K-plus range, refusing to use AI will require a new economic rationale.
2. AI targets the most expensive waiting in IP development
KiSA entered screenwriting around 2015 and worked on Medical Examiner Qin Ming. His shorthand for the job at the time: a vast pool of novel IP on the left, platforms on the right, and screenwriters moving the material across. An S+ series could spend years in development while mobilizing huge amounts of capital.
Film and TV companies review hundreds or thousands of IP packages each year and must select a handful within days, yet the fastest route from text to release still takes 2–3 years. Platform-director disagreements, actor availability, production resources and role rewrites can all trigger another round of work; rework is “almost a constant.”
In 2023, KiSA left the company to work on AI content. By August, no more than 10 people may have been making AI shorts on the Chinese internet. He used Midjourney to produce Overtime Night, but because generated motion could not guarantee editing continuity, he had to rewrite the script as he went. The project was completed over a single weekend, in about 15 hours. “That was when I felt my expression was infinite.”
3. Six months of VFX previs compressed into 3 days
陆川 breaks down a conventional action sequence: storyboards typically take 1–2 weeks, or up to 1.5 months at a slower pace; then comes the motion board, followed by a 7–8-person or even 10-person previs team rebuilding the models. Once the project enters a VFX house, dozens more people reconstruct a monster’s skeleton, muscles, skin, hair, materials and lighting. Designing a full monster asset alone can take 1–4 months.
These teams operate independently, and the rough models from one stage often cannot be reused directly in the next. A conventional VFX process therefore takes roughly 6 months. The cost of effects-heavy films is not an abstract “technology cost,” but the accumulated bill for specialized labor, repeated modeling and cross-team sign-offs.
The AI workflow reduces creative direction to core commands, then uses tools such as Midjourney to generate keyframes. At the pace 陆川 cites—4 images in 16 seconds—the system can produce hundreds or even thousands of images a day. The results do not always follow the director’s instructions, but occasionally are “better than you expected.”
For a 1–2 minute action sequence, keyframes can be extended into moving footage to complete a visualization in 48–72 hours, or no more than 3 days. 泓君 describes that as close to a 100x gain in time efficiency, but consistency, controllability and final quality in a feature-length film do not resolve themselves automatically.
4. General-purpose models produce average taste; film requires industrial-grade delivery
While designing the male lead for an animated film set in the Ming dynasty, 陆川 found that domestic models defaulted to similar popular handsome stars no matter how he changed the prompt. The training data contained a large volume of costume idol dramas, making it difficult to produce a Peaky Blinders-style character—unconventionally handsome, but with tension and personality.
His criticism goes to the objective function: general-purpose models align with the majority’s taste, while art requires authorship, independence and sometimes a critical edge. If a model mainly serves vertical short dramas and mass-produced content, its capabilities may be narrowed by pandering rather than elevated by better work.
For ordinary users, animating old photos or automatically cutting footage from a child’s ages 1 to 6 into a 15–20 second video can be enough; flashy and polished does the job. But films, long-form series and mid-length dramas operate against 130 years of industrial standards and must carry commercial-release value. 陆川 therefore argues for a film-specific model capable of “industrial-grade delivery.”
Existing agent workflows resemble placing the artillery in Israel, the infantry in Turkey, the air force in Spain and the navy in France: capabilities and aesthetics vary widely, and a final color grade cannot easily unify the work. 陆川’s quality floor is simple: “You can’t call it a film just because it was made with AI.”(你不能因为是用AI做的,就叫电影)
5. The moat in voice lies in the subconscious and in imperfection
In 2002, 黄莺 was 1 of 4 people selected from tens of thousands to join the Shanghai Film Dubbing Studio. Translation had to meet the standard of “faithful, expressive and elegant,” while the actor also had to understand the character and re-embed the rhythms of English, Italian and Japanese into Chinese lip movements. Film, tape and film-to-tape workflows left almost no room for repeated trial and error.
For now, she believes AI lacks the “subconscious.” A pause can signal thought, a stutter, unwillingness to answer or an attempt to find an excuse. The same line—“The sky is so blue today”—requires different stress and intent if the preceding line was “The sky was so gray yesterday” rather than “The sun feels so warm.”
Performance must also yield space to collaborators. When 黄莺 voiced “Open Heaven” for the sword spirit in Jian Lai, she did not push the emotion to its limit because the full emotional amplitude had to be reserved for the protagonist, 陈平安. The same scene might end restrained on screen but need to build toward a climax on stage; that difference reflects both artistic judgment and audience psychology.
A voice crack, clipping, mic pop, nasal tone, tremor or uncontrolled physical reaction is not necessarily an error. Such details can make the audience feel as if “saliva is spraying onto my own face.” 黄莺 warns that excessive pressure to remove these traces moves performance toward AI—“but human beings are not perfect.”
6. Voice rights have a legal principle; provenance and blended cloning still lack a yardstick
黄莺 draws a line on replacement: “This is not happening yet—but who knows about the future.” If an actor knows only how to pronounce each word, not what they are expressing, the performance has lost its personal emotion and hardened into a formula, making it “very easy for AI to replace.”
Two days earlier, 泓君 had spoken with 郑大圣, 杨浩宇 and others. Their greater concern was not AI itself, but whether audiences shaped by short videos, short dramas and “Xiaomei” and “Xiaoshuai”-type content would accept generated work. “What audiences can accept, the market will produce.” Demand-side taste may rewrite the industry before supply-side technology does.
Since 2016, 黄莺 has negotiated AI training terms in commercial contracts. Some companies agree that her voice will not be used for AI learning, training or deployment; others insist on using it for training and deployment. The terms are negotiated separately, with different limits in each case. What she cannot accept is use of her voice without her authorization, informed consent or a signed agreement.
Such cases were still rare in 2016–2018, but the past 2 years have brought professional voice actors’ voices being infringed and large volumes of related content being produced. A Beijing case was won because the commissioning party admitted the violation. If the counterparty denies it, off-the-shelf technology may not clearly prove that a voice is yours; “50% my voice plus 50% your voice” makes ownership even less clear. The Civil Code gives voice rights the same level of protection as image rights, but enforcement still lacks technical and legal yardsticks.
7. Live action is better reserved for irreproducible moments
While filming Nanjing! Nanjing!, 陆川 spent more than 1 month deciding what 范伟’s interpreter should say before execution. He rejected 100-plus options, then came up with a line while traveling to work on the final day of shooting. He was moved to tears, but still was not sure whether it worked.
On set, he quietly told 范伟 the line. 范伟 immediately began to cry and went aside, cigarette in mouth, to compose himself: “My wife is pregnant. My wife is pregnant again.” Two days later, 范伟, who normally did not drink, had someone bring a special-vintage bottle of Wuliangye from home, saying he wanted to share a drink with the director for that line. He eventually had to be carried away drunk.
For 陆川, this is the clearest example of where to draw the human-machine boundary: give AI the work AI can do and do not waste investors’ money. But creators’ genuinely original ideas and the performance moments produced by human collisions should remain with people.
Young directors are already using AI shorts in project submissions to show more precisely what they want to make. The next step is not to shoot everything live, but to deploy 100-plus-person human crews to capture “irreproducible, unrepeatable performance moments.” Repetitive clichés can go to AI.
8. Live action may become as expensive as horseback riding; film school still matters
陆川 compares the future of live-action production to horseback riding: once everyday transportation, now an elite pursuit. The norm may be a super-director at a computer coordinating a large number of agents; organizing dozens of people to haul cameras into the wilderness and build real sets could instead become a top-tier, luxury creative experience.
To those insisting on the old method, he compares the moment to a tractor already waiting at the edge of the field. The question is not how to preserve subsistence farming forever, but whether to learn to become the driver and start the truck—or hold up “2 charred sticks” to block the tractor from entering the field. When an era ends, “it will not give you any breathing room.”
AI has also eased 陆川’s earlier creative fears. He once worried that some films would be carried into his grave because of content restrictions, finances or physical stamina. Now, even without funding or a public release, he can make them and leave them on his computer for his children to watch: “I won’t carry a single film into my grave.”(我不会把任何一部电影带进棺材里)
Widespread tools do not make film school obsolete. 陆川 says a real film-school education is like taking a Yangcheng Lake crab to Yangcheng Lake to be dipped, rather than serving up a “bath crab.” At the Central Academy of Drama, he watched students sweat while putting up a set for a Japanese troupe, with 刘烨 and 胡军 playing basketball nearby. That afternoon made one thing clear to him: “You can never learn online what those 2 words—drama—mean.”