Pioneers Insight Method Research Author
Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs
Back to Episodes

Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs

Summary

  • Cerebras says AI infrastructure is serving booked demand, not betting that customers will eventually appear. Andrew Feldman disclosed a $25 billion backlog and described customers as “trying to capture yesterday’s demand,” with football-field-sized facilities drawing more power than midsize cities. For investors, the operational challenge is keeping customers while demand outstrips the ability to build and fill data centers.

  • Reasoning models make inference speed and token availability part of product capability. Rombach says models such as “Fable” and “5/6” increasingly infer intent without a “prompt whisperer,” while his Hermes-agent experiment used a Bittensor project running Z.ai’s GLM-5.2 and debated its own research strategy before acting. Long-running jobs can already produce “amazing things”; Rombach’s hypothetical was that 15-times-faster Cerebras inference could compress “weeks or months’ worth of thinking” into a day.

  • Cerebras expects its young architecture to improve substantially faster than conventional Moore’s law in the next 18 months. Feldman’s view is “way over 2x,” because a new architecture still offers workload-specific optimization unavailable to a 20-year-old GPU design reliant on smaller fabrication geometries. Hyperscalers’ custom chips do not invalidate merchant silicon demand: their primary objective is avoiding total dependence on any supplier.

  • Open models are becoming the cost, control, and sovereignty layer beneath frontier AI. Feldman’s analogy is that users should not “take your Ferrari to the grocery store”: frontier models handle the hardest problems, while open models absorb routine G&A and the “cutting and pasting economy.” Regulated enterprises also want domestic, on-prem deployment; Cerebras therefore runs gpt-oss-120B, GLM, Kimmy, the “Quincy” set of models, proprietary OpenAI models, and customer-built models.

  • Staged release and red-teaming become defensible when a model can expose serious software flaws in hours. Rombach relayed that Palo Alto Networks found previously unknown bugs and paused other work for six weeks of patching; Feldman said giving government defenders “2 or 3 weeks to patch any obvious holes” is reasonable. He simultaneously warned against reflexive regulation and treated a future massive data breach as inevitable: “Something will happen.”

  • By every AGI definition used 20 years ago, Feldman believes the threshold has already been crossed; the open question is where recursive improvement stops. Repeatedly asking models to learn, retry, and broaden their search yields “not a little bit better answers, but vastly better answers,” compressing thousands of human-style generations into machine-speed iteration. His ledger includes real labor dislocation, but also a chance that children and their loved ones do not die of cancer and that every child receives an adaptive tutor.

  • Black Forest Labs sees generative media as the foundation of multimodal world models, not merely an image-and-video tool business. Robin Rombach traced FLUX back to latent diffusion, then toward models pretrained on images, video, and audio and combined with action prediction—the same model could help make a movie or become “a brain on a robot.” Near-term value remains human-directed: Martin Scorsese used the models to externalize a scene in his head, while IP owners can combine controlled generation, customized models, and potentially licensed fan creativity.

Deep dive

1. AI infrastructure is chasing yesterday’s demand

  • Feldman’s scale claim is deliberately physical: upcoming data centers may consume more power than the previous 50 years on Earth, while individual football-field-sized buildings receive more electricity than midsize cities. Construction spans the US, Canada, the Nordics, France, Europe, the Middle East, Kazakhstan, Tajikistan, Georgia, and Armenia.

  • Cerebras has a $25 billion backlog, but Feldman says it is not alone. OpenAI, Anthropic, Google, Microsoft, and AWS also want more data centers, while Rombach names SpaceX and xAI among the insatiable capacity buyers. These customers are not pursuing “if you build it, they will come”; they are “trying to capture yesterday’s demand.”

  • On token-maxing, Feldman says experimentation can waste some resources, but the net value is enormous. He compares early token access to inexperienced Costco shoppers buying “four things you didn’t need and each was $22” before learning to shop strategically.

  • Enterprise behavior is already maturing from “Everybody, as many tokens as you want” toward differentiated allocation: highly productive teams get what they need, while cheaper or open models serve other work. Rombach says individuals “start playing with the tool, and then the tool starts playing with them,” forcing users to articulate goals, systems, and requirements.

2. Reasoning turns token volume into a hardware problem

  • Early computers and prompts did “exactly what you tell them”; tiny wording changes could transform the output. Feldman says “Fable” and “5/6” now increasingly infer the desired chart, layout, or solution without requiring users to be a “prompt whisperer”—a material shift from summarization toward understanding intent.

  • Rombach’s Hermes-agent experiment used Z.ai’s GLM-5.2 through a Bittensor project with unlimited capacity. Asked to become the world’s best trend hunter, the agent debated whether to search Hacker News, Reddit, social media, or Instagram: “You were watching a reasoning model work out.”

  • Unlimited tokens might mean “unlimited reasoning,” because models can spend 25 or 48 hours exploring a problem. Rombach’s explicitly hypothetical Cerebras case—15-times-faster execution sustained for 24 hours—could produce “weeks or months’ worth of thinking,” making inference latency part of answer quality rather than merely user experience.

  • Feldman says Cerebras broke from the traditional 18-month doubling curve and expects “way over 2x” improvement during the next 18 months. His logic: new architectures retain large workload-specific optimization opportunities, whereas a 20-year-old GPU architecture depends more heavily on shrinking fabrication geometry.

3. Open models become the minivan of enterprise AI

  • Hyperscalers building silicon does not mean every chip must beat the merchant frontier. Feldman’s explanation is institutional memory: cloud providers once depended on Intel, while GPU vendors depended on a few hyperscalers. “You just can’t be entirely dependent on other people’s chips.”

  • Model routing follows the same diversification logic. “You don’t want to take your Ferrari to the grocery store”: OpenAI, Anthropic, and Gemini may serve frontier problems, while routine Workday manipulation, G&A, and the “cutting and pasting economy” need reliable open-source capability rather than “gold medal math.”

  • Sovereignty adds a second demand driver. Finance, healthcare, and other regulated industries may require domestic, on-prem systems with tighter control over leakage and intelligence; Feldman calls gpt-oss-120B a good move but argues the US needs “more domestic open-source models” to offer an alternative to Chinese models.

  • Feldman presents Cerebras as able to run a wide range of models: GLM, Kimmy, the “Quincy” set, closed OpenAI models, GlaxoSmithKline’s internally developed model, and models from G42 and MBZUAI. Government actions regarding “Fable” and “56” were, he says, particularly in Europe, a wake-up call.

4. Staged releases look reasonable once models find serious holes

  • Asked whether a powerful model should roll out gradually, Feldman offers an honest limitation: “I hadn’t seen it before,” so he cannot judge that specific release. In principle, though, staged access resembles precautions for powerful pharmaceuticals and could give government red teams “2 or 3 weeks to patch any obvious holes.”

  • His balancing point is that polarization “hurts clear thinking.” The US should preserve fierce competition among Dario, Sam, and their peers and avoid becoming a region whose first instinct is regulation, while recognizing that developers are inventing both the technology and its guardrails without a playbook.

  • Guardrails themselves add latency. Cerebras discovered during the previous six weeks that faster chips can make those protections “less painful,” illustrating how safety and performance are coupled rather than separate product layers.

  • Rombach’s sharpest evidence came from Palo Alto Networks’ Nikesh: a tested model reportedly found bugs the company did not know about, forcing six weeks of patching. Feldman expects a massive breach eventually—an unspecified “unknown unknown”—and argues institutions should prepare their response before “something will happen.”

5. Old AGI tests are obsolete; recursive learning is the live question

  • Feldman agrees that “by any definition we had 20 years ago we’ve hit it.” Earlier Turing tests and science-fiction waypoints have been surpassed; the harder task is generating questions appropriate to capabilities that previous observers did not anticipate.

  • The recursive mechanism is simple but potentially exponential: ask, learn from the answer, retry with more information, and broaden the search. These loops produce “not a little bit better answers, but vastly better answers”; what remains unknown is whether improvement stops at compute, token, or budget limits—or keeps climbing.

  • His large-project analogy carries the argument: Versailles and other major structures accumulated learning across three or four generations of specialist families, while humans often wait 20 or 40 years for paradigms to change because “paradigms don’t die. People do.” AI instead approaches Drosophila-like iteration—“two a day”—compressing thousands of equivalent generations.

  • Feldman does not deny economic dislocation; cars were bad for horseshoers and carriage builders. His counterweight is a chance that cancer research might spare children and their loved ones, plus adaptive tutors that teach each child individually rather than today’s “factory farming” classroom model. “Thoughtfully and fairly” written, he thinks the ledger can come out okay.

6. Black Forest Labs is building multimodal world models

  • Rombach traces Black Forest Labs’ foundation to latent diffusion, developed by him and collaborators as PhD students in Munich. The method compresses natural data into efficient representations—analogous in principle to JPEG or MP3—then trains transformers over them; the team subsequently built Stable Diffusion and FLUX on that base.

  • The roadmap now combines models pretrained on images, video, and audio with action prediction. Rombach distinguishes intuitive visual intelligence from a deep reasoning layer: a complete intelligence needs both, but video pretraining can already provide implicit knowledge of physics and real-world interaction.

  • Control has expanded progressively from text-to-image, to text-plus-image editing, to multiple images combined semantically through a prompt. Applying the same pattern across video, audio, and action makes every modality a potential input or output, exposing more “manipulation layers” to users and developers.

7. Scorsese’s use case is communication, not one-click cinema

  • Rombach rejects prescribing how generative models should be used: “These AI models, they are a medium.” With Martin Scorsese, the team explored an Eastern European village, iterating from the director’s description until images communicated the mental scene more directly than language could.

  • The mechanism matters more than celebrity endorsement. Language is a “lossy communication medium,” whereas an image or video contains unusually rich signal; the technology can “parallelize your brainstorming,” extending the familiar storyboard and miniature workflow rather than removing the filmmaker.

  • Rombach is unsure that generating an entire movie is the ultimate goal. He favors a human-in-the-loop workflow and will not predict when high-end cinema arrives, though capability has progressed from 64-by-64-pixel images during his PhD to high-resolution, multi-minute videos.

  • Robin supplied the production-economics case: Gal Gadot described a Bitcoin film shot on a soundstage with generative scenery, reportedly costing $30 million rather than the $150 million physical sets would have required. Rombach confirms some comparable production use, but stresses that the technology remains on a rapidly improving trajectory.

8. The same visual stack reaches robots, IP libraries, and customers

  • Rombach’s most expansive thesis is that “the same kind of AI model” could make a movie and serve as a robot’s brain. Generation becomes simulation; action prediction requires perception and understanding. Today, varied factory hardware still needs a few hours of task-specific fine-tuning data, while the longer-term goal is natural in-context instructions such as picking up a glass.

  • On intellectual property, Black Forest Labs blocks generation of certain IP in its public tools. It also works directly with IP holders on customized models, sometimes built from open-source foundations and sometimes from more powerful proprietary systems—a control-plus-capability proposition for owners of large libraries.

  • Robin’s proposed consumer model is licensed fan creation, extending Star Wars fan fiction and films into AI-made untold stories. Rombach agrees that a model acceptable to rights holders while enabling “super-creative customization” could let viewers visualize alternate events and interpretations they already imagine.

  • The business was two years old and had just crossed 100 employees. Black Forest Labs is hiring in Germany and San Francisco for large-scale and diffusion-model researchers, customer-facing engineers building physical-AI or IP solutions, and infrastructure specialists focused on stable training and maximizing MFU.