The Man Who Invented Prompt Engineering on AI, AGI & Humanoids w/ Richard Socher & Salim Ismail
Summary
- Foundational AI is becoming telecom-like infrastructure, pushing durable value toward customer access, proprietary data loops, and model-agnostic trust layers. Ismail and Diamandis argue that open source is gaining; Socher warns that a pure model provider may resemble a capital-intensive telco whose infrastructure enables others to capture the economics: “You can’t build an Uber without the Internet being everywhere, but Verizon doesn’t get a cut of Uber.”
- Socher believes a couple of billion dollars and “a year and a half to two years” could produce digital superintelligence, but rejects AGI as a single benchmark or robotics test. His pragmatic threshold is automating perhaps 80% of all digitized work and maybe 80% of those workflows; a rigorous definition must also measure sample efficiency, reasoning, knowledge, visual, and social intelligence. “You can have a superintelligence that’s purely digital.”
- Cheaper intelligence may increase total compute and energy consumption even as each unit becomes radically more efficient. Ismail argues data centers are being designed around six-month-old cost assumptions while the incremental effort to create the next generation is falling roughly 10x each cycle; Socher “mostly disagree[s],” invoking Jevons’s paradox as personal assistants, tutors, health teams, and new energy-intensive uses multiply.
- AI-for-science is the episode’s most consequential upside case, with medicine, proteins, materials, and quantum simulation all moving into faster experimental loops. Socher’s Salesforce team generated functional proteins 40% different from natural ones versus 3% after Frances Arnold’s Nobel-winning directed-evolution process; Peter also cites an AI co-scientist replicating 10 years of antibiotic-resistance work in 48 hours. Socher’s governing claim: “Anything you can simulate, AI can solve pretty much every problem in that domain.”
- Current AI valuations can combine “seed-stage risk” with “late-stage returns,” creating a steep revenue hurdle and making defensibility central. Thinking Machines is discussed at a $30 billion starting valuation, while Socher’s preferred pattern is either horizontal infrastructure or vertical applications with deep buyer knowledge and a “virtuous data cycle” that improves as customers use the product.
- Humanoid robotics is advancing quickly, but unconstrained homes, capital intensity, privacy, and unclear form-factor advantages may delay broad deployment. Diamandis cites a projection of as many as 10 billion robots by 2040; Ismail keeps asking why a household machine needs two arms and legs rather than “wheels and seven arms.” Socher’s balanced answer is that humanoids and specialized machines are “not a zero-sum game.”
- Knowledge-work agents are already useful, while a true Jarvis remains blocked more by context, trust, law, and Internet economics than raw model intelligence. You.com says users have built more than 50,000 custom agents, including marketing, journalism, and venture-capital workflows; fully autonomous agents still need intimate preference data and may be resisted by websites whose advertising they ignore. “Eventually we’re going to have more AI agents surfing the web than people.”
- Bitcoin’s durability is improving, but custody and transaction design still compare poorly with consumer finance’s reversibility and insurance. Ismail reads the billion-dollar Bybit hack and promised customer restitution as evidence of ecosystem robustness, yet recommends holding rather than trading because roughly 80% of annual upside can arrive in five unknowable days. Strategy’s purchase of another 20,000 Bitcoin for about $2 billion underscores institutional conviction, while the discussion also highlights Michael Saylor’s evangelism.
Deep dive
1. Frontier models are outrunning the benchmarks used to rank them
Elon Musk’s $6 billion xAI buildout anchors the discussion: a reportedly coherent, largest-scale GPT cluster assembled in 122 days, according to Diamandis. Socher calls the speed surprising but plausible when exponential technology meets capital, noting Musk’s unusual ability to integrate hardware and software rather than treating AI as a software-only problem.
Ismail questions claims that Grok 3 outperforms every rival, saying early reports place it somewhat lower, though its speed of arrival remains “unbelievable.” Socher’s broader point is that ordinary informational needs are already largely served; the frontier has shifted toward “programming, science, research,” including Anthropic’s newly announced Claude 3.7 model.
You.com federates more than 40 models and routes requests by inferred intent and user feedback. OpenAI’s o1 and o3 remain popular, while Socher calls Claude Sonnet 3.5 one of the best models for programming. DeepSeek gained extraordinary mindshare without a large marketing budget—but Socher stresses how frequently the preferred model changes.
A single intelligence score now obscures the role of test-time compute. Simply telling a model to “wait before you answer this” can improve accuracy, making speed another intelligence dimension; even the Turing test breaks when “the best way to fail” is answering impossibly well, such as writing an app in 30 seconds.
2. AGI is better defined by capabilities than one theatrical test
Asked whether a couple of billion dollars would let him build digital superintelligence, Socher answers, “Probably like a year and a half to two years.” He admits the question touches a personal pull back toward intensive research: commercial products and revenue are meaningful, but several research bottlenecks remain open.
His pragmatic financial definition of AGI is automation of perhaps 80% of digitized work, then perhaps 80% of the associated workflows—already “a huge amount of GDP.” It is useful precisely because broader definitions are so inconsistent, but Socher does not present it as a complete academic account of intelligence.
A fuller definition must separate visual, linguistic, mathematical, reasoning, knowledge, and social capabilities. It should also test sample efficiency: humans can learn from one or two examples, so a genuinely intelligent system ought to acquire some abilities with far less data along particular dimensions.
Ismail raises physical tests such as making coffee or assembling IKEA furniture. Socher rejects embodiment as a prerequisite: blind, deaf, or paraplegic people can be highly intelligent, so “you can have a superintelligence that’s purely digital”—physical manipulation is another capability group, not the definition itself.
3. Open models commoditize intelligence while trust captures value
Ismail is categorical that open source is gaining: mainstream excitement attracts too much distributed energy for closed systems to compete easily in the long term. DeepSeek is the evidence at hand, and Ismail imagines an eventual Wikipedia-like system to which many people contribute; Diamandis compares the trajectory with open-source web servers, asserting that 99.9% of web servers are now open source.
A pure foundation-model company may ultimately resemble a telco: enormous capex creates indispensable infrastructure but does not guarantee value capture. The episode’s cleanest analogy is, “You can’t build an Uber without the Internet being everywhere, but Verizon doesn’t get a cut of Uber”; applications built above intelligence infrastructure can own the customer and economics.
You.com therefore fine-tunes open models and federates outside ones rather than spending heavily to train everything from scratch. Its “trust layer” combines public and internal company data, citations that jump directly to highlighted source text, user certifications and training, and models taught to say “I don’t know” instead of fabricating missing information.
That model-agnostic layer “future-proof[s] organizations” against getting trapped in a one-year contract when a better model arrives two months later. Current deployments include cybersecurity companies such as Mimecast, publisher-specific assistants, and universities bringing cohorts as large as 30,000 students onto the platform—forcing professors to rethink assignments models can solve immediately.
4. AI-for-science turns discovery into a guided search loop
Socher says science and medicine enjoy unusually broad support because people do not want more human labor in those fields; “they just want more breakthroughs and cool discoveries.” For now, humans guide AI toward valued questions, then increasingly let it generate and test hypotheses in automated loops.
Peter cites an AI co-scientist reportedly replicating 10 years of antibiotic-resistance studies in 48 hours, then relays Dario Amodei’s forecast of “a century’s worth of biomedical research in the next 5 to 10 years.” The associated lifespan doubling is presented as a possibility, not a promise.
Socher’s strongest specimen is his Salesforce protein-language project, begun in 2018. The team synthesized proteins 40% different from naturally occurring ones; Frances Arnold’s Nobel-winning directed-evolution process had reached 3%. Proper folding and the predicted desired properties indicated that the model had learned “the syntax, the grammar of these proteins.”
Robotic laboratories can connect generated hypotheses to physical experiments, while MatterGen-style prompts ask for materials with specified elements, cost, manufacturability, or superconducting behavior. Quantum computing could expand what is simulatable, though Socher retains Sabine Hossenfelder’s warning—“we’ll see if they really can scale it”—amid competing trapped-ion, neutral-atom, and topological-qubit approaches.
5. The compute glut debate turns on Jevons’s paradox
Ismail believes data centers are being overbuilt because each facility reflects the model sizes and costs expected when planning began, perhaps six months earlier. DeepSeek suggests the incremental effort to produce the next generation is falling roughly 10x each cycle, so completed capacity may target a training requirement that no longer exists.
Socher “mostly disagree[s]”: greater efficiency lowers the price of intelligence and therefore expands its uses. Everyone may consume compute through a personal assistant, tutor, and health team; abundant energy also converts apparent resource shortages into engineering problems—salt water, for example, becomes usable water through desalination.
Their partial convergence separates total energy from particular data-center assets. Ismail expects society to use every available unit of energy but fewer conventional data centers than forecast; Socher concedes that facilities still require enough data and workloads to occupy them, or the industry risks “a real-estate crisis” of buildings without tenants.
6. AI startup pricing is outrunning proof of product-market fit
Ismail offers three explanations for OpenAI’s executive departures: newly valuable researchers can pursue personal missions; some dislike a “move fast and break things” culture; others fear capability is advancing without enough wisdom. Socher zooms out to California’s unenforceable noncompetes: once costly research proves something possible, knowledgeable employees can reproduce it more cheaply elsewhere.
Thinking Machines brings together Mira Murati, John Schulman, and other experienced builders, but Socher hopes it does not “just build another LLM.” At a discussed starting valuation of $30 billion, his investor framing is stark: “seed-stage risk combined with late-stage returns,” with a correspondingly extreme revenue hurdle.
Socher divides opportunity between a small horizontal infrastructure layer and thousands of vertical applications. He reports his first fund at roughly 5x TVPI after about four years and recalls investing in Hugging Face at a $5 million valuation before it reached $4.5 billion; vertical bets require both genuine AI skill and intimate knowledge of what buyers need.
The strongest moat is a “virtuous data cycle.” Autonomous-driving startups paid humans for every training mile, while Tesla owners bought cars and generated data for free; You.com similarly learns from explicit reactions to good answers, bad answers, and disliked passages, letting product usage continually improve the system.
7. Robots will diversify before one humanoid form wins
Ismail’s pushback is persistent: he asks why a robot needs human form and favors wheels and seven arms, arguing that one fast robot could do the work of seven. Socher similarly asks why a musculoskeletal humanoid is needed when specialized machines do particular jobs, quipping, “If you want a musculoskeletal humanoid robot, you get a man and a woman, have a baby, and grow the baby.” Diamandis imagines owning multiple household robots.
Socher refuses the binary. Machina Labs’ massive arms can form sheet metal and ship something like a factory into the field where a part is needed 200 rather than millions of times; humanoids may coexist with such specialized machines because robotics “doesn’t have to be zero-sum.”
Diamandis calls Unitree a possible “black horse,” with wheeled quadrupeds that jump, climb, and move rapidly. Socher jokes that everyone is building the original Terminator while nobody attempts a T-1000, then says a recent discussion with a hardware hacker made his own shape-changing concept seem potentially workable.
Peter highlights NEO Gamma and Figure’s Helix, while acknowledging polished demos omit “the 37,000 shots that went wrong.” Homes are far less standardized than roads, making manipulation harder and more capital-intensive; early teleoperation may also mean a remote worker can see children, doors, and private spaces while collecting training data.
8. Crypto resilience is improving faster than crypto usability
Socher holds only limited crypto exposure and sees it mostly as a distraction from AI. Agents may use cryptocurrencies, but they can also use credit cards; when he experimented, gas fees felt nearly as expensive as card fees, undermining the case unless transaction costs become almost free.
Ismail sees the billion-dollar Bybit hack as a resilience test the ecosystem handled better than in earlier years. A cold wallet holding Ethereum was drained, but on-chain movement remained visible and the exchange promised to make customers whole; debate even reached whether Ethereum could roll the transactions back before the theft.
Socher identifies the consumer-finance trade-off: a stolen credit card can be disputed and reimbursed, whereas decentralization also decentralizes “the risk, the security…and the liability.” Diamandis says that using a Trezor or Ledger feels nontrivial even to him, while Mt. Gox is cited as an early warning about keeping major wealth on centralized exchanges.
Strategy acquired another 20,000 Bitcoin for roughly $2 billion. Ismail says around 80% of a year’s upside may occur on five trading days—and admits he “spectacularly” missed four of five—so his advice is to buy what one can and “close your eyes for 10 years.” Peter and Ismail also emphasize Saylor’s unusually compelling evangelism.
9. Agents automate workflows before they become autonomous proxies
Diamandis frames actions as another sequence that can be learned through imitation and exploration; Socher says users have already created more than 50,000 custom agents on You.com. Diamandis’s compact definition is “white-collar job description”: once a repeated knowledge workflow can be explained clearly, a model can often execute much of it.
The clearest example is a marketer turning each product-release PDF into competitive research, two industry email campaigns, and three LinkedIn posts. The agent repeats those steps whenever a new file arrives; journalists can automate research across many specified sources and synthesis, while venture firms can encode recurring data-room analysis.
Booking travel exposes the autonomy gap. As a poorly paid graduate student, Socher would accept a 10-hour layover to save $200; today he may spend thousands extra for a direct flight. An agent must learn these shifting, contextual preferences before “boom, boom, boom, now it’s done” becomes credible.
A true Jarvis faces four blockers: consent laws restrict recording others; Microsoft’s screen-watching concept triggered a privacy backlash; young AI companies lack the trust required to observe everything; and websites may block agents that bypass advertising. “Your AI assistant doesn’t get distracted” by an advertisement for your next vacation while booking a work flight to Utah—threatening how Expedia, Amazon, and much of the web monetize attention.