Which Industries Survive AI, The New AI Benchmarks, and the 2026 Recursive Learning Timeline | #218
Summary
AI will disrupt documentation-heavy sectors first, not flatten every industry at once. Fitzpatrick puts media, legal services, and business-process outsourcing in the blast zone, while oil and gas or real estate retain much of their underlying function; the competitive question is whether “the startups get distribution before the big companies build the technology,” especially in banking, where much of the application footprint is over 20 years old.
The missing enterprise-AI asset is not another broad model score but task-level benchmarks. Public coding benchmarks show models improving 50% to 100% across many dimensions over three years, yet businesses need “accuracy or human equivalent on a specific task”—from claims processing to title insurance—and an “80% accurate very smart deployment” still carries too much production risk. The discussion points toward thousands of narrow benchmarks, whose ownership could create visibility and leverage in neglected industries.
The practical 2026 playbook is to follow value, pick two or three use cases, and put an operator—not the technology department—in charge. Fitzpatrick recommends reaching a working prototype in roughly a month, tying it to metrics such as CSAT, inventory days, stockouts, or cost per call, and asking: “Would you bet your annual bonus that whatever use case you deploy works?” For the first project, he favors an outcomes-based RFP to a third-party vendor.
Full autonomy is the wrong initial architecture for most enterprises. Klarna reportedly said its AI handled 2.3 million calls a month, did the work of 700 full-time agents, and could save $40 million annually, only to reverse course 8 to 12 months later; Fitzpatrick’s lesson is to route routine work to agents while escalating complex refunds, system write-backs, and customers seeking human contact. “Human in the loop is going to be a feature, not a bug.”
General-purpose models may commoditize model-building, but company context, custom evaluation, and data security remain enterprise concerns. Fitzpatrick expects enterprises to tailor frontier models rather than pre-train their own, with swappable models sitting beneath company-specific documents, workflows, and benchmarks. The key is separating genuinely proprietary information—such as Jane Street’s trading data—from ordinary back-office data, using on-premise or small language models where warranted instead of treating everything identically.
The strongest deployments begin with narrow data assembly and already show measurable operating gains. Invisible combined 750 SwissGear tables, expanded overall inventory coverage about 30%, and doubled the number of SKUs with reliable predictions in a couple of months; it also built player-movement models for the Charlotte Hornets, a HIPAA-compliant control tower for Lifespan MD, and underwater-drone decisioning for the U.S. Navy. The repeated pattern is specific data, a bounded decision, and an observable result.
The episode’s central disagreement is whether recursive self-improvement soon eliminates specialized human feedback. Wissner-Gross gives a “two to three years max” conservative outer bound for an AI researcher matching or surpassing human model researchers, while Fitzpatrick argues that tacit expertise, company-specific work, new modalities, languages, robotics, and multi-step reasoning keep creating fresh evaluation needs. He leaves open the possibility that useful training frontiers could run out 10 to 15 years from now, but argues that human expertise and human-in-the-loop work remain important for a long time.
The 2026 architecture call is multi-agent, multimodal, and simulation-heavy, with government process automation among the largest potential beneficiaries. Fitzpatrick expects task-specific agents orchestrated by an LLM, more audio/video/image interaction, and “mirror worlds” or RL gyms that test systems before real deployment; cited estimates suggest AI-assisted permitting could cut energy and data-center timelines 50%, while licensing, benefits, and compliance cycles might shrink 70%.
Deep dive
1. AI disruption will concentrate where documentation is the product
Fitzpatrick rejects the premise that every sector will be hit equally. Media, legal services, and business-process outsourcing produce large volumes of knowledge work and documentation—the activities most directly exposed—while Wissner-Gross carefully narrows his own claim: “Knowledge work as we currently know it” is cooked, not necessarily knowledge workers or entire companies.
Oil and gas or real estate look materially different because their core functions remain. Fitzpatrick argues that deciding which apartment or office building to buy will work much as it did five or six years ago, even if AI improves supporting analysis; companies should identify the portions of their operations that can genuinely change rather than declare the whole business “AI-first.”
The competitive race is whether “the startups get distribution before the big companies build the technology.” Banking embodies the tension: much of its application footprint is north of 20 years old, while newer fintechs such as Revolut can build differently without the same modernization burden. Fitzpatrick does not claim to know which side wins.
Capability is a separate constraint from industry exposure. A 50-person company may lack even a CTO, while established IT teams may not possess relevant skills; even knowing Python can leave gaps. Fitzpatrick’s advice is unsentimental: companies unable to hire or develop the expertise should rent it through partners rather than pretend every firm can build internally.
2. Measurable baselines determine where adoption can move safely
Mortgage underwriting advanced because banks can back-test decisions against a statistically valid baseline and check whether the resulting credit decisions work without redlining. Contact centers offer similarly legible measures—time per call, cost per call, and CSAT—which should have made them unusually favorable territory for AI.
Legal work splits between advice and commodity production. Fitzpatrick expects high-end counsel to persist for a large M&A transaction, while routine NDAs and standardized documents compress sharply; Diamandis notes that venture documents repeatedly carry a $50,000 legal cap, contain “about eight knobs,” and somehow run to “$49,999.99” despite being nearly identical.
Klarna became the cautionary example. The company reportedly said its AI handled 2.3 million calls monthly, replaced the work of 700 full-time agents, and would save $40 million a year; roughly 8 to 12 months after becoming a flagship agentic-success story, it announced a return to human contact-center agents.
Fitzpatrick offers hypotheses, not inside knowledge: some customers simply demand another person, while refunds and other non-first-line cases require complex write-backs into source systems. His confusion is architectural—the sensible design was always a changing mixture of agents and humans, not “all humans, all agents, back to all humans.”
3. Enterprise AI needs a value-led operating system, not a strategy document
A CEO facing the board’s “What’s your AI plan?” should “follow the value.” Fitzpatrick would select two or three levers capable of materially moving the business—customer service, FP&A forecasting, inventory management, or digital marketing—then take only one or two into a real pilot rather than invite unfocused experimentation.
Generative AI reverses the older machine-learning deployment pattern. A prototype can be running in about a month rather than after months of construction, but reliability emerges through intensive testing and validation afterward. A strategy deck is not progress; the decisive test is, “Would you bet your annual bonus that whatever use case you deploy works?”
For a first implementation, Fitzpatrick recommends an RFP to a third party compensated according to outcomes. An internal team may lack experience and cannot be held to the same “you get paid if it works” standard; an outcome-linked contract transfers some execution risk while forcing everyone to define success in advance.
Citing an MIT report that only 5% of enterprise models had reached production, Fitzpatrick identifies organization as a central failure point. Put the best operator in charge, outside the technology organization, and assign a business KPI: CSAT or call time for service, inventory days and stockouts for forecasting. Otherwise, “a thousand flowers bloom” into unaccountable science projects.
4. Narrow benchmarks become the control layer for enterprise AI
Broad public benchmarks, especially coding benchmarks, remain useful indicators of general progress; by Fitzpatrick’s reading, models improved 50% to 100% on most observable dimensions over three years. Their limitation is relevance: a business does not need abstract intelligence so much as “accuracy or human equivalent on a specific task.”
A contact-center benchmark should compare AI agents with the company’s own expert agents across representative calls. Claims processing needs its own human-equivalence set. An “80% accurate very smart deployment” may still be unacceptably risky, so each workflow requires a custom eval built around its actual errors, thresholds, and escalation conditions.
Diamandis hears an entrepreneurial opening: practitioners who understand both AI and a neglected domain such as title insurance can define and broadcast the benchmark before anyone else. In his framing, declaring credible ownership of an unclaimed evaluation category can make someone “an instant star,” because the benchmark may be harder to create than the post-trained model.
Invisible already builds customer-specific benchmarks around individual tasks. A generic sales agent cannot simply be purchased like conventional SaaS; it must learn the company’s products, knowledge corpus, selling method, and “way of speaking,” then be tested against a local eval that determines whether the tailored behavior is actually good.
5. Generalist models win the base layer, while enterprise context stays local
Wissner-Gross invokes the “infamous BloombergGPT moment”: Bloomberg possessed proprietary financial data and sought domain-superior performance, yet frontier labs’ generalist models reportedly leapfrogged the project within months. His challenge is whether internal data and post-training have a durable future if general models keep absorbing specialized capabilities.
Fitzpatrick distinguishes building a language model from adding company context. He does not expect individual institutions to pre-train their own compute-intensive LLMs; he expects them to tailor leading models using private documents, preferred outputs, and workflows, with an enterprise layer designed so newer foundation models can later be “dropped in.”
A law firm’s preferred M&A documents illustrate what the public model cannot know. That information still requires local tailoring, post-training, or evaluation. Ismail extends the argument: the trade-secret knowledge of “how do we do things?” may become a company’s most valuable edge, increasing the need for protection between internal data and the broader AI world.
Fitzpatrick resists treating every byte as sacred. Banks, hospitals, and trading firms may keep sensitive information on-premise or use small language models, but Jane Street’s trading data and its back-office forecasting data are not equally proprietary. The workable policy classifies data by actual competitive or regulatory sensitivity rather than refusing every external model.
6. Data readiness means assembling the minimum viable truth
An agent built on fragmented customer and product records “is going to break by definition.” Fitzpatrick therefore starts every use case with its inputs, but he does not endorse a five-year attempt to perfect the entire corporate data lake—the path many large organizations have already pursued without making every repository accurate, accessible, and coherent.
Credit underwriting may need five or six central categories: the credit itself, market conditions, company financials, the security of the credit, and related variables. It does not require every record across the commercial bank. The practical question is, “What data do I need for this specific use case?” followed by focused remediation.
Generative AI also elevates information that never lived in a system of record: video, images, free text, and other unstructured files. These are assets that people have not historically tried to master, so the first step is identifying the task and making the relevant data ready.
Blundin’s Vestmark example sharpens that distinction: reconciliation records show how an account ended up reconciled, not what the employee did to resolve it. An AI assistant can first observe and accelerate the workflow, then produce the human-feedback or tuning data needed for automation. Another bank CIO faced 300 guarded customer databases, each owned by a different product silo.
7. Domain deployments work when the decision and data are tightly bounded
For the Charlotte Hornets, Invisible fine-tuned computer-vision models over single-point footage from multiple college and international venues. Conventional statistics capture points, rebounds, or plus-minus; the model instead measured player movement, spacing, and who creates space across inconsistent camera angles, giving draft evaluators evidence on characteristics and player fit that transactional stats miss.
Lifespan MD began with data architecture, not automated diagnosis. Invisible’s Neuron platform assembled a HIPAA-compliant, multi-tenant view of patients, providers, and practice performance, enabling questions such as which longevity tests men aged 35 to 50 use most. Patient data remains at individual practices while clinicians and central operators receive only the access appropriate to them.
Fitzpatrick calls clinical decisioning a murkier target than administration. The U.S. spends roughly $13,000 to $14,000 per capita on health care versus $2,500 to $3,000 in Germany or Canada, with something like 30% to 40% going to administration. The nearer opportunity is removing scheduling and paperwork while making physicians “even more empowered.”
Other projects widen the pattern without changing it: with SAIC Vantour and the U.S. Navy, Invisible worked on decisioning around sensor-rich underwater-drone swarms; at SwissGear, it combined 750 tables, expanded overall inventory coverage about 30%, and doubled the number of SKUs receiving reliable forecasts in a couple of months. The inventory work was aimed at minimizing stockouts and excess inventory.
8. Human feedback is becoming more expert, not simply disappearing
Invisible’s Meridial business trains models, while its enterprise side builds custom applications. Wissner-Gross asks whether an AI researcher capable of constructing models, datasets, and benchmarks would erase the need for a marketplace of human ML contributors; Fitzpatrick replies that predictions of RLHF’s imminent disappearance have persisted for five years without matching deployment reality.
The work is changing from commodity “cat dog, cat dog labeling” toward PhD- and master’s-level evaluation, controlled RL environments, simulations, and RL gyms. Fitzpatrick says studies and operational experience favor pairing synthetic and human data, particularly for multi-step reasoning where hallucinations can compound across the chain.
His best specimen is deliberately narrow: a model studying evolutions in 17th-century French architecture, in French, still needs qualified humans to validate it. The same logic applies to a new legal dataset, where an associate or M&A lawyer must determine whether the resulting document is genuinely comparable to expert work.
Wissner-Gross’s pushback is about efficiency, not whether feedback has any value. Reinforcement fine-tuning may require fewer human hours than large annotation workforces, and an RL environment can scale once built. Fitzpatrick concedes the form will evolve but notes that post-training feedback is a small share of total compute cost and among its most valuable inputs.
9. Recursive self-improvement may outrun corporate absorption
Wissner-Gross gives “two to three years” as a conservative outer bound for an AI researcher as good as—or stronger than—the human researchers building ML models. He expects recursive self-improvement capabilities to accelerate in 2026 while corporations move “at a snail’s pace,” leaving implementation, change management, and workflow bottlenecks as the binding constraint.
Fitzpatrick leaves open that machines might exhaust useful training frontiers “10, 15 years from now,” but sees no near-term shortage: new languages, modalities, robotics tasks, legal subfields, and company-specific workflows continuously create narrower evaluation demands. General competence does not manufacture private precedent or tacit expertise that exists only in experienced people’s heads.
Sales is his counterexample to total generalization. The best sellers’ patterns are often undocumented, and a market full of 500 email-based SDR vendors may make genuine interaction scarcer and more valuable. “The human-touch elements become more and more important,” particularly where trust, exception handling, or authenticity drives the outcome.
Wissner-Gross frames the commercial opportunity as the gap between capability and adoption: companies such as Invisible can be “the lubricant” between frontier labs and enterprises. In his tests for contact centers, 80% of people preferred the AI, while the dissatisfied 20% could “torture the whole thing to death,” creating lucrative work in routing, evaluation, data, and exception handling.
10. AI-native challengers will redesign flows rather than automate job boxes
Ismail’s Canon thought experiment replaces departmental thinking with a functional flow: marketing, retail sales, registration, ink replenishment, predictive repair, upselling, and accounting become one AI-managed printer lifecycle. The product reports its own state and triggers the next action, potentially leaving humans “90% out of the loop.”
He calls today’s employee-by-employee automation “radio over TV”—like putting radio announcers on television to read the old scripts instead of adapting to the new medium. An AI-native company starts from the desired function and reconstructs the entire operating system, while incumbents tend to attach assistants to inherited roles and handoffs.
Diamandis returns to distribution as the incumbent’s advantage. His proposal is to invite global AI entrepreneurs to pitch how they would disrupt the company, fund the best five, grant them data and access, then acquire or take a majority stake in the ventures that work and make them the new company. The objective is “innovation on the edge” followed by displacement of the legacy center.
Ismail points to Apple’s small, secret teams as the organizational model, citing roughly 18 groups examining different industries and patiently iterating until an entry is ready. Large operators may struggle to attack their optimized core, but their knowledge of adjacent industries can support AI-native ventures aimed at neighboring markets.
11. Multi-agent teams, multimodality, and mirror worlds define 2026
Fitzpatrick’s first architecture call is the multi-agent team. Rather than one agent making every decision, enterprises train task-specific agents to high accuracy and place them under an LLM that orchestrates the broader logic. Contact centers are a natural example.
The second call is a multimodal leap. Audio, video, and images should become much larger parts of how people engage with models, moving interaction beyond historically text-heavy interfaces. The earlier basketball, health-care, and drone examples show why the enterprise data surface is already multimodal even when its software interface is not.
The third is the “mirror world” or RL gym: simulated environments and digital twins in which a coding, contact-center, or manufacturing system can execute tasks and be tested before touching production. Simulation supplies repeatable tests for systems whose real-world mistakes would be costly or dangerous.
Avatars are another likely 2026 development. Fitzpatrick says training one from his public statements would not be difficult and notes sports-related avatar work already underway; he expects people may prefer conversing with recognizable people via avatars over generic chatbots, making synthetic personalities a more natural part of interaction.
12. Human work shifts toward physical presence, authenticity, and new categories
Asked for the “last expert standing,” Fitzpatrick names broad categories rather than three precise occupations. Seismic work, drilling-site operations, real-estate selection, physical trades, and jobs built around human interaction survive longer because their function is not simply searching documents. Nearer-term disruption remains concentrated in BPO, legal services, and media.
He does not equate task disruption with falling employment. Media’s economics and channels changed, yet Substack, Medium, blogs, and other formats created more media entrepreneurs. Fitzpatrick cites estimates that roughly 25% of each U.S. high-school class eventually enters a field that did not exist during high school, while 20% of U.S. employment now consists of digital-ecosystem jobs.
Wissner-Gross offers three competing candidates for the final roles: politicians, because they make the laws; the greatest physicists and mathematicians, as the peak of human intellectual work; or occupations where customers demand an authentic human counterparty. Diamandis summarizes the third hypothesis as “tastemakers will dominate.”
Government may provide the clearest social return. Fitzpatrick cites a study suggesting AI-assisted permitting could cut energy and data-center implementation timelines 50%, and an OECD estimate that licensing, benefits approval, and compliance cycles could shrink 70%. He identifies project management and timelines for spending and infrastructure deployment as a simple, high-value use.