Matt Fitzpatrick: Who Wins the Data Labelling Race & Why Al Needs Forward-Deployed Engineers
Matt Fitzpatrick: Who Wins the Data Labelling Race & Why Al Needs Forward-Deployed Engineers
Summary
- The core call: enterprise AI adoption is “a decade, not two years.” Models improved 40–60% on public benchmarks in two years and KPMG says 60% of consumers use genAI weekly — yet MIT finds only 5% of genAI deployments working in any form and Gartner sees 40% of enterprise projects canceled by 2027. Fitzpatrick’s diagnosis: the bottleneck is data infrastructure, workflow redesign, ownership and “most importantly, trust” — banks will run genAI through model-risk-management-style validation first, the rest of the enterprise five to six years after.
- “You just cannot do this with out-of-the-box SaaS.” Invisible sells nothing upfront — free 8-week solution sprints, payment only at user acceptance testing, and zero charge for forward-deployed engineers (450 people, 8 offices). The structural point for SaaS investors: “the minute you had to bring in FDEs in a SaaS context, your economics broke instantly,” and Fitzpatrick argues out-of-the-box software “has always been a lie to some degree” — the endgame is hyperpersonalized software, not boxes.
- Internal builds are losing: MIT’s data shows externally driven builds are 2x as effective as internal ones. Exhibit A: an e-commerce retailer spent $25M on a returns agent, built its own eval on call-resolution speed plus sentiment (a hallucinated “$2 million refund” scores perfectly), then shut it down and reverted to a deterministic flow. His fix: three to four initiatives, led by operational leaders — “don’t locate it in the tech function” — paid as it works.
- The biggest industry misnomer: synthetic data will not replace human feedback. Synthetic works for base-truth domains like math; multi-step reasoning across 45 languages and multimodal contexts is “in the first inning,” and legal-grade data sits inside big law firms, not public corpora. Unlike ML, genAI requires humans in the loop for statistical validation — “you are going to need humans in the loop for decades to come.”
- Data labeling shakeout: 3–5 players, not winner-take-all. Concentration is structural (few LLM builders exist), but “people are willing to pay for good data” because one week of bad data burns enormous compute; claims of total price insensitivity are “an exaggeration.” The moat is Helmer-style institutional memory — a digital assembly line that sources 26,000 selected experts on 24 hours’ notice — and yes, he insists the big numbers are revenue, not GMV.
- Capital posture flipped: Invisible raised only $7M primary in nine years; it has now raised $130M and will not be profitable this year — “the greatest environment for growth that has ever existed… I hope we never get to the harvest stage.” Next frontier: physical-world data (FDEs dropping Starlink terminals on farms for herd-safety computer vision) and robotics.
- Agent skepticism, quantified: Salesforce AI research puts out-of-box agents at 58% accuracy single-turn and 33% multi-turn; AWS reported 70% of “agents” are traditional scripting. His contrarian allocation for a $400M fund: look beyond out-of-box agents and back AI-native businesses that serve the customer need directly — YC’s recent class, he thinks, did 2x the revenue of any prior class doing exactly that.
Deep dive
1. Enterprise AI Adoption Lags
- Fitzpatrick sits on both sides of the market — Invisible trains all the large language models via RLHF and deploys enterprise use cases on its own platform — and his framing of the “cognitive dissonance” of the last two years: model performance up 40–60% on public benchmarks, consumer adoption exponential (KPMG: 60% of consumers use genAI weekly), while MIT reports 5% of genAI deployments are working in any form and Gartner expects 40% of enterprise projects canceled by 2027.
- The reason is everything around the model: “It’s the data infrastructure to support those models. It’s the redesign of workflows… and most importantly, it’s trust. It’s observability.” From a decade building credit models in banking, he expects enterprises to replicate what banks call model risk management — testing, training, validation — “that whole process is in the first inning.”
- The timeline call: “it’s going to take a decade, not two years” — banks and healthcare do the testing-and-validation phase first, “then the rest of the enterprise will be over the next five, six years after that.”
- Stebbings’ own anecdote lands the point: after speaking at one of the world’s largest banks, he messaged his team “they’re toast” — because the CTO laughed off his off-the-shelf tool over data, security and permissions. His concession: “everything that you just said there, I listened to.”
2. Internal AI Builds Fail
- The MIT report’s buried stat, per Fitzpatrick: externally driven builds are 2x as effective as internal team builds. His 10-year pattern: enterprises bought 15 apps, then cloud brought custom wrappers, and “genAI has 5x’d that” — internal teams handed enormous budgets with none of the discipline applied to vendors: deliverables, timelines, ROI, milestones.
- The talent problem, stated carefully: “the amount of talent that knows how to do this well is not large” — and it works at AI startups and big tech, so enterprises reasoning from first principles carry real risk.
- His best specimen: an e-commerce retailer spent $25 million building a returns agent and defined success with a home-built eval mixing speed of call resolution and sentiment. “The problem with that is what if the agent hallucinates and says, ‘Here’s $2 million.’ That actually gets resolved quickly and the person’s happy.” A couple of months later they shut it down and moved back to a deterministic flow — “that’s not surprising to me at all.”
- The correction he expects over two years: the CFO function imposes guardrails — ROI, metrics, milestones. His playbook: pick the three to four things that move the needle, assign your best four operational leaders, “don’t locate it in the tech function,” anchor payment to outcomes. The leader “doesn’t have to be highly technical” — just the same muscle memory used on any vendor.
3. Forward-Deployed Engineering Wins
- The buyer’s reality: 250 vendors a week pitching, all sounding similar — a customer literally opened a meeting with “how are you different than the other 250 people that have pitched me this week?” And many don’t work: Salesforce AI research testing out-of-the-box agents found 58% accuracy on single-turn and 33% on multi-turn workflows — “which means they don’t really work.”
- Invisible’s answer: “We don’t actually sell anything.” We meet a customer, we say we will do it for free for eight weeks and prove to you the tech works." Asked whether he can serve enterprise without an intense forward-deployed mechanism: “I don’t think you can… you just cannot do this with out-of-the-box SaaS.” The company has doubled down — 450 people across eight offices — and charges nothing for FDEs, unlike competitors: “I spend less on sellers and more on forward-deployed engineers. That’s my simple math.”
- His definition matters: FDE done well is “executing a very specific workflow build” in about three months — not solutions engineering, and not Accenture-style three-year builds. For startup founders the razor is simple: a knowledge repository (public filings, healthcare reference) doesn’t need FDEs; “if you’re trying to change workflows, you do need FDEs” — and adoption/workflow embedding is the hardest part.
- The company he most respects is Palantir: “they realized 10 years before the rest of the tech market that forward-deployed engineering customization would be important… a very countercultural leap” when tech services was where nobody wanted to play.
4. AI Software Requires Customization
- Founder Francis’s founding principle: “If there’s an app for everything, how come nothing works?” The Accenture paradigm of the last 20 years: buy 50 apps, pay $200M over two years to layer them together, end up with “no working data, no linkages” — layers of sediment. Insurtechs like Duck Creek did really well with momentum and push from the SIs that got them going.
- His heresy for SaaS investors: look at how much of any large public software company’s revenue is actually services — “out-of-the-box software has always been a lie to some degree… they just dressed it up.” GenAI breaks the model further because what gets built is hyperspecific to each customer: “it’s not a box.” Payment happens at user acceptance testing, two to three months in — “machine learning has been around the enterprise… that’s always been a motion that looked like this.”
- The proof case: Lifespan MD, a concierge medicine business with data fragmented across EHRs, CRM, ERP and notes. Invisible’s five modular platforms — Neuron (data), Axon (agent builder), Atomic (process builder), an expert marketplace, Synapse (evaluation) — unify it in two to three months where “Accenture would take two years,” then layer conversational querying (“who’s used peptides, male between 36 and 50, and what have been the results”) and custom scheduling agents. His thesis: “you move from out-of-the-box SaaS to much more hyperpersonalization using the specific data of an individual customer.”
5. Data Labeling Is Specialized
- On being lumped with talent marketplaces, his reframe: “You have to be able to source a PhD in astrophysics from Oxford [on 24 hours’ notice], put them into a digital assembly line, and four days later generate perfect, statistically validated data that will be compared head-to-head to somebody else’s.” Invisible sees 1.3 million experts a year and keeps 26,000 selected experts who must start within 24 hours and produce perfect data.
- The moat, in Hamilton Helmer’s terms, is institutional memory — his favorite example being the Toyota production system, which Toyota could describe openly and nobody could replicate. “It is a digital assembly line no different than an auto factory,” plus five years of data on who’s been good at what task.
- The sector’s arc: five years ago it was “cat-dog commodity labeling” run on Google Sheets; now it’s validating “an architectural expert on 17th-century French architecture who speaks French” on a day’s notice. Specialization and unbundling into insanely niche supply pools is real — the board member Stebbings quoted didn’t see it coming, and Fitzpatrick says “absolutely.”
- On pay: “I think of our business like Uber” — price discovery, where the rate depends on market context such as rain or location. The nuance: “you can overpay a really bad expert and that is a total waste of everyone’s time” — the value is knowing the difference between a $150 and a $130 expert. And finite supply doesn’t matter: “the expertise needed varies so much month-to-month that if you bottled up whatever supply it is, it would change in three months.”
6. Data Markets Stay Concentrated
- On the two-players-over-50%-of-revenue pattern: he won’t disagree — “there are not that many players that are actually building LLMs, so by definition the whole space has concentration.” His counter is diversification: 2024 was materially weighted to AI training (the talent-marketplace side was “a pretty material percentage”), but Invisible has confirmed 12 enterprise deals in the last 45 days.
- The negotiation staring contest resolves on quality, not leverage: “people are willing to pay for good data… one week of bad data burns a lot of compute.” Bake-offs are routine — a multimodal audio model comes in, “we go head-to-head with somebody that week, and at the end of it we win or we lose.” But the board member’s claim of drastic price insensitivity gets pushback: “I think that’s an exaggeration… there’s actually fairly standard price bounds across all the players.”
- The shakeout call: “I don’t think the answer is one player… most of these markets end up with three, four or five players. I don’t actually think it’s even two” — with specialization by task (coding, specialist tasks, PhDs). In enterprise, historically “it’s been Palantir, not many others,” which is why alternatives generate so much excitement.
- On Stebbings’ revenue-vs-GMV challenge (he “got battered” for calling it revenue on prior shows): “I think it is revenue.” The Airbnb distinction: Airbnb takes one consistent fee; here “there’s huge variety depending on the project, the expertise type” — no consistent rate relative to booking.
7. Human Feedback Remains Essential
- The biggest misnomer in the industry, as he tells it: that synthetic data takes over in two to three years. “From first principles that actually doesn’t make very much sense.” Synthetic data suits base-truth domains like math; multi-step reasoning across audio, video and 45 languages — “computational biology in Hindi versus French versus English with a southern accent” — is a permutation space “we’re still in the first inning of.” His categorical version: “human feedback is going to be important… for the next decade. I have a strong belief on that.”
- Legal is his carrying example: “a lot of the legal data in the world exists with big law firms — it doesn’t even exist in the public,” and the public corpus “has been commoditized for years.” The deeper distinction from his QuantumBlack days: ML models can be backtested to statistical validity without human intervention; “on the genAI side, you are going to need humans in the loop for decades to come.”
- On the Gemini 3/Opus 4.5 benchmark churn Stebbings mocked: models are “unequivocally” improving, but they’re moving to hyperspecific tasks “where there’s not a public benchmark by definition” — how well does the model build an LBO model, does an IC memo hit “99% precision” for one specific PE firm. Hence the call investors should note: benchmark progress is “almost orthogonal to enterprise uptake,” which depends on trust and precision on specific tasks, not generalizability. Built expertise compounds — the likely SAIC/Vantor/U.S. Navy partnership fine-tuning a model for underwater drone swarms (react, pull back, alert, engage) is his example of switching costs earned through logic.
- On junior-talent hollowing-out, he’s unworried: the 5-to-10-year adoption curve leaves time to react, and new grads are “some of the highest adopters” — he’s hiring more of them. His analogy: accountants went from slide rules to Excel and headcount didn’t fall — Jevons paradox, “you actually had way more accounting,” and every FP&A function is probably larger than 25 years ago.
8. Capital Funds Physical-World Expansion
- The capital flip: Invisible raised only $7 million of primary capital in its entire nine-year journey; it has now raised $130M (initially announced as $100M) and will not be profitable this year. His logic: “you can either harvest capital or invest capital… I think we’re in the greatest environment for growth that has ever existed… I hope we never get to the harvest stage.” The decision he’s scared of but thinks about often: whether to pursue hyperscale toward “$50 to $100 billion” — which requires investing far more per customer. On Mercor’s $2B raise: no envy — his spend goes to enterprise and core software platforms, “a little bit different than what others are focused on.”
- Where he isn’t investing but wants to: physical-world interactions. The specimen: serving one of the largest US agricultural conglomerates on herd safety — FDEs visiting farms, “dropping Starlink terminals into those farms and building out custom computer vision models” to decide when to send a vet. Oil rigs next; robotics “will take longer but will be really interesting when it works” — and needs task-specific robots, not broad-based ones.
- Handed Stebbings’ hypothetical $400M fund, his contrarian allocation: the model layer produces returns, the agent layer is complicated, the application layer “tricky too.” The interesting question for society: “whether new companies built around AI get distribution faster than big companies figure out how to adopt AI.” His bet: AI-native businesses serving the customer need directly — genAI-native services in tax and accounting, physical-world operators — noting YC’s recent class did “2x the revenue of any prior class,” he thinks.
- On whether 70–80% software margins are over: “First of all, challenge that 70–80% software margins actually ever existed.” Public software multiples went from 20x to 10x in two years as profitability exposed slowing growth; meanwhile “the integrated units will be very, very profitable” — faster customer acquisition, faster builds, no box stickiness.
9. Trust Beats Faking It
- When he took over, “if you looked at the entire public internet, I think there was one article available” on Invisible. Now he’s on the road 70% of the time, guided by a Marc Andreessen idea: “when private and public narrative diverge, that is the risk or the opportunity” — hypothetically claiming an out-of-the-box agent that does everything, when it isn’t true, creates opportunity for rivals. Stebbings’ pushback — worth keeping: “Is that not our industry? … our job is to sell and then deliver later.” Fitzpatrick’s line: “I want a company we work with to know that if I say this will work, it will work. You only get one chance to do that.”
- Non-determinism is why fake-it-till-you-make-it got riskier: contracts get signed for “50 agents to be delivered — but then the question is, do you deliver the agents? Do they work?” He cites an AWS report from that day: 70% of “agents” are actually traditional script-writing and automation — “that’s why I don’t self-identify as an agent company at all… agents are one tool in the toolkit.”
- His own confession, from the McKinsey years 12 years ago when it was still called “data analytics”: he set a vision without knowing what he’d build, on conviction that fragmented data made enterprise decisioning broken — 70% of software in America is over 20 years old, the average bank spends 93% of tech cost on maintenance (Stebbings adds: one bank has 6,500 people in KYC alone). The counterintuitive lesson: “I didn’t fake it… my entire approach would be to say I think this would work, this is my reasoning why — and let’s build this.” People trust that; “an out-of-the-box AI that solves all your problems” triggers skepticism.
10. Strategy Changes; Recruiting Matters
- The topic he thinks about most: “if you get amazing people, everything else will follow.” Not just hire — “hire, retain, and evolve great people,” ignoring roles: “really good people will run five to six different roles… hire great all-around athletes.” His sports analogy invokes likely Nick Saban: “He did not build Alabama football with the process. He built it with recruiting the best football players in the country.”
- Stebbings’ Revolut pushback — Nik’s doctrine that winning is culture and “brutality in bounds drives humans” — draws a notable concession: “No, I think it’s actually right” for scaling a consistent business model. The caveat: “a lot of what we do is research and exploration fundamentally… it is a research culture as much as an implementation culture” — though delivery and ops are “in war mode quite a bit of the time,” and great engineers should get 30% of time on new projects.
- Remote is no longer the default at Invisible: fully remote for nine years, now offices in New York, San Francisco (the old Pinterest space), London, Paris, Poland, DC and Austin. Productivity up “exponentially”; he tripled the engineering team this year and found “the vast majority wanted to be in person” — especially younger tenures — without mandated attendance. His balance: seven-day availability is separate from collocation, and six days a week in-office “is overkill and you lose great people.”
- Two management beliefs he abandoned: central control (“a bit of a fallacy” — flatten hierarchy, empower teams at the edge with consistent tooling, as armies do), and strategy itself: “in the AI world at least, strategy is a somewhat overrated concept… every three months the entire world changes.” Hold core beliefs plus 30–40% constant iteration; “five-year strategic planning is not a useful exercise right now” — five years is for culture and institutional memory. His optimist close, with numbers: AI data centers are just 0.25–0.5% of data centers’ 1% share of global electricity (air conditioning is 14–20%); US healthcare spends $14,000 per capita — 2.5–3x Germany or Canada — with 250,000 deaths a year from avoidable errors (Johns Hopkins) that AI can attack; and education excites him most: Invisible assesses “an enormous amount” of hires who never went to college “on cognitive aptitude and skill.”