Pioneers Insight Method Research Author
From Job Displacement to AI Trainers, Brendan Foody on Work in the AI Age
Back to Episodes

From Job Displacement to AI Trainers, Brendan Foody on Work in the AI Age

Summary

  • Mercor has raised $100 million and surpassed a $100 million revenue run rate while recruiting thousands of people for top AI labs. Founded in 2023, it began as a general talent-matching business, but human data is transitioning from crowdsourcing low- and medium-skilled workers toward vetting highly capable people who can work directly with researchers at the capability frontier.

  • Agent evaluations are the bottleneck between impressive benchmarks and economically useful automation. Acing SWE-bench is far removed from replacing a software engineer who coordinates with product teams, chooses tools, and exercises taste. Foody expects a years-long, industry-specific build-out because “the evals are upstream” of capabilities that labs and application companies want.

  • Mercor’s two compounding loops are talent supply and customer performance feedback. Its free career tools address labor marketplaces’ typical 50-to-1 supply-demand imbalance, while outcome data improves predictions about who will excel. Foody believes the less-obvious data flywheel may ultimately matter more than the marketplace network effect.

  • Foody expects knowledge-work displacement to arrive quickly, painfully, and politically. Customer support and recruiting are already areas where displacement is being reported, though much has not happened yet; physical work and roles valued for human interaction should automate more slowly. His survival heuristic is versatility, because verifiable skills such as math and soon code “will get solved very quickly,” while taste and founder judgment have sparse feedback.

  • Reinforcement fine-tuning could unlock enterprise agents with only hundreds to thousands of examples. Unlike supervised fine-tuning, RFT specifies the desired outcome and rewards the model for discovering how to produce it. Foody argues that base models already have the necessary reasoning capabilities; what they need is company-specific knowledge of tool use and “what good looks like in that role.”

  • Creating evaluations could become the world’s most common knowledge-work job—even though workers are helping automate themselves. The economic logic shifts labor from repeatedly performing a task, a variable cost, toward defining an evaluation once, a fixed cost. That fixed-cost opportunity lasts as long as there is a frontier for human evaluation; if models become superhuman, humans may become unnecessary in many other parts of the economy as well.

Deep dive

1. Expert evaluations replaced commodity labeling as the scarce input

  • The operating snapshot is unusually compressed: Mercor, founded in 2023 by three college dropouts and Thiel Fellows, had raised $100 million and passed $100 million in revenue run rate. Foody says its LLMs review résumés, conduct interviews, and predict job performance “better than a human can.”

  • Mercor began as a general matching business for talented people who were not getting opportunities through manual hiring. Its wedge changed when human data began transitioning from crowdsourcing low- and medium-skilled workers producing “barely grammatically correct sentences” toward vetting the most capable people to work directly with frontier researchers.

  • Demand ranges from consulting and software engineering to hobbyists and video games. Reinforcement learning can improve capabilities once there are evaluations for them, making Foody’s market map simple: “the evals are upstream of all of that.”

2. Talent prediction gets valuable where performance follows a power law

  • Mercor already sees models outperform human hiring managers on most of its evaluations. Foody expects it may become “almost irrational” not to listen to model recommendations, even if legal reasons leave a human responsible for final approval.

  • The economics are strongest in power-law professions: identifying a 90th-percentile engineer—or someone costing half as much who performs in the top quartile—creates disproportionate value. Investing is more power-law than software engineering; factory work is considerably more commoditized.

  • Text-heavy assessment should automate first because models can interrogate candidates and analyze transcripts at superhuman scale. Passion, persuasion, and sales ability depend more on multimodal signals, while repeated hiring for one role produces cleaner performance attribution than comparing 20 people doing 20 different jobs.

  • Hiring also leaves online evidence underused: GitHub repositories, college blog posts, personal projects, and design portfolios contain substantial signal that managers lack time to review. Foody cites subtler signals too, such as international candidates who studied in a Western country tending to collaborate or communicate better, and intrinsic motivation visible in a thesis or side project. Mercor ideally tests the work itself—such as a scoped-down MVP—and uses proxies when the real outcome takes too long to observe.

3. Verifiability determines which human skills lose value first

  • Foody expects displacement in many roles to happen “very quickly” and become both painful and a large populist political problem. Customer support and recruiting are already areas where displacement is being reported, though much of it has not happened yet; he expects more to occur imminently, especially as companies focus on efficiency during contractions.

  • He expects more work in the physical world: robotics-data collection, restaurant service, therapy, and other roles where human interaction matters. Physical automation should move more slowly because it lacks the virtual world’s rapid, self-reinforcing improvement loops.

  • His durable-skill advice is versatility: models repeatedly master activities people assumed would remain difficult. Math and soon code should progress fastest because answers are verifiable; taste in a founder is harder because the signal is sparse, though cross-domain reasoning still transfers once a new field has enough data to start learning.

  • Foody would probably not push a young child toward computer science, though he is not totally against it. He favors intellectually stimulating passions, general reasoning, entrepreneurial experimentation, contrarian views of missing markets, and product taste—not a wager that people who can code remain uniquely valuable in five years.

4. Economic work requires agent evals, not another academic benchmark

  • Historical evaluations resemble zero-shot academic questions; real jobs are end-to-end systems. A software engineer must interpret product needs, coordinate across teams, use tools, and translate competing priorities into output—leaving “a lot between being really good at SWE-bench and replacing software engineering.”

  • Evaluation creation therefore needs to be industry-specific. Customer support is a tractable starting point because workers share an interface and limited tools; software engineering is harder because architecture, coordination, and taste make even some otherwise verifiable domains a years-long build-out.

  • Foody “would not be surprised” if evaluation creation became the world’s most common knowledge-work job. Existing employees could encode what good looks like while marketplace contractors expand coverage, moving economic activity from repeatedly doing tasks toward the one-time fixed cost of teaching agents.

  • Elad relays a friend’s Nyquist analogy: humans may recognize that something is smarter than they are without being able to measure how much smarter it is. He then asks whether physician feedback could eventually degrade a model already better than individual physicians, citing Med-PaLM 2. Foody replies that models should learn to separate experts’ valuable knowledge from their mistakes.

5. RFT turns enterprise customization into an outcome-design problem

  • Models may help propose their own evaluation criteria, with humans validating them, but Foody still expects domain experts to ground the process. Frontier specialists could earn extraordinary sums for narrow problems where only a handful remain better than the model.

  • Broader agent workflows preserve a role for more typical professionals: communicating a diagnosis, coordinating tools, and sending follow-ups require more than answering one medical question. That distinction makes expert compensation increasingly power-law without immediately excluding the middle of the distribution.

  • Foody also suggests models may soon become much better managers: breaking large problems apart, assigning work, and performance-managing humans. Sarah Guo describes the product gap as an assistant that, given enough context and objectives, perfectly prioritizes, tasks, and sequences her work; Elad compresses the desired behavior to “tell me what to do for the next 3 minutes.”

  • The underlying reasoning may already be sufficient. Reinforcement fine-tuning teaches tool choice and company-specific definitions of success by rewarding desired outcomes; unlike SFT’s input-output pairs, Foody calls it “profoundly data efficient.” Elad characterizes the practical scale as hundreds to thousands of examples rather than a billion tokens, and Foody agrees.

6. Mercor is building toward one market for humans and agents

  • Mercor’s priorities are attracting “all of the smartest people in the world” and predicting their job performance. Because ordinary labor marketplaces have roughly a 50-to-1 supply-side-to-demand-side ratio, free mock interviews, career advice, and shareable profiles can create value for candidates who never receive a job; the other side of the business funds those tools.

  • The marketplace network effect is visible, but Foody expects the performance-data flywheel to become more important: customer feedback reveals who succeeded and why, improving subsequent selection. The long-term product could match each problem with some coordination of human workers and AI agents.

  • Foody says earlier shared hiring systems such as Hired could aggregate résumés and connections but could not record and conduct interviews at scale or analyze the data needed to explain job performance. LLMs create the “why now” for a more complete, automated labor marketplace.

  • His endpoint is a global unified labor market. Today candidates apply to a dozen jobs while a San Francisco company considers a fraction of a percent of global talent; software-cost assessment could let every candidate access one market and every company hire from it.

  • For startups, Foody recommends uncompromising talent density early, then data-driven standards as the organization scales. Elad relays Valley-era rumors that Google hired well but took years to remove poor performers, while Facebook had a mixed early talent pool but was better at removing underperformers.