Apr 10, 2025 · 41m · no-priors
No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Mercor CEO Brendan Foody joins Sarah Guo and Elad Gil on No Priors to discuss how large language models are transforming talent evaluation, the impending economic impact of white-collar automation, and the critical need for agentic evaluation benchmarks in frontier AI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 30.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Foody firmly counters the popular narrative that rapid SWE-bench progress means software engineers are about to be replaced, emphasizing the massive unaddressed coordination and tool-use challenges.
Hardest push from the hosts ▶ 22:07 Challenging human feedback when models surpass expert panelsElad uses Google's Med-PaLM 2 to push back against the necessity of human evaluators, arguing that human scoring actively degrades model performance once AI surpasses median practitioners.
Biggest teaching moment ▶ 31:33 Explaining why RFT supersedes SFT in enterprise customizationFoody provides a deep technical explanation contrasting SFT data inefficiency with outcome-based Reinforcement Fine-Tuning (RFT) across enterprise customer use cases.
The host holds their own ▶ 37:31 Elad analyzing hiring versus firing dichotomies at Google and FacebookElad demonstrates seasoned startup expertise by contrasting Google's high hiring bar with Facebook's swift performance management during their formative growth stages.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Mercor's Core Mission and AI Lab Talent Demand | 4 | 6 | 1 | 1 | Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers. | |
| Power Law Distributions in Knowledge Work Performance | 5 | 5 | 2 | 2 | Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy. | |
| Discovering Unconventional Signals and Contextual Matching | 5 | 6 | 1 | 1 | Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication. | |
| Labor Displacement and Physical vs Digital Automation | 4 | 6 | 1 | 1 | Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles. | |
| Enduring Human Skills, Verifiability, and Generalization | 5 | 5 | 1 | 2 | Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification. | |
| Overcoming the AI Evaluation Crisis with Agents | 5 | 6 | 2 | 2 | Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination. | |
| Cultivating Taste and Designing Effective Work Assessments | 4 | 5 | 2 | 1 | Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments. | |
| Scale of Data Collection and Fixed-Cost Evals | 6 | 5 | 2 | 3 | Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments. | |
| Human Feedback Limits and Complex Agentic Benchmarks | 7 | 5 | 2 | 3 | Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows. | |
| Incentivizing Frontier Experts and Knowledge Worker Compensation | 6 | 5 | 1 | 2 | Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management. | |
| AI Agents as Managers and Personal Assistants | 5 | 4 | 1 | 1 | Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows. | |
| Reinforcement Fine-Tuning and Enterprise Agent Customization | 5 | 7 | 1 | 1 | Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows. | |
| Mercor's Strategic Priorities: Network Effects and Flywheels | 4 | 6 | 1 | 1 | Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel. | |
| Hiring vs Firing Dynamics and Evaluating Proxies | 7 | 4 | 1 | 2 | Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities. | |
| LLM Hiring Capabilities and Scaled Thiel Heuristics | 4 | 5 | 1 | 1 | Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion. |