Apr 10, 2025 · 41m · no-priors

No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody

Brendan Foody · 27m spoken Sarah Guo · 6m spoken Elad Gil · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Mercor CEO Brendan Foody joins Sarah Guo and Elad Gil on No Priors to discuss how large language models are transforming talent evaluation, the impending economic impact of white-collar automation, and the critical need for agentic evaluation benchmarks in frontier AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 30.2% of the talking time here. How this is scored →

The hosts as informed peer 5.1 Guest teaching 5.3 Guest disagreement 1.3 The hosts pushing back 1.6
05100:0015:0030:000:37–3:30 · The hosts as informed peer 4/10 Mercor's Core Mission and AI Lab Talent Demand Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers.3:30–6:58 · The hosts as informed peer 5/10 Power Law Distributions in Knowledge Work Performance Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy.6:59–9:37 · The hosts as informed peer 5/10 Discovering Unconventional Signals and Contextual Matching Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication.9:37–11:43 · The hosts as informed peer 4/10 Labor Displacement and Physical vs Digital Automation Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles.11:44–14:03 · The hosts as informed peer 5/10 Enduring Human Skills, Verifiability, and Generalization Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification.14:03–16:34 · The hosts as informed peer 5/10 Overcoming the AI Evaluation Crisis with Agents Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination.16:35–18:58 · The hosts as informed peer 4/10 Cultivating Taste and Designing Effective Work Assessments Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments.18:58–21:25 · The hosts as informed peer 6/10 Scale of Data Collection and Fixed-Cost Evals Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments.21:25–23:42 · The hosts as informed peer 7/10 Human Feedback Limits and Complex Agentic Benchmarks Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows.23:42–28:42 · The hosts as informed peer 6/10 Incentivizing Frontier Experts and Knowledge Worker Compensation Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management.28:43–30:55 · The hosts as informed peer 5/10 AI Agents as Managers and Personal Assistants Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows.30:55–33:28 · The hosts as informed peer 5/10 Reinforcement Fine-Tuning and Enterprise Agent Customization Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows.33:29–37:31 · The hosts as informed peer 4/10 Mercor's Strategic Priorities: Network Effects and Flywheels Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel.37:31–40:02 · The hosts as informed peer 7/10 Hiring vs Firing Dynamics and Evaluating Proxies Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities.40:02–41:32 · The hosts as informed peer 4/10 LLM Hiring Capabilities and Scaled Thiel Heuristics Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion.0:37–3:30 · Guest teaching 6/10 Mercor's Core Mission and AI Lab Talent Demand Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers.3:30–6:58 · Guest teaching 5/10 Power Law Distributions in Knowledge Work Performance Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy.6:59–9:37 · Guest teaching 6/10 Discovering Unconventional Signals and Contextual Matching Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication.9:37–11:43 · Guest teaching 6/10 Labor Displacement and Physical vs Digital Automation Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles.11:44–14:03 · Guest teaching 5/10 Enduring Human Skills, Verifiability, and Generalization Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification.14:03–16:34 · Guest teaching 6/10 Overcoming the AI Evaluation Crisis with Agents Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination.16:35–18:58 · Guest teaching 5/10 Cultivating Taste and Designing Effective Work Assessments Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments.18:58–21:25 · Guest teaching 5/10 Scale of Data Collection and Fixed-Cost Evals Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments.21:25–23:42 · Guest teaching 5/10 Human Feedback Limits and Complex Agentic Benchmarks Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows.23:42–28:42 · Guest teaching 5/10 Incentivizing Frontier Experts and Knowledge Worker Compensation Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management.28:43–30:55 · Guest teaching 4/10 AI Agents as Managers and Personal Assistants Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows.30:55–33:28 · Guest teaching 7/10 Reinforcement Fine-Tuning and Enterprise Agent Customization Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows.33:29–37:31 · Guest teaching 6/10 Mercor's Strategic Priorities: Network Effects and Flywheels Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel.37:31–40:02 · Guest teaching 4/10 Hiring vs Firing Dynamics and Evaluating Proxies Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities.40:02–41:32 · Guest teaching 5/10 LLM Hiring Capabilities and Scaled Thiel Heuristics Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion.0:37–3:30 · Guest disagreement 1/10 Mercor's Core Mission and AI Lab Talent Demand Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers.3:30–6:58 · Guest disagreement 2/10 Power Law Distributions in Knowledge Work Performance Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy.6:59–9:37 · Guest disagreement 1/10 Discovering Unconventional Signals and Contextual Matching Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication.9:37–11:43 · Guest disagreement 1/10 Labor Displacement and Physical vs Digital Automation Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles.11:44–14:03 · Guest disagreement 1/10 Enduring Human Skills, Verifiability, and Generalization Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification.14:03–16:34 · Guest disagreement 2/10 Overcoming the AI Evaluation Crisis with Agents Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination.16:35–18:58 · Guest disagreement 2/10 Cultivating Taste and Designing Effective Work Assessments Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments.18:58–21:25 · Guest disagreement 2/10 Scale of Data Collection and Fixed-Cost Evals Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments.21:25–23:42 · Guest disagreement 2/10 Human Feedback Limits and Complex Agentic Benchmarks Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows.23:42–28:42 · Guest disagreement 1/10 Incentivizing Frontier Experts and Knowledge Worker Compensation Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management.28:43–30:55 · Guest disagreement 1/10 AI Agents as Managers and Personal Assistants Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows.30:55–33:28 · Guest disagreement 1/10 Reinforcement Fine-Tuning and Enterprise Agent Customization Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows.33:29–37:31 · Guest disagreement 1/10 Mercor's Strategic Priorities: Network Effects and Flywheels Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel.37:31–40:02 · Guest disagreement 1/10 Hiring vs Firing Dynamics and Evaluating Proxies Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities.40:02–41:32 · Guest disagreement 1/10 LLM Hiring Capabilities and Scaled Thiel Heuristics Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion.0:37–3:30 · The hosts pushing back 1/10 Mercor's Core Mission and AI Lab Talent Demand Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers.3:30–6:58 · The hosts pushing back 2/10 Power Law Distributions in Knowledge Work Performance Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy.6:59–9:37 · The hosts pushing back 1/10 Discovering Unconventional Signals and Contextual Matching Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication.9:37–11:43 · The hosts pushing back 1/10 Labor Displacement and Physical vs Digital Automation Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles.11:44–14:03 · The hosts pushing back 2/10 Enduring Human Skills, Verifiability, and Generalization Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification.14:03–16:34 · The hosts pushing back 2/10 Overcoming the AI Evaluation Crisis with Agents Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination.16:35–18:58 · The hosts pushing back 1/10 Cultivating Taste and Designing Effective Work Assessments Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments.18:58–21:25 · The hosts pushing back 3/10 Scale of Data Collection and Fixed-Cost Evals Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments.21:25–23:42 · The hosts pushing back 3/10 Human Feedback Limits and Complex Agentic Benchmarks Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows.23:42–28:42 · The hosts pushing back 2/10 Incentivizing Frontier Experts and Knowledge Worker Compensation Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management.28:43–30:55 · The hosts pushing back 1/10 AI Agents as Managers and Personal Assistants Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows.30:55–33:28 · The hosts pushing back 1/10 Reinforcement Fine-Tuning and Enterprise Agent Customization Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows.33:29–37:31 · The hosts pushing back 1/10 Mercor's Strategic Priorities: Network Effects and Flywheels Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel.37:31–40:02 · The hosts pushing back 2/10 Hiring vs Firing Dynamics and Evaluating Proxies Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities.40:02–41:32 · The hosts pushing back 1/10 LLM Hiring Capabilities and Scaled Thiel Heuristics Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 31.9% · guest 68.1%0:00 · the hosts 31.9% · guest 68.1%3:00 · the hosts 22.8% · guest 77.2%3:00 · the hosts 22.8% · guest 77.2%6:00 · the hosts 20% · guest 80%6:00 · the hosts 20% · guest 80%9:00 · the hosts 19.6% · guest 80.4%9:00 · the hosts 19.6% · guest 80.4%12:00 · the hosts 38% · guest 62%12:00 · the hosts 38% · guest 62%15:00 · the hosts 13.5% · guest 86.5%15:00 · the hosts 13.5% · guest 86.5%18:00 · the hosts 41.1% · guest 58.9%18:00 · the hosts 41.1% · guest 58.9%21:00 · the hosts 32.4% · guest 67.6%21:00 · the hosts 32.4% · guest 67.6%24:00 · the hosts 17.6% · guest 82.4%24:00 · the hosts 17.6% · guest 82.4%27:00 · the hosts 62.5% · guest 37.5%27:00 · the hosts 62.5% · guest 37.5%30:00 · the hosts 23.5% · guest 76.5%30:00 · the hosts 23.5% · guest 76.5%33:00 · the hosts 7.3% · guest 92.7%33:00 · the hosts 7.3% · guest 92.7%36:00 · the hosts 59.1% · guest 40.9%36:00 · the hosts 59.1% · guest 40.9%39:00 · the hosts 33.5% · guest 66.5%39:00 · the hosts 33.5% · guest 66.5%
Sharpest disagreement ▶ 22:43 Pushing back on premature AGI timelines

Foody firmly counters the popular narrative that rapid SWE-bench progress means software engineers are about to be replaced, emphasizing the massive unaddressed coordination and tool-use challenges.

Hardest push from the hosts ▶ 22:07 Challenging human feedback when models surpass expert panels

Elad uses Google's Med-PaLM 2 to push back against the necessity of human evaluators, arguing that human scoring actively degrades model performance once AI surpasses median practitioners.

Biggest teaching moment ▶ 31:33 Explaining why RFT supersedes SFT in enterprise customization

Foody provides a deep technical explanation contrasting SFT data inefficiency with outcome-based Reinforcement Fine-Tuning (RFT) across enterprise customer use cases.

The host holds their own ▶ 37:31 Elad analyzing hiring versus firing dichotomies at Google and Facebook

Elad demonstrates seasoned startup expertise by contrasting Google's high hiring bar with Facebook's swift performance management during their formative growth stages.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Mercor's Core Mission and AI Lab Talent Demand 4611 Sarah and Elad ask foundational questions about Mercor's pivot from general matching to high-end RL human data. Foody educates them on the shift in AI lab demand from crowdsourced low-skill labelers to elite domain experts pushing model frontiers.
Power Law Distributions in Knowledge Work Performance 5522 Elad probes whether knowledge work follows a power law versus a standard bell curve. Foody tailors the answer across industries, noting VC is extreme power law while software engineering sits in the middle, and explains how text-based evaluation reaches superhuman accuracy.
Discovering Unconventional Signals and Contextual Matching 5611 Sarah asks about unconventional discovered features in candidate evaluations. Foody shares specific non-intuitive empirical signals, such as international candidates with Western study abroad backgrounds showing superior collaborative communication.
Labor Displacement and Physical vs Digital Automation 4611 Elad asks about the timeline and scale of white-collar labor displacement. Foody lays out a stark scenario predicting rapid displacement in digital workflows, resulting in political unrest and a shift of human labor toward physical and emotional support roles.
Enduring Human Skills, Verifiability, and Generalization 5512 Sarah and Elad drill into which skills remain resilient, pinpointing verifiable domains like code and math versus unverifiable domains like founder evaluation. Foody agrees, explaining how autograders structure unstructured verification.
Overcoming the AI Evaluation Crisis with Agents 5622 Sarah raises the AI eval crisis and benchmark saturation. Foody explains that zero-shot benchmarks are obsolete and argues that the new frontier requires multi-step agentic evaluations reflecting realistic cross-functional coordination.
Cultivating Taste and Designing Effective Work Assessments 4521 Elad asks if young kids should learn computer science today. Foody discourages pure syntax training in favor of taste, reasoning, and entrepreneurial problem-solving, advising against using weak proxies in hiring assessments.
Scale of Data Collection and Fixed-Cost Evals 6523 Sarah and Elad question the long-term sustainability of eval-creation jobs, pointing out that labelers essentially train their own automated replacements. Elad invokes the Nyquist theorem to challenge how humans can eval superhuman intelligence, while Foody breaks knowledge work into variable tasks versus one-time fixed eval investments.
Human Feedback Limits and Complex Agentic Benchmarks 7523 Elad cites Google's Med-PaLM 2 to argue that human panel scoring eventually degrades superhuman models. Foody responds that models learn to selectively filter out human errors and highlights the vast remaining gap between narrow benchmark scores and end-to-end professional workflows.
Incentivizing Frontier Experts and Knowledge Worker Compensation 6512 Sarah highlights the high opportunity cost of recruiting top software engineers and doctors for labeling. Elad offers a critique of institutional waste in government and big tech as existing disguised forms of UBI, which Foody connects to upcoming algorithmic performance management.
AI Agents as Managers and Personal Assistants 5411 Foody suggests AI agents will excel faster as performance managers than individual contributors. Sarah relates this to personal assistant bottlenecks, and Foody explains why base models need specialized grounding to handle organizational workflows.
Reinforcement Fine-Tuning and Enterprise Agent Customization 5711 Foody delivers a clear technical explanation of Reinforcement Fine-Tuning (RFT) compared to supervised fine-tuning (SFT), explaining how defining target reward outcomes dramatically increases data efficiency for enterprise workflows.
Mercor's Strategic Priorities: Network Effects and Flywheels 4611 Foody details Mercor's strategic priorities: building candidate supply through free AI tooling to solve the 50:1 marketplace rejection problem, and compounding downstream customer performance data into an enduring predictive flywheel.
Hiring vs Firing Dynamics and Evaluating Proxies 7412 Elad draws on historical Silicon Valley lore comparing Google's hire-well/cannot-fire culture against Facebook's aggressive early performance management. Sarah discusses work trials as outcome proxies, and Foody outlines cross-company data aggregation opportunities.
LLM Hiring Capabilities and Scaled Thiel Heuristics 4511 Foody reflects on scaling Peter Thiel's interview heuristics to global candidate pools using LLMs to unlock non-traditional talent, bringing the episode to a collaborative conclusion.

Statements from this episode (37)

Assertion Not checkable as stated
Foody: Top AI Labs Use Mercor to Hire Thousands of Model Trainers
“It's used by all of the top AI labs to hire thousands of people that train the next generation of models.”
Brendan Foody Apr 10, 2025 ▶ 1:03
Insight
Foody: Model Improvement via RL Is Gated Entirely by Evaluation Benchmarks
“Reinforcement learning is becoming so effective that once you create evals, the models can learn them and how to you know, improve capabilities. And so for everything that we want alums to be good at, we need evals for those things.”
Brendan Foody Apr 10, 2025 ▶ 1:14
Assertion Not checkable as stated
Foody: AI models beat human hiring managers on most talent evals
“We're already seeing on most of our evals that models are better than human hiring managers at assessing talent, and it's still like the very early innings.”
Brendan Foody Apr 10, 2025 ▶ 2:58
Prediction Not checkable as stated
Foody: Hiring decisions will eventually rely primarily on AI recommendations
“I think we'll get to a point where I'll almost be irrational to not listen to the model, right, where people trust the model's recommendation. And like, maybe for legal reasons, we'll still have the human pressing the button and making the final sign off. But …”
Brendan Foody Apr 10, 2025 ▶ 3:08
Insight
Foody: Software talent is power-law distributed, but less extreme than venture investing
“It's very industry by industry, right? Like for you and in investing, right? It's like the most power law thing imaginable. And where it's just like the top handful of companies each decade are the ones that matter such a disproportionate amount. And it's the …”
Brendan Foody Apr 10, 2025 ▶ 4:30
Assertion Not checkable as stated
Foody: AI models are superhuman at evaluating text-based interview transcripts
“Everything that you can measure with text, the models are really good at. Like if you can ask questions in an interview and read through the transcript, the models are superhuman at that across many more domains than one would think.”
Brendan Foody Apr 10, 2025 ▶ 5:21
Prediction Not checkable as stated
Foody: AI will be slower to evaluate multimodal signals like candidate passion
“I think the things where models are going to be slower is on the multimodal signals and understanding, like, How passionate is this person about what they're working on, right? Like how persuasive are they or good at sales? And those capabilities will come, bu…”
Brendan Foody Apr 10, 2025 ▶ 5:42
Prediction Not checkable as stated
Foody: High-volume hiring processes will be automated first by AI
“And so I think it'll be those higher volume processes that also get automated first.”
Brendan Foody Apr 10, 2025 ▶ 6:55
Insight
Foody: Hiring managers under-index on engineers' public online artifacts
“One of the really interesting things for engineering is that there's so much signal about a lot of the best engineers online that I don't think people properly tap into, right? It's everything ranging from their GitHub's to the personal projects on their websi…”
Brendan Foody Apr 10, 2025 ▶ 7:17
Assertion Not checkable as stated
Foody: International candidates who studied in the West communicate better
“One interesting one we've seen in the past is that people who are based internationally but study abroad in a Western country tend to, like, work much more collaboratively or communicate better with people”
Brendan Foody Apr 10, 2025 ▶ 8:29
Prediction Not checkable as stated
Foody: AI job displacement will happen quickly and spark populist movement
“I think displacement in a lot of roles is going to happen very quickly and it's going to be Very painful, ah, and a large political problem. Like, I think we're gonna have a big populist movement around this and all the displacement that's gonna happen”
Brendan Foody Apr 10, 2025 ▶ 9:59
Prediction Not checkable as stated
Foody: Physical world automation will progress much slower than digital automation
“I think that automation In the physical world is going to happen a lot slower than what's happening in the digital world just because of so many of the, like, self-reinforcing gains and a lot of, yeah, self-improvement that can happen in, in the virtual world,…”
Brendan Foody Apr 10, 2025 ▶ 11:22
Prediction Not checkable as stated
Foody: Verifiable domains like math and code will be solved quickly by AI
“For things like math or soon code that are verifiable, they will get solved very quickly.”
Brendan Foody Apr 10, 2025 ▶ 12:24
Prediction Not checkable as stated
Foody: AI labs will specialize by industry rather than one lab dominating
“I think it's going to be hard for one lab to do everything there and there's going to be, you know, more specialization as we progress further and further and marginal gains in each industry become more challenging.”
Brendan Foody Apr 10, 2025 ▶ 13:18
Prediction Not checkable as stated
Foody: AI reasoning from code and math will generalize via transfer learning
“Yeah, I generally believe in it but to a certain extent, like, you still need a reasonable amount of data for the new domain and to kickstart it but there's gonna be a lot of transfer learning.”
Brendan Foody Apr 10, 2025 ▶ 13:41
Insight
Foody: Agent eval creation is biggest barrier to automating knowledge work
“And so I think we're going to see an immense amount of eval creation for, like, agents and that is the largest barrier to Automating most knowledge work in the economy.”
Brendan Foody Apr 10, 2025 ▶ 15:11
Prediction Not checkable as stated
Foody: Software engineer evals will take years to build
“All the things that go into making a good software engineer, that's gonna be really hard to do. Like, I think it's going to be a years long build out for even some of the verifiable domains. Cause there's so much that goes into a good software engineer of like…”
Brendan Foody Apr 10, 2025 ▶ 16:14
Prediction Not checkable as stated
Foody: Pure coding skills will not be primary value driver in five years
“I am skeptical that, like, the really valuable thing is just people who can code in five years. I think it's much more likely, like, the people that have these contrarian ideas around what's missing in markets and have the taste of What like features and nuanc…”
Brendan Foody Apr 10, 2025 ▶ 17:13
Prediction Not checkable as stated
Foody: Eval creation could become the most common knowledge job globally
“It Would not surprise me if that becomes the most common knowledge work job in the world.”
Brendan Foody Apr 10, 2025 ▶ 19:43
Insight
Foody: Superintelligence cannot be recognized without comprehensive human evals
“You don't even know that you have super intelligence without having evals for everything. Cause it's like, you sort of need to understand what is the human baseline and like, what is good? It's like grounded in this like understanding of human behavior.”
Brendan Foody Apr 10, 2025 ▶ 20:09
Insight
Foody: Knowledge work will shift from repetitive tasks to fixed-cost eval building
“It does seem structurally more efficient for work to trend away from the variable cost of like doing it repeatedly towards this fixed cost of how do we build out the evals and the processes for models to do this themselves.”
Brendan Foody Apr 10, 2025 ▶ 21:13
Prediction Not checkable as stated
Foody: AI models will propose evaluation criteria for domain experts to validate
“I think that they'll play a role in creating their own evals, where they, like, yeah, where they might come up with certain criteria for what a good response looks like, and humans validate that criteria. However, I think you often need to ground this in, like…”
Brendan Foody Apr 10, 2025 ▶ 21:49
Assertion Supported
Gil: Google's Med-PaLM 2 outperformed individual physicians in panel evaluations
“Where MedPalm II, where the output of the model was better than the average physician. It was basically like a health model, the Google Belt, and they would use physician panels to rate Outputs of the model versus individual physicians, and the model did bette…”
Elad Gil Apr 10, 2025 ▶ 22:07
Prediction Not checkable as stated
Foody: AI models will identify and ignore mistakes in human evaluator data
“Well, I think the models will be able to delineate between the valuable human knowledge and the human knowledge that's not valuable. And that maybe you have doctors that create like a bunch of evals for this particular task and the model realizes like, wow, li…”
Brendan Foody Apr 10, 2025 ▶ 22:43
Prediction Not checkable as stated
Foody: Top AI Evaluators Will See Power-Law Compensation Growth
“I think that it'll definitely become more power law over time, which means that, like, the best people are going to, of course, make an incredible amount of money.”
Brendan Foody Apr 10, 2025 ▶ 24:16
Opinion
Gil: Government, academia, and big tech have effectively functioned as UBI
“The degree to which we effectively had different forms of UBI or universal basic income in different sectors of the economy. Government is a clear example where there's enormous waste, fraud, grift, et cetera, happening. Parts of academia, if you just look at …”
Elad Gil Apr 10, 2025 ▶ 27:10
Prediction Not checkable as stated
Foody: Better employee productivity analytics will trigger more corporate layoffs
“I think that as we have better analytics around the value of employees, it seems intuitive that these companies will become you know, start doing more layoffs, more cuts, et cetera.”
Brendan Foody Apr 10, 2025 ▶ 28:01
Prediction Not checkable as stated
Foody: AI may soon be better as a manager than an individual contributor
“Also, one thing that I think is very interesting is that a lot of people are in the mindset of AI being really good as an independent contributor when actually it may soon become much better at being a manager, right? And like taking a large problem, breaking …”
Brendan Foody Apr 10, 2025 ▶ 29:15
Assertion Not checkable as stated
Foody: AI models ace math tests but still fail at basic assistant tasks
“We have these models that are like incredibly good at math, right? Like you give them a test and they can ace the test, but they still can't do like basic personal assistant work, right?”
Brendan Foody Apr 10, 2025 ▶ 30:19
Opinion
Foody: Reinforcement fine-tuning makes application-layer AI customization viable
“The reason I'm so optimistic about it taking off is that it's, like, profoundly data efficient, right? And it finally makes sense to customize models at the application layer.”
Brendan Foody Apr 10, 2025 ▶ 32:40
Assertion Not checkable as stated
Foody: Labor marketplaces average a 50-to-1 supply-to-demand ratio
“The average labor marketplace has a 50 to one ratio of supply side relative to demand side, which means the average person that applies talks to their friend who also applied and neither of them got jobs.”
Brendan Foody Apr 10, 2025 ▶ 34:02
Opinion
Foody: Global Unified Labor Market Is World's Largest Economic Opportunity
“When you're able to solve this matching problem at the cost of software, It makes way for a global unified labor market that every candidate applies to and every company hires from. And I believe that that's not only the largest economic opportunity in the wor…”
Brendan Foody Apr 10, 2025 ▶ 35:44
Prediction Not checkable as stated
Foody: Future Labor Markets Will Coordinate Humans and AI Agents
“I think so, because customers ultimately come with, like, a problem to be solved, right? And ideally, it's some coordination of how those two fit together.”
Brendan Foody Apr 10, 2025 ▶ 36:13
Insight
Gil: Early-stage companies rarely master both hiring and firing well
“I find that almost every great company either hires well, like what you're talking about, or fires well, which is sort of your phase two. But I think often they do that, one of those things really well early. For some reason, most people don't seem to get both…”
Elad Gil Apr 10, 2025 ▶ 37:32
Assertion Not checkable as stated
Gil: Early Google hired well but fired poorly, unlike early Facebook
“Google was a good example of a organization that would always hire well, but couldn't fire well. It took them a really long time to clean people out. Years, like literally years. Facebook, on the other hand, was kind of known for a more mixed early talent pool…”
Elad Gil Apr 10, 2025 ▶ 37:49
Insight
Foody: LLMs unlock hiring automation that LinkedIn could never achieve
“I think that LinkedIn centralizes And aggregates the very first layer of the application process of, like, what are the things that this person has done and, like, who are they connected to? The challenge historically has been that the rest of the process to f…”
Brendan Foody Apr 10, 2025 ▶ 40:03
Prediction Not checkable as stated
Foody: AI Will Scale Peter Thiel's Interview Heuristics to Everyone Globally
“Imagine if you could have Peter Thiel as a heuristic interview everyone in the world when they're 18, right? And, like, and maybe he could go through and, like, meticulously spend time determining, like, you know, who is actually going to be good at what job. …”
Brendan Foody Apr 10, 2025 ▶ 41:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.