Apr 25, 2023 · 1h 0m · no-priors

No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta

Noam Brown · 44m spoken Elad Gil · 6m spoken Sarah Guo · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In the debut episode of No Priors, Meta research scientist Noam Brown discusses his breakthroughs in game theory AI—from solving poker to mastering human negotiation in Diplomacy—and explains why scaling inference-time compute is the key to artificial general intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 18% of the talking time here. How this is scored →

The hosts as informed peer 5.1 Guest teaching 4.6 Guest disagreement 1.4 The hosts pushing back 1.7
05100:0015:0030:0045:001:00:000:02–4:18 · The hosts as informed peer 4/10 Noam Brown's Journey into Algorithmic Game Theory Elad opens with a detailed question contextualizing prompt-based LLMs versus game-theoretic agents. Noam amicably shares his background moving from finance and economics into algorithmic game theory. The conversation is entirely cordial and biographical.4:18–6:48 · The hosts as informed peer 5/10 AlphaGo, Deep Blue, and the Limits of Scaling Search Sarah prompts Noam to explain search spaces in Go versus chess, sharing her own background as a former Go player. Noam outlines the difference in handcrafted evaluation functions versus pattern matching. The exchange is cooperative and educational.6:48–11:06 · The hosts as informed peer 4/10 Tackling Diplomacy and Passing Undetected Among Humans Elad asks what guided the selection of Diplomacy as a benchmark and what surprising outcomes emerged. Noam explains how Cicero operated undetected across 40 games by taking advantage of human assumptions in natural language. The tone is informative and collaborative.11:06–18:01 · The hosts as informed peer 5/10 Cicero's Empathetic Dialogue and the Death of the Turing Test Elad asks whether data bottlenecks will limit AI agent development. Noam gently pushes back on the conventional data-scarcity worry, pointing out that inference-time compute and reasoning architectures are the true bottlenecks. The dialogue is intellectually rich without hostility.18:01–24:43 · The hosts as informed peer 5/10 Training Cicero with Human Data and Cooperative Self-Play Sarah asks about training Cicero using self-play and limited human data from webdiplomacy.net. Noam delivers an in-depth explanation of why pure self-play creates alien conventions and how human behavioral priors are necessary in mixed-motive games. Sarah synthesizes the implications for real-world agent interactions.24:43–30:02 · The hosts as informed peer 6/10 The Next Frontier: Moving Beyond Recreational Games to General AI Elad actively challenges Noam's assertion that generality is humans' main remaining advantage, arguing that human domain specialization means the AI generality bar is artificially inflated. Noam clarifies that the key differentiator is sample efficiency rather than broad multi-task performance. This marks the most intellectually sharp exchange of the episode.30:02–33:26 · The hosts as informed peer 5/10 Evaluating AI in Financial Markets and Real-World Negotiations Elad inquires about AI applications in financial markets and programmatic smart contracts, linking it back to inference compute. Noam candidly notes his limitations regarding proprietary finance setups while explaining the challenges of non-stationary environments. Sarah then explores automated commercial negotiations.33:26–39:20 · The hosts as informed peer 5/10 General Reasoning, Theorem Proving, and Inference-Time Thinking Noam highlights the necessity of inference-time search over purely scaled neural networks, citing the 100,000x compute equivalence in AlphaGo. Sarah adds context regarding the Riemann hypothesis and code generation scope. The dynamic is respectful with mutual technical understanding.39:20–43:49 · The hosts as informed peer 7/10 Architectural Generality vs Specialized Modules and Research Risk-Taking Elad demonstrates substantial domain knowledge by citing neurobiology, brain ablation cases, and evolutionary local maxima to question whether general architectures should be preferred over modular sub-models. Noam acknowledges the validity of modular systems while clarifying his focus on general outcome capabilities. Sarah closes the segment highlighting high-risk research strategy.43:49–49:17 · The hosts as informed peer 5/10 The Mechanics of Diplomacy and the Evolution of Centaur Play Noam provides a rules breakdown of Diplomacy, explaining its unique private negotiation dynamics. Sarah and Elad discuss the durability of human-AI centaur play, with Noam playfully noting that in pure chess AI has already made the human partner redundant.49:17–1:00:28 · The hosts as informed peer 5/10 Cracking Poker AI: Abstraction, Real-Time Search, and Nash Equilibrium Noam details his PhD breakthrough with Libratus and Pluribus, explaining hand abstraction, real-time depth-limited search, overbets, and Nash equilibrium exploitation. Sarah notes how this upended traditional poker psychology. The hosts provide engaged, admiring wrap-up commentary.0:02–4:18 · Guest teaching 3/10 Noam Brown's Journey into Algorithmic Game Theory Elad opens with a detailed question contextualizing prompt-based LLMs versus game-theoretic agents. Noam amicably shares his background moving from finance and economics into algorithmic game theory. The conversation is entirely cordial and biographical.4:18–6:48 · Guest teaching 4/10 AlphaGo, Deep Blue, and the Limits of Scaling Search Sarah prompts Noam to explain search spaces in Go versus chess, sharing her own background as a former Go player. Noam outlines the difference in handcrafted evaluation functions versus pattern matching. The exchange is cooperative and educational.6:48–11:06 · Guest teaching 5/10 Tackling Diplomacy and Passing Undetected Among Humans Elad asks what guided the selection of Diplomacy as a benchmark and what surprising outcomes emerged. Noam explains how Cicero operated undetected across 40 games by taking advantage of human assumptions in natural language. The tone is informative and collaborative.11:06–18:01 · Guest teaching 6/10 Cicero's Empathetic Dialogue and the Death of the Turing Test Elad asks whether data bottlenecks will limit AI agent development. Noam gently pushes back on the conventional data-scarcity worry, pointing out that inference-time compute and reasoning architectures are the true bottlenecks. The dialogue is intellectually rich without hostility.18:01–24:43 · Guest teaching 6/10 Training Cicero with Human Data and Cooperative Self-Play Sarah asks about training Cicero using self-play and limited human data from webdiplomacy.net. Noam delivers an in-depth explanation of why pure self-play creates alien conventions and how human behavioral priors are necessary in mixed-motive games. Sarah synthesizes the implications for real-world agent interactions.24:43–30:02 · Guest teaching 5/10 The Next Frontier: Moving Beyond Recreational Games to General AI Elad actively challenges Noam's assertion that generality is humans' main remaining advantage, arguing that human domain specialization means the AI generality bar is artificially inflated. Noam clarifies that the key differentiator is sample efficiency rather than broad multi-task performance. This marks the most intellectually sharp exchange of the episode.30:02–33:26 · Guest teaching 4/10 Evaluating AI in Financial Markets and Real-World Negotiations Elad inquires about AI applications in financial markets and programmatic smart contracts, linking it back to inference compute. Noam candidly notes his limitations regarding proprietary finance setups while explaining the challenges of non-stationary environments. Sarah then explores automated commercial negotiations.33:26–39:20 · Guest teaching 5/10 General Reasoning, Theorem Proving, and Inference-Time Thinking Noam highlights the necessity of inference-time search over purely scaled neural networks, citing the 100,000x compute equivalence in AlphaGo. Sarah adds context regarding the Riemann hypothesis and code generation scope. The dynamic is respectful with mutual technical understanding.39:20–43:49 · Guest teaching 3/10 Architectural Generality vs Specialized Modules and Research Risk-Taking Elad demonstrates substantial domain knowledge by citing neurobiology, brain ablation cases, and evolutionary local maxima to question whether general architectures should be preferred over modular sub-models. Noam acknowledges the validity of modular systems while clarifying his focus on general outcome capabilities. Sarah closes the segment highlighting high-risk research strategy.43:49–49:17 · Guest teaching 4/10 The Mechanics of Diplomacy and the Evolution of Centaur Play Noam provides a rules breakdown of Diplomacy, explaining its unique private negotiation dynamics. Sarah and Elad discuss the durability of human-AI centaur play, with Noam playfully noting that in pure chess AI has already made the human partner redundant.49:17–1:00:28 · Guest teaching 6/10 Cracking Poker AI: Abstraction, Real-Time Search, and Nash Equilibrium Noam details his PhD breakthrough with Libratus and Pluribus, explaining hand abstraction, real-time depth-limited search, overbets, and Nash equilibrium exploitation. Sarah notes how this upended traditional poker psychology. The hosts provide engaged, admiring wrap-up commentary.0:02–4:18 · Guest disagreement 1/10 Noam Brown's Journey into Algorithmic Game Theory Elad opens with a detailed question contextualizing prompt-based LLMs versus game-theoretic agents. Noam amicably shares his background moving from finance and economics into algorithmic game theory. The conversation is entirely cordial and biographical.4:18–6:48 · Guest disagreement 1/10 AlphaGo, Deep Blue, and the Limits of Scaling Search Sarah prompts Noam to explain search spaces in Go versus chess, sharing her own background as a former Go player. Noam outlines the difference in handcrafted evaluation functions versus pattern matching. The exchange is cooperative and educational.6:48–11:06 · Guest disagreement 1/10 Tackling Diplomacy and Passing Undetected Among Humans Elad asks what guided the selection of Diplomacy as a benchmark and what surprising outcomes emerged. Noam explains how Cicero operated undetected across 40 games by taking advantage of human assumptions in natural language. The tone is informative and collaborative.11:06–18:01 · Guest disagreement 2/10 Cicero's Empathetic Dialogue and the Death of the Turing Test Elad asks whether data bottlenecks will limit AI agent development. Noam gently pushes back on the conventional data-scarcity worry, pointing out that inference-time compute and reasoning architectures are the true bottlenecks. The dialogue is intellectually rich without hostility.18:01–24:43 · Guest disagreement 1/10 Training Cicero with Human Data and Cooperative Self-Play Sarah asks about training Cicero using self-play and limited human data from webdiplomacy.net. Noam delivers an in-depth explanation of why pure self-play creates alien conventions and how human behavioral priors are necessary in mixed-motive games. Sarah synthesizes the implications for real-world agent interactions.24:43–30:02 · Guest disagreement 3/10 The Next Frontier: Moving Beyond Recreational Games to General AI Elad actively challenges Noam's assertion that generality is humans' main remaining advantage, arguing that human domain specialization means the AI generality bar is artificially inflated. Noam clarifies that the key differentiator is sample efficiency rather than broad multi-task performance. This marks the most intellectually sharp exchange of the episode.30:02–33:26 · Guest disagreement 1/10 Evaluating AI in Financial Markets and Real-World Negotiations Elad inquires about AI applications in financial markets and programmatic smart contracts, linking it back to inference compute. Noam candidly notes his limitations regarding proprietary finance setups while explaining the challenges of non-stationary environments. Sarah then explores automated commercial negotiations.33:26–39:20 · Guest disagreement 1/10 General Reasoning, Theorem Proving, and Inference-Time Thinking Noam highlights the necessity of inference-time search over purely scaled neural networks, citing the 100,000x compute equivalence in AlphaGo. Sarah adds context regarding the Riemann hypothesis and code generation scope. The dynamic is respectful with mutual technical understanding.39:20–43:49 · Guest disagreement 2/10 Architectural Generality vs Specialized Modules and Research Risk-Taking Elad demonstrates substantial domain knowledge by citing neurobiology, brain ablation cases, and evolutionary local maxima to question whether general architectures should be preferred over modular sub-models. Noam acknowledges the validity of modular systems while clarifying his focus on general outcome capabilities. Sarah closes the segment highlighting high-risk research strategy.43:49–49:17 · Guest disagreement 2/10 The Mechanics of Diplomacy and the Evolution of Centaur Play Noam provides a rules breakdown of Diplomacy, explaining its unique private negotiation dynamics. Sarah and Elad discuss the durability of human-AI centaur play, with Noam playfully noting that in pure chess AI has already made the human partner redundant.49:17–1:00:28 · Guest disagreement 1/10 Cracking Poker AI: Abstraction, Real-Time Search, and Nash Equilibrium Noam details his PhD breakthrough with Libratus and Pluribus, explaining hand abstraction, real-time depth-limited search, overbets, and Nash equilibrium exploitation. Sarah notes how this upended traditional poker psychology. The hosts provide engaged, admiring wrap-up commentary.0:02–4:18 · The hosts pushing back 1/10 Noam Brown's Journey into Algorithmic Game Theory Elad opens with a detailed question contextualizing prompt-based LLMs versus game-theoretic agents. Noam amicably shares his background moving from finance and economics into algorithmic game theory. The conversation is entirely cordial and biographical.4:18–6:48 · The hosts pushing back 1/10 AlphaGo, Deep Blue, and the Limits of Scaling Search Sarah prompts Noam to explain search spaces in Go versus chess, sharing her own background as a former Go player. Noam outlines the difference in handcrafted evaluation functions versus pattern matching. The exchange is cooperative and educational.6:48–11:06 · The hosts pushing back 1/10 Tackling Diplomacy and Passing Undetected Among Humans Elad asks what guided the selection of Diplomacy as a benchmark and what surprising outcomes emerged. Noam explains how Cicero operated undetected across 40 games by taking advantage of human assumptions in natural language. The tone is informative and collaborative.11:06–18:01 · The hosts pushing back 2/10 Cicero's Empathetic Dialogue and the Death of the Turing Test Elad asks whether data bottlenecks will limit AI agent development. Noam gently pushes back on the conventional data-scarcity worry, pointing out that inference-time compute and reasoning architectures are the true bottlenecks. The dialogue is intellectually rich without hostility.18:01–24:43 · The hosts pushing back 1/10 Training Cicero with Human Data and Cooperative Self-Play Sarah asks about training Cicero using self-play and limited human data from webdiplomacy.net. Noam delivers an in-depth explanation of why pure self-play creates alien conventions and how human behavioral priors are necessary in mixed-motive games. Sarah synthesizes the implications for real-world agent interactions.24:43–30:02 · The hosts pushing back 4/10 The Next Frontier: Moving Beyond Recreational Games to General AI Elad actively challenges Noam's assertion that generality is humans' main remaining advantage, arguing that human domain specialization means the AI generality bar is artificially inflated. Noam clarifies that the key differentiator is sample efficiency rather than broad multi-task performance. This marks the most intellectually sharp exchange of the episode.30:02–33:26 · The hosts pushing back 2/10 Evaluating AI in Financial Markets and Real-World Negotiations Elad inquires about AI applications in financial markets and programmatic smart contracts, linking it back to inference compute. Noam candidly notes his limitations regarding proprietary finance setups while explaining the challenges of non-stationary environments. Sarah then explores automated commercial negotiations.33:26–39:20 · The hosts pushing back 1/10 General Reasoning, Theorem Proving, and Inference-Time Thinking Noam highlights the necessity of inference-time search over purely scaled neural networks, citing the 100,000x compute equivalence in AlphaGo. Sarah adds context regarding the Riemann hypothesis and code generation scope. The dynamic is respectful with mutual technical understanding.39:20–43:49 · The hosts pushing back 3/10 Architectural Generality vs Specialized Modules and Research Risk-Taking Elad demonstrates substantial domain knowledge by citing neurobiology, brain ablation cases, and evolutionary local maxima to question whether general architectures should be preferred over modular sub-models. Noam acknowledges the validity of modular systems while clarifying his focus on general outcome capabilities. Sarah closes the segment highlighting high-risk research strategy.43:49–49:17 · The hosts pushing back 2/10 The Mechanics of Diplomacy and the Evolution of Centaur Play Noam provides a rules breakdown of Diplomacy, explaining its unique private negotiation dynamics. Sarah and Elad discuss the durability of human-AI centaur play, with Noam playfully noting that in pure chess AI has already made the human partner redundant.49:17–1:00:28 · The hosts pushing back 1/10 Cracking Poker AI: Abstraction, Real-Time Search, and Nash Equilibrium Noam details his PhD breakthrough with Libratus and Pluribus, explaining hand abstraction, real-time depth-limited search, overbets, and Nash equilibrium exploitation. Sarah notes how this upended traditional poker psychology. The hosts provide engaged, admiring wrap-up commentary.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 27.6% · guest 72.4%0:00 · the hosts 27.6% · guest 72.4%3:00 · the hosts 21.8% · guest 78.2%3:00 · the hosts 21.8% · guest 78.2%6:00 · the hosts 15.2% · guest 84.8%6:00 · the hosts 15.2% · guest 84.8%9:00 · the hosts 10.1% · guest 89.9%9:00 · the hosts 10.1% · guest 89.9%12:00 · the hosts 7.4% · guest 92.6%12:00 · the hosts 7.4% · guest 92.6%15:00 · the hosts 24.8% · guest 75.2%15:00 · the hosts 24.8% · guest 75.2%18:00 · the hosts 9.4% · guest 90.6%18:00 · the hosts 9.4% · guest 90.6%21:00 · the hosts 13.5% · guest 86.5%21:00 · the hosts 13.5% · guest 86.5%24:00 · the hosts 8.5% · guest 91.5%24:00 · the hosts 8.5% · guest 91.5%27:00 · the hosts 35.4% · guest 64.6%27:00 · the hosts 35.4% · guest 64.6%30:00 · the hosts 27.1% · guest 72.9%30:00 · the hosts 27.1% · guest 72.9%33:00 · the hosts 13.2% · guest 86.8%33:00 · the hosts 13.2% · guest 86.8%36:00 · the hosts 22.1% · guest 77.9%36:00 · the hosts 22.1% · guest 77.9%39:00 · the hosts 49.9% · guest 50.1%39:00 · the hosts 49.9% · guest 50.1%42:00 · the hosts 9.5% · guest 90.5%42:00 · the hosts 9.5% · guest 90.5%45:00 · the hosts 20% · guest 80%45:00 · the hosts 20% · guest 80%48:00 · the hosts 22.9% · guest 77.1%48:00 · the hosts 22.9% · guest 77.1%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 8.5% · guest 91.5%54:00 · the hosts 8.5% · guest 91.5%57:00 · the hosts 14% · guest 86%57:00 · the hosts 14% · guest 86%1:00:00 · the hosts 13% · guest 87%1:00:00 · the hosts 13% · guest 87%
Sharpest disagreement ▶ 29:32 Noam counters Elad's claim on generality by pointing to sample efficiency

When Elad argues that the standard for AI generality is set unrealistically higher than human competence, Noam counters that the fundamental gap is sample efficiency.

Hardest push from the hosts ▶ 29:07 Elad challenges the human generality premise

Elad directly pushes back against Noam's framing of human versatility, pointing out that individuals are rarely good at everything and questioning if AI benchmarks demand too much.

Biggest teaching moment ▶ 37:47 Noam calculates the 100,000x scaling penalty of omitting search

Noam breaks down the quantitative reality of planning algorithms, proving that matching MCTS performance without test-time compute would require an unrealistic 100,000x model scaling.

The host holds their own ▶ 39:20 Elad cites neurobiological modularity to question monolithic general architectures

Elad demonstrates strong independent knowledge by referencing brain modularity, ablation studies, and evolutionary local maxima to interrogate the prevailing pursuit of single general architectures.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Noam Brown's Journey into Algorithmic Game Theory 4311 Elad opens with a detailed question contextualizing prompt-based LLMs versus game-theoretic agents. Noam amicably shares his background moving from finance and economics into algorithmic game theory. The conversation is entirely cordial and biographical.
AlphaGo, Deep Blue, and the Limits of Scaling Search 5411 Sarah prompts Noam to explain search spaces in Go versus chess, sharing her own background as a former Go player. Noam outlines the difference in handcrafted evaluation functions versus pattern matching. The exchange is cooperative and educational.
Tackling Diplomacy and Passing Undetected Among Humans 4511 Elad asks what guided the selection of Diplomacy as a benchmark and what surprising outcomes emerged. Noam explains how Cicero operated undetected across 40 games by taking advantage of human assumptions in natural language. The tone is informative and collaborative.
Cicero's Empathetic Dialogue and the Death of the Turing Test 5622 Elad asks whether data bottlenecks will limit AI agent development. Noam gently pushes back on the conventional data-scarcity worry, pointing out that inference-time compute and reasoning architectures are the true bottlenecks. The dialogue is intellectually rich without hostility.
Training Cicero with Human Data and Cooperative Self-Play 5611 Sarah asks about training Cicero using self-play and limited human data from webdiplomacy.net. Noam delivers an in-depth explanation of why pure self-play creates alien conventions and how human behavioral priors are necessary in mixed-motive games. Sarah synthesizes the implications for real-world agent interactions.
The Next Frontier: Moving Beyond Recreational Games to General AI 6534 Elad actively challenges Noam's assertion that generality is humans' main remaining advantage, arguing that human domain specialization means the AI generality bar is artificially inflated. Noam clarifies that the key differentiator is sample efficiency rather than broad multi-task performance. This marks the most intellectually sharp exchange of the episode.
Evaluating AI in Financial Markets and Real-World Negotiations 5412 Elad inquires about AI applications in financial markets and programmatic smart contracts, linking it back to inference compute. Noam candidly notes his limitations regarding proprietary finance setups while explaining the challenges of non-stationary environments. Sarah then explores automated commercial negotiations.
General Reasoning, Theorem Proving, and Inference-Time Thinking 5511 Noam highlights the necessity of inference-time search over purely scaled neural networks, citing the 100,000x compute equivalence in AlphaGo. Sarah adds context regarding the Riemann hypothesis and code generation scope. The dynamic is respectful with mutual technical understanding.
Architectural Generality vs Specialized Modules and Research Risk-Taking 7323 Elad demonstrates substantial domain knowledge by citing neurobiology, brain ablation cases, and evolutionary local maxima to question whether general architectures should be preferred over modular sub-models. Noam acknowledges the validity of modular systems while clarifying his focus on general outcome capabilities. Sarah closes the segment highlighting high-risk research strategy.
The Mechanics of Diplomacy and the Evolution of Centaur Play 5422 Noam provides a rules breakdown of Diplomacy, explaining its unique private negotiation dynamics. Sarah and Elad discuss the durability of human-AI centaur play, with Noam playfully noting that in pure chess AI has already made the human partner redundant.
Cracking Poker AI: Abstraction, Real-Time Search, and Nash Equilibrium 5611 Noam details his PhD breakthrough with Libratus and Pluribus, explaining hand abstraction, real-time depth-limited search, overbets, and Nash equilibrium exploitation. Sarah notes how this upended traditional poker psychology. The hosts provide engaged, admiring wrap-up commentary.

Statements from this episode (27)

Insight
Computer science iterates faster than economics because building requires no permission
“If you come up with an idea, you have to get it passed through legislation and it's a very long process. Computer science is much more exciting in that way because you can just build something. You don't really need permission to do it.”
Noam Brown Apr 25, 2023 ▶ 1:36
Assertion Not checkable as stated
AI was widely considered a dead field when Noam Brown began studying
“The idea of AGI was really science fiction. There were some people that were, you know serious about it, but very few, the majority opinion was that AI was, if anything, it was kind of a dead field.”
Noam Brown Apr 25, 2023 ▶ 3:23
Insight
IBM's Deep Blue proved that scaling search works in AI
“We learned that scale really does work. And in that case, it wasn't scaling, you know, training and neural nets, it was scaling search.”
Noam Brown Apr 25, 2023 ▶ 5:17
Opinion
A five-person team could solve Settlers of Catan in a year
“To then, like, go to a game like Sotos of Catan, it just felt, like, too easy. Like, you could just take a team of five people, spend a year on that, and you'd have it cracked.”
Noam Brown Apr 25, 2023 ▶ 7:42
Assertion Supported
Meta's Cicero played 40 Diplomacy games without detection as a bot
“But surprisingly, we managed to go, like, the full 40 games without being detected as a bot.”
Noam Brown Apr 25, 2023 ▶ 10:07
Insight
People assume weird online text is human error before suspecting AI
“If somebody's saying something a little weird, because the bot does say weird things every once in a while, their first instinct is not going to be like, oh, I'm talking to a bot. Their first instinct is going to be like, oh, this person is like dumb or distra…”
Noam Brown Apr 25, 2023 ▶ 10:28
Opinion
The Turing test is no longer a useful measure for AI
“I think the Turing test is no longer really a useful measure the way it was intended to be. Certainly Just because we have bots that can, I wouldn't say they can pass the Turing test, but I mean, like they're getting close enough that it's no longer that usefu…”
Noam Brown Apr 25, 2023 ▶ 12:18
Opinion
Data availability is not the true bottleneck for AI scaling
“It's not clear that data really is the bottleneck on performance here. And I've talked to AI researchers about this, and I think there isn't as much of a worry about this as people might think. Probably that's because there's a lot more data that's out there t…”
Noam Brown Apr 25, 2023 ▶ 15:54
Prediction Didn’t hold up
A $500 million AI model will likely be trained by 2025
“You can probably easily 10 X that, you know, I wouldn't be surprised if there's a five hundred million dollar model that's trained in the next year or two.”
Noam Brown Apr 25, 2023 ▶ 16:29
Insight
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Noam Brown Apr 25, 2023 ▶ 17:12
Assertion Supported
Supervised learning on human games fails to produce expert players
“Like we also found in chess and go, we actually ran this experiment. If you do. Just pure supervised learning on a giant data set of human chess and go games. The bot that you get out from that is not an expert chess or go player. Even if it's like conditioned…”
Noam Brown Apr 25, 2023 ▶ 20:08
Insight
Self-play negotiation bots trained without human data invent unintelligible languages
“If you train that bot from scratch with no human data it's going to, it could learn to negotiate, but it could learn to negotiate in a language that's not English. It could learn to negotiate in some like gibberish robot language. And then when you stick it in…”
Noam Brown Apr 25, 2023 ▶ 22:15
Prediction Not checkable as stated
An AI-generated novel rivaling Harry Potter could arrive by 2028
“I don't think you can get an AI to output like the next Harry Potter just yet. That might not be that far off. Maybe it's like five years away or something. But I don't think it's happening just yet.”
Noam Brown Apr 25, 2023 ▶ 28:22
Assertion Supported
Humans require orders of magnitude less data than AI to achieve mastery
“Like how many games does it take for an AI, for a human to become a good chess player or a good diplomacy player or a good artist? The answer is orders of magnitude less than it takes for an AI.”
Noam Brown Apr 25, 2023 ▶ 29:35
Insight
Reinforcement learning struggles in trading because financial markets are non-stationary
“I think the major challenge with Using things like reinforcement learning for trading is that it's a non-stationary environment. So you can have all this historical data, but it's not a stationary system and it's gonna like the markets respond to world events,…”
Noam Brown Apr 25, 2023 ▶ 30:54
Prediction Not checkable as stated
AI models could likely beat humans today in constrained business negotiations
“I think if you were to look at constrained domains certain negotiation tasks, I think that AIs could probably do better than humans in that today. I mean, I'm trying to think of like specific examples, but things like you know, if you wanted to negotiate over …”
Noam Brown Apr 25, 2023 ▶ 32:34
Assertion Supported
AlphaGo's raw neural network performs substantially below top human players
“If you take out the planning that's being done in AlphaGo and just use the raw Policy network, the raw neural network, it's actually substantially below top human performance.”
Noam Brown Apr 25, 2023 ▶ 34:21
Assertion Supported
Monte Carlo Tree Search fails in imperfect-information games like poker
“And that planning algorithm that's used in AlphaGo, Monte Carlo Tree Search, is very domain specific. I think people don't appreciate just how domain specific it is because it works in chess, it works in Go, and these have been like the classic domains that pe…”
Noam Brown Apr 25, 2023 ▶ 34:52
Prediction Not checkable as stated
An AI model could prove the Riemann hypothesis by 2028
“You know, it doesn't seem crazy to me that you could have a model that can prove the Riemann hypothesis within the next five years. If you can solve the reasoning problem in a truly general way.”
Noam Brown Apr 25, 2023 ▶ 35:45
Prediction Not checkable as stated
Next-token prediction will not replace big-company software engineers
“Like next, next token prediction is going to, is getting you surprisingly far. But I don't think it's gonna get you all the way there to like replacing, you know engineers at big companies.”
Noam Brown Apr 25, 2023 ▶ 37:25
Assertion Not checkable as stated
Cicero is the first major game AI breakthrough involving cooperation
“What's really interesting about diplomacy, aside from just the natural language component, is that it really is the first major game AI breakthrough in a game that involves cooperation.”
Noam Brown Apr 25, 2023 ▶ 47:22
Prediction Not checkable as stated
AI will eventually make human partners marginal in centaur Diplomacy play
“Eventually I'd imagine that these systems become so strong that, like, it kind of goes the way of chess, where, like, the human's just kind of, like, adding a marginal difference at the end.”
Noam Brown Apr 25, 2023 ▶ 48:56
Assertion Contradicted
Inference-time search improved Noam Brown's poker AI performance by 100,000x
“If we were to add this search, this planning algorithm that would come up with a better strategy when it's actually in the hand, how much better could it do? And the answer was it improved the performance by about a 100,000 X. It was the equivalent of scaling …”
Noam Brown Apr 25, 2023 ▶ 53:31
Assertion Supported
Noam Brown's six-player poker bot cost under $150 to train
“We did another competition that bought one and that bought Cost under a 150 dollars to train if you were to run it on like a cloud computing service.”
Noam Brown Apr 25, 2023 ▶ 55:18
What-if
The multiplayer poker AI breakthrough was algorithms, not just scaling compute
“This wasn't just a matter of scaling compute. It really was an algorithmic breakthrough, and this kind of result would have been doable 20 years ago if people knew the approach to dig.”
Noam Brown Apr 25, 2023 ▶ 55:26
Assertion Not checkable as stated
All professional poker players now use AI bots for training
“And I should also say like the way professional poker players train now, They all use bots to assist them. It's a lot like chess where you play the game and then you have a bot analyze your play at the afterwards and see like, okay, did you make mistakes? Wher…”
Noam Brown Apr 25, 2023 ▶ 58:10
Insight
Poker is essentially high-dimensional chess over action probabilities
“I kind of describe poker as essentially high dimensional chess. It's chess. It's like chess where you have to reason about like a probability distribution over actions instead of just like discrete actions.”
Noam Brown Apr 25, 2023 ▶ 58:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.