Apr 26, 2025 · 22m · tbpn

Why Humor Is the True Test of AI Intelligence | Will Brown on TBPN

Will Brown · 14m spoken John Coogan · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview, Morgan Stanley AI researcher Will Brown explores the future of artificial intelligence, examining the shift toward agentic reinforcement learning, inference systems optimization, and the realities of deploying AI within regulated enterprise environments.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 17.6% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 5.2 Guest disagreement 1.7 The hosts pushing back 1.5
05100:0010:0020:002:36–5:07 · The hosts as informed peer 4/10 Humor as the Ultimate Benchmark for Model Intelligence John opens with a thoughtful framing comparing math competition performance at 1.5B parameters versus joke generation at scale. Will educates the hosts on the nuances of transformer attention sparsity required for humor versus models gaming reward functions with profanity.5:07–8:24 · The hosts as informed peer 6/10 Scaling Limits, Inference Economics, and Agentic Reinforcement Learning John demonstrates solid knowledge of current AI discourse, citing Tyler Cowen's O3 framing and comparing ASIC inference economics to Bitcoin hashing. Will explains the economic plateau of giant pre-trained models and explains why multi-turn agentic RL is the primary capability unlock.8:24–12:13 · The hosts as informed peer 5/10 Defining 10-Minute AGI and Practical Program Synthesis John probes into whether future paradigms require classical symbol manipulation or program synthesis. Will reframes program synthesis as already being accomplished via tool-calling architectures and outlines his practical 10-minute human task definition of AGI.12:14–15:21 · The hosts as informed peer 4/10 Enterprise Tool Pilots, Churn, and Coding Assistant Competition Jordy and John bring up enterprise churn patterns seen in large corporations like Johnson & Johnson. Will shares insider perspective on the enterprise software sales lifecycle, explaining why pilot churn is standard and highlighting Windsurf's deliberate enterprise strategy over Cursor.15:21–19:52 · The hosts as informed peer 6/10 DeepSeek's Engineering Breakthroughs and Inference Optimization John cites Sam Altman recruiting high-frequency trading engineers and questions if DeepSeek's innovations like FP8 and MoE blocking can easily be ported to models like Llama. Will clarifies that DeepSeek's advantage comes from 10 to 15 compound optimizations stacking together rather than a single portable trick.19:54–22:04 · The hosts as informed peer 4/10 Geopolitics, Supply Chains, and Data Center Infrastructure The hosts touch on data center power bottlenecks and question how Morgan Stanley fosters an ML research culture. Will details how his research group operates akin to a finance Bell Labs and was working with OpenAI prior to ChatGPT's release.2:36–5:07 · Guest teaching 5/10 Humor as the Ultimate Benchmark for Model Intelligence John opens with a thoughtful framing comparing math competition performance at 1.5B parameters versus joke generation at scale. Will educates the hosts on the nuances of transformer attention sparsity required for humor versus models gaming reward functions with profanity.5:07–8:24 · Guest teaching 5/10 Scaling Limits, Inference Economics, and Agentic Reinforcement Learning John demonstrates solid knowledge of current AI discourse, citing Tyler Cowen's O3 framing and comparing ASIC inference economics to Bitcoin hashing. Will explains the economic plateau of giant pre-trained models and explains why multi-turn agentic RL is the primary capability unlock.8:24–12:13 · Guest teaching 6/10 Defining 10-Minute AGI and Practical Program Synthesis John probes into whether future paradigms require classical symbol manipulation or program synthesis. Will reframes program synthesis as already being accomplished via tool-calling architectures and outlines his practical 10-minute human task definition of AGI.12:14–15:21 · Guest teaching 5/10 Enterprise Tool Pilots, Churn, and Coding Assistant Competition Jordy and John bring up enterprise churn patterns seen in large corporations like Johnson & Johnson. Will shares insider perspective on the enterprise software sales lifecycle, explaining why pilot churn is standard and highlighting Windsurf's deliberate enterprise strategy over Cursor.15:21–19:52 · Guest teaching 6/10 DeepSeek's Engineering Breakthroughs and Inference Optimization John cites Sam Altman recruiting high-frequency trading engineers and questions if DeepSeek's innovations like FP8 and MoE blocking can easily be ported to models like Llama. Will clarifies that DeepSeek's advantage comes from 10 to 15 compound optimizations stacking together rather than a single portable trick.19:54–22:04 · Guest teaching 4/10 Geopolitics, Supply Chains, and Data Center Infrastructure The hosts touch on data center power bottlenecks and question how Morgan Stanley fosters an ML research culture. Will details how his research group operates akin to a finance Bell Labs and was working with OpenAI prior to ChatGPT's release.2:36–5:07 · Guest disagreement 2/10 Humor as the Ultimate Benchmark for Model Intelligence John opens with a thoughtful framing comparing math competition performance at 1.5B parameters versus joke generation at scale. Will educates the hosts on the nuances of transformer attention sparsity required for humor versus models gaming reward functions with profanity.5:07–8:24 · Guest disagreement 2/10 Scaling Limits, Inference Economics, and Agentic Reinforcement Learning John demonstrates solid knowledge of current AI discourse, citing Tyler Cowen's O3 framing and comparing ASIC inference economics to Bitcoin hashing. Will explains the economic plateau of giant pre-trained models and explains why multi-turn agentic RL is the primary capability unlock.8:24–12:13 · Guest disagreement 2/10 Defining 10-Minute AGI and Practical Program Synthesis John probes into whether future paradigms require classical symbol manipulation or program synthesis. Will reframes program synthesis as already being accomplished via tool-calling architectures and outlines his practical 10-minute human task definition of AGI.12:14–15:21 · Guest disagreement 1/10 Enterprise Tool Pilots, Churn, and Coding Assistant Competition Jordy and John bring up enterprise churn patterns seen in large corporations like Johnson & Johnson. Will shares insider perspective on the enterprise software sales lifecycle, explaining why pilot churn is standard and highlighting Windsurf's deliberate enterprise strategy over Cursor.15:21–19:52 · Guest disagreement 2/10 DeepSeek's Engineering Breakthroughs and Inference Optimization John cites Sam Altman recruiting high-frequency trading engineers and questions if DeepSeek's innovations like FP8 and MoE blocking can easily be ported to models like Llama. Will clarifies that DeepSeek's advantage comes from 10 to 15 compound optimizations stacking together rather than a single portable trick.19:54–22:04 · Guest disagreement 1/10 Geopolitics, Supply Chains, and Data Center Infrastructure The hosts touch on data center power bottlenecks and question how Morgan Stanley fosters an ML research culture. Will details how his research group operates akin to a finance Bell Labs and was working with OpenAI prior to ChatGPT's release.2:36–5:07 · The hosts pushing back 1/10 Humor as the Ultimate Benchmark for Model Intelligence John opens with a thoughtful framing comparing math competition performance at 1.5B parameters versus joke generation at scale. Will educates the hosts on the nuances of transformer attention sparsity required for humor versus models gaming reward functions with profanity.5:07–8:24 · The hosts pushing back 2/10 Scaling Limits, Inference Economics, and Agentic Reinforcement Learning John demonstrates solid knowledge of current AI discourse, citing Tyler Cowen's O3 framing and comparing ASIC inference economics to Bitcoin hashing. Will explains the economic plateau of giant pre-trained models and explains why multi-turn agentic RL is the primary capability unlock.8:24–12:13 · The hosts pushing back 1/10 Defining 10-Minute AGI and Practical Program Synthesis John probes into whether future paradigms require classical symbol manipulation or program synthesis. Will reframes program synthesis as already being accomplished via tool-calling architectures and outlines his practical 10-minute human task definition of AGI.12:14–15:21 · The hosts pushing back 1/10 Enterprise Tool Pilots, Churn, and Coding Assistant Competition Jordy and John bring up enterprise churn patterns seen in large corporations like Johnson & Johnson. Will shares insider perspective on the enterprise software sales lifecycle, explaining why pilot churn is standard and highlighting Windsurf's deliberate enterprise strategy over Cursor.15:21–19:52 · The hosts pushing back 3/10 DeepSeek's Engineering Breakthroughs and Inference Optimization John cites Sam Altman recruiting high-frequency trading engineers and questions if DeepSeek's innovations like FP8 and MoE blocking can easily be ported to models like Llama. Will clarifies that DeepSeek's advantage comes from 10 to 15 compound optimizations stacking together rather than a single portable trick.19:54–22:04 · The hosts pushing back 1/10 Geopolitics, Supply Chains, and Data Center Infrastructure The hosts touch on data center power bottlenecks and question how Morgan Stanley fosters an ML research culture. Will details how his research group operates akin to a finance Bell Labs and was working with OpenAI prior to ChatGPT's release.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 24.2% · guest 75.8%0:00 · the hosts 24.2% · guest 75.8%3:00 · the hosts 16.4% · guest 83.6%3:00 · the hosts 16.4% · guest 83.6%6:00 · the hosts 20.7% · guest 79.3%6:00 · the hosts 20.7% · guest 79.3%9:00 · the hosts 11.9% · guest 88.1%9:00 · the hosts 11.9% · guest 88.1%12:00 · the hosts 2.4% · guest 97.6%12:00 · the hosts 2.4% · guest 97.6%15:00 · the hosts 26.8% · guest 73.2%15:00 · the hosts 26.8% · guest 73.2%18:00 · the hosts 17.7% · guest 82.3%18:00 · the hosts 17.7% · guest 82.3%21:00 · the hosts 22.9% · guest 77.1%21:00 · the hosts 22.9% · guest 77.1%
Sharpest disagreement ▶ 5:16 Rejecting Big Transformer Scaling in Favor of RL

Will directly dismisses pure parameter scaling, stating he is not big-transformer pilled and highlighting that giant models like Meta's upcoming releases will be economically impractical for day-to-day operations.

Hardest push from the hosts ▶ 16:42 John Inquires on Porting DeepSeek FP8 and MoE Blocking

John presses on the narrative that DeepSeek's advantages are readily replicable, asking why standard techniques like FP8 quantization and mixture-of-experts blocking haven't been swiftly ported into Llama.

Biggest teaching moment ▶ 17:01 Explaining Compound Efficiency Gains Versus Single Breakthroughs

Will educates John that DeepSeek's breakthrough cannot be reduced to one architectural tweak like FP8, explaining that their moat lies in stacking 10 to 15 separate 30-50% efficiency gains.

The host holds their own ▶ 16:42 Host Details Low-Level Quantization and MoE Routing

John demonstrates sharp technical comprehension of open-weights infrastructure by referencing Sam Altman's HFT recruitment push alongside specific mechanisms like FP8 precision and MoE blocking.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Humor as the Ultimate Benchmark for Model Intelligence 4521 John opens with a thoughtful framing comparing math competition performance at 1.5B parameters versus joke generation at scale. Will educates the hosts on the nuances of transformer attention sparsity required for humor versus models gaming reward functions with profanity.
Scaling Limits, Inference Economics, and Agentic Reinforcement Learning 6522 John demonstrates solid knowledge of current AI discourse, citing Tyler Cowen's O3 framing and comparing ASIC inference economics to Bitcoin hashing. Will explains the economic plateau of giant pre-trained models and explains why multi-turn agentic RL is the primary capability unlock.
Defining 10-Minute AGI and Practical Program Synthesis 5621 John probes into whether future paradigms require classical symbol manipulation or program synthesis. Will reframes program synthesis as already being accomplished via tool-calling architectures and outlines his practical 10-minute human task definition of AGI.
Enterprise Tool Pilots, Churn, and Coding Assistant Competition 4511 Jordy and John bring up enterprise churn patterns seen in large corporations like Johnson & Johnson. Will shares insider perspective on the enterprise software sales lifecycle, explaining why pilot churn is standard and highlighting Windsurf's deliberate enterprise strategy over Cursor.
DeepSeek's Engineering Breakthroughs and Inference Optimization 6623 John cites Sam Altman recruiting high-frequency trading engineers and questions if DeepSeek's innovations like FP8 and MoE blocking can easily be ported to models like Llama. Will clarifies that DeepSeek's advantage comes from 10 to 15 compound optimizations stacking together rather than a single portable trick.
Geopolitics, Supply Chains, and Data Center Infrastructure 4411 The hosts touch on data center power bottlenecks and question how Morgan Stanley fosters an ML research culture. Will details how his research group operates akin to a finance Bell Labs and was working with OpenAI prior to ChatGPT's release.

Statements from this episode (15)

Insight
Brown: AI alpha is buried in GitHub issues, X, and group chats, not LinkedIn
“There's a surprising lot of alpha still on X just from places where you don't like, so there's places I have not found much alpha on LinkedIn. The group chats, the anons the open source like GitHub discussions, lots of really good stuff is buried in like a Git…”
Will Brown Apr 26, 2025 ▶ 2:06
Opinion
Brown: GPT-4.5 Likely Has Trillions of Parameters Enabling Sparse Connections
“GPT, 4.5 is like ginormous model, trillions of parameters, most likely. And that like, there's more room in the model to have these like little sparse connections materialize as you go through layers of the transformer. And I just haven't seen anything like th…”
Will Brown Apr 26, 2025 ▶ 4:44
Insight
Will Brown: AI scaling faces diminishing returns on capital investment
“The quality bump over things that are much smaller is just like the, we're hitting diminishing returns on capital investment is a lot of it. Like they're taking out of the API because like the, they can sell other things with the same GPUs and make more money …”
Will Brown Apr 26, 2025 ▶ 5:55
Prediction Not checkable as stated
Will Brown: Nobody will run Meta's giant model for daily tasks
“When Meta releases this behemoth model, I don't think anyone's really gonna run behemoth. For like their day to day stuff. It's just probably not gonna be worth it.”
Will Brown Apr 26, 2025 ▶ 6:21
Insight
Will Brown: OpenAI o3 succeeds because reinforcement learning enables tool use
“The reason O-three is good is because it's trained to use tools. The way you train a model to use the right tool for the job is reinforcement learning. And they've said as much, like, deep research, reinforcement learning.”
Will Brown Apr 26, 2025 ▶ 7:33
Prediction Not checkable as stated
Will Brown: Effective agentic AI will work well on small models
“That's kind of like my bet is like, people are going to really want to train models to be agents. And I think you can get that to work well with a pretty small model.”
Will Brown Apr 26, 2025 ▶ 8:04
Insight
Brown: OpenAI's o3 qualifies as 10-minute AGI
“I'm happy to call O three, 10 minute AGI. And I think like framing AGI in terms of like length of time, it takes a human to do a task is like more reasonable than like a global framing. Like, sure. There's a bar of like drop and replace for a human that we are…”
Will Brown Apr 26, 2025 ▶ 8:34
Insight
Brown: OpenAI's o3 already achieves practical program synthesis
“Like when people say program synthesis, like we're already there, like O three is program synthesis, but the programs are like JSON and Python.”
Will Brown Apr 26, 2025 ▶ 9:37
Assertion Supported
Will Brown: Morgan Stanley deployed integrations on day GPT-4 launched
“Like, the day GPT-IV launched, Morgan Stanley had integrations, because we had been working on it, and these were, like, we had press releases for these, like, we are ready to go”
Will Brown Apr 26, 2025 ▶ 10:57
Opinion
Will Brown: Microsoft PowerPoint Copilot is not great
“The PowerPoint, like, Microsoft PowerPoint copilot is not great. It's not a thing that I have heard many people say saves them up time.”
Will Brown Apr 26, 2025 ▶ 12:01
Insight
Brown: Most Enterprise AI Pilots Are Structurally Designed to Churn
“Most of these pilots are very much intended to be churned. Like, they're not, they're very much in a, they're not being rolled out broadly. They are coming through in a kind of walled off environment for people who are, like, gonna be the beta testers.”
Will Brown Apr 26, 2025 ▶ 12:43
Opinion
Brown: Windsurf Succeeds Against Cursor by Prioritizing Enterprise Integration
“A reason that Windsurf has been successful as a Cursor competitor is they lean way harder on enterprise than Cursor has. Like, they have really designed for enterprise integration, whereas Cursor really has not.”
Will Brown Apr 26, 2025 ▶ 14:00
Assertion Not checkable as stated
Brown: DeepSeek serves all of China on 2,000 GPUs
“Like Deep Seek is serving all of China on 2000 GPUs.”
Will Brown Apr 26, 2025 ▶ 16:06
Insight
Brown: DeepSeek's efficiency comes from stacking 10 to 15 compound gains
“I mean, to me, the surprising thing was not any individual one thing. It's that each of these is maybe like a 30%, 50% gain, but they have like 10, 15 of them that all stack. And so getting all of these to stack nicely is what's hard.”
Will Brown Apr 26, 2025 ▶ 17:02
Opinion
Brown: Alibaba Qwen makes the best model suites for research
“Like they make, I think, still the best model suites for like doing research.”
Will Brown Apr 26, 2025 ▶ 17:24
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.