Jan 1, 2025 · 1h 51m · latent-space

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents

Shawn Wang · 1h 17m spoken Alessio Fanelli · 22m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In their milestone 100th episode, Latent Space co-hosts Alessio Fanelli and Swyx deliver a comprehensive 2024 retrospective analyzing the rise of AI engineering, test-time reasoning compute, the Four Wars of AI, and dramatic inference cost deflation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 99.9% of the talking time here. How this is scored →

The hosts as informed peer 6.1 Guest teaching 0.8 Guest disagreement 0.9 The hosts pushing back 0.9
05100:0020:0040:001:00:001:20:001:40:000:03–3:14 · The hosts as informed peer 6/10 Celebrating 100 Episodes and the AI Engineer Era Alessio and Swyx open their 100th episode by celebrating the validation of the AI Engineer movement and discussing its industry recognition via Gartner's hype curve.3:15–9:16 · The hosts as informed peer 7/10 Research vs Production and the Big Scaling Debate at NeurIPS Swyx recounts the NeurIPS debates on pre-training walls, contrasting test-time compute with inference-time compute optimality alongside papers from Chinchilla to Noam Brown.9:17–18:18 · The hosts as informed peer 8/10 Frontier Model Competition, Market Shifts, and Inference Economics Swyx challenges claims that open source is closing the gap with frontier reasoning models, dissecting market share shifts across OpenAI, Anthropic, and Gemini.18:18–26:53 · The hosts as informed peer 7/10 Agent Frontiers, Systems ML, and Ilya's Keynote Insights The hosts discuss agent frontier challenges, Jeff Dean's systems ML approaches with AlphaChip, and Ilya Sutskever's historical scaling predictions.26:58–32:19 · The hosts as informed peer 7/10 Long-Tail NeurIPS Discoveries: Steganography and Dataset Trends Swyx outlines long-tail research from NeurIPS, including DeepMind's steganography paper on agent collusion and the shift toward dataset tracks.32:22–38:18 · The hosts as informed peer 7/10 The Four Wars of AI: Data Rights, Synthetic Data, and Reasoning Moats The hosts analyze the data wars, synthetic reasoning distillation via STaR methods, and the durability of reasoning moats like o1.38:18–45:31 · The hosts as informed peer 7/10 The Compute War: GPU Super-Rich Clusters vs GPU-Poor Practicality Swyx introduces the GPU smiling curve, comparing 100k-cluster mega-labs with lean customer-facing wrappers like Suno and Bolt.45:31–52:37 · The hosts as informed peer 7/10 The Multimodal War: Specialist Startups vs Integrated God Models Swyx and Alessio debate specialist multi-modal startups against integrated omni-models, bantering over Flux, Midjourney, and Sora.52:37–1:05:31 · The hosts as informed peer 8/10 The LLMOS & Agent War: Sandboxes, Memory Architecture, and Protocols Alessio evaluates PyPI metrics for LangChain and CrewAI while Swyx elaborates on why memory architectures need temporal decay and sleep consolidation.1:05:32–1:11:59 · The hosts as informed peer 8/10 The Shifting Evaluation Landscape: Benchmarks and Capabilities Tiering Swyx reviews benchmark degradation from MMLU and GPQA to SWE-bench verified, setting up the framework for mature versus emerging capabilities.1:12:01–1:22:48 · The hosts as informed peer 0/10 Editing Room Interlude: Capability Breakdown and the Price Collapse Swyx delivers an editing room monologue charting the 3-orders-of-magnitude price-intelligence collapse throughout 2024.1:22:48–1:30:37 · The hosts as informed peer 0/10 2024 Chronological AI Rewind: January to June Milestones Swyx continues his monologue rewind from January to June, touching on Perplexity, Devin's launch, Suno/Udio, Llama 3, and GPT-4o.1:30:39–1:43:11 · The hosts as informed peer 8/10 2024 Chronological AI Rewind: July to December & The Rise of Agents The hosts reunite to discuss SSI's single-product mission, o1's launch velocity, Canvas vs Google Docs, and how AI will establish workplace skill floors.1:43:11–1:50:52 · The hosts as informed peer 6/10 Top Latent Space Episodes of 2024 and Reflections Alessio and Swyx wrap up with reflections on their top 2024 episodes, the lack of a real AI winter, and community growth.0:03–3:14 · Guest teaching 0/10 Celebrating 100 Episodes and the AI Engineer Era Alessio and Swyx open their 100th episode by celebrating the validation of the AI Engineer movement and discussing its industry recognition via Gartner's hype curve.3:15–9:16 · Guest teaching 1/10 Research vs Production and the Big Scaling Debate at NeurIPS Swyx recounts the NeurIPS debates on pre-training walls, contrasting test-time compute with inference-time compute optimality alongside papers from Chinchilla to Noam Brown.9:17–18:18 · Guest teaching 1/10 Frontier Model Competition, Market Shifts, and Inference Economics Swyx challenges claims that open source is closing the gap with frontier reasoning models, dissecting market share shifts across OpenAI, Anthropic, and Gemini.18:18–26:53 · Guest teaching 1/10 Agent Frontiers, Systems ML, and Ilya's Keynote Insights The hosts discuss agent frontier challenges, Jeff Dean's systems ML approaches with AlphaChip, and Ilya Sutskever's historical scaling predictions.26:58–32:19 · Guest teaching 0/10 Long-Tail NeurIPS Discoveries: Steganography and Dataset Trends Swyx outlines long-tail research from NeurIPS, including DeepMind's steganography paper on agent collusion and the shift toward dataset tracks.32:22–38:18 · Guest teaching 1/10 The Four Wars of AI: Data Rights, Synthetic Data, and Reasoning Moats The hosts analyze the data wars, synthetic reasoning distillation via STaR methods, and the durability of reasoning moats like o1.38:18–45:31 · Guest teaching 1/10 The Compute War: GPU Super-Rich Clusters vs GPU-Poor Practicality Swyx introduces the GPU smiling curve, comparing 100k-cluster mega-labs with lean customer-facing wrappers like Suno and Bolt.45:31–52:37 · Guest teaching 2/10 The Multimodal War: Specialist Startups vs Integrated God Models Swyx and Alessio debate specialist multi-modal startups against integrated omni-models, bantering over Flux, Midjourney, and Sora.52:37–1:05:31 · Guest teaching 2/10 The LLMOS & Agent War: Sandboxes, Memory Architecture, and Protocols Alessio evaluates PyPI metrics for LangChain and CrewAI while Swyx elaborates on why memory architectures need temporal decay and sleep consolidation.1:05:32–1:11:59 · Guest teaching 1/10 The Shifting Evaluation Landscape: Benchmarks and Capabilities Tiering Swyx reviews benchmark degradation from MMLU and GPQA to SWE-bench verified, setting up the framework for mature versus emerging capabilities.1:12:01–1:22:48 · Guest teaching 0/10 Editing Room Interlude: Capability Breakdown and the Price Collapse Swyx delivers an editing room monologue charting the 3-orders-of-magnitude price-intelligence collapse throughout 2024.1:22:48–1:30:37 · Guest teaching 0/10 2024 Chronological AI Rewind: January to June Milestones Swyx continues his monologue rewind from January to June, touching on Perplexity, Devin's launch, Suno/Udio, Llama 3, and GPT-4o.1:30:39–1:43:11 · Guest teaching 1/10 2024 Chronological AI Rewind: July to December & The Rise of Agents The hosts reunite to discuss SSI's single-product mission, o1's launch velocity, Canvas vs Google Docs, and how AI will establish workplace skill floors.1:43:11–1:50:52 · Guest teaching 0/10 Top Latent Space Episodes of 2024 and Reflections Alessio and Swyx wrap up with reflections on their top 2024 episodes, the lack of a real AI winter, and community growth.0:03–3:14 · Guest disagreement 0/10 Celebrating 100 Episodes and the AI Engineer Era Alessio and Swyx open their 100th episode by celebrating the validation of the AI Engineer movement and discussing its industry recognition via Gartner's hype curve.3:15–9:16 · Guest disagreement 1/10 Research vs Production and the Big Scaling Debate at NeurIPS Swyx recounts the NeurIPS debates on pre-training walls, contrasting test-time compute with inference-time compute optimality alongside papers from Chinchilla to Noam Brown.9:17–18:18 · Guest disagreement 2/10 Frontier Model Competition, Market Shifts, and Inference Economics Swyx challenges claims that open source is closing the gap with frontier reasoning models, dissecting market share shifts across OpenAI, Anthropic, and Gemini.18:18–26:53 · Guest disagreement 1/10 Agent Frontiers, Systems ML, and Ilya's Keynote Insights The hosts discuss agent frontier challenges, Jeff Dean's systems ML approaches with AlphaChip, and Ilya Sutskever's historical scaling predictions.26:58–32:19 · Guest disagreement 0/10 Long-Tail NeurIPS Discoveries: Steganography and Dataset Trends Swyx outlines long-tail research from NeurIPS, including DeepMind's steganography paper on agent collusion and the shift toward dataset tracks.32:22–38:18 · Guest disagreement 1/10 The Four Wars of AI: Data Rights, Synthetic Data, and Reasoning Moats The hosts analyze the data wars, synthetic reasoning distillation via STaR methods, and the durability of reasoning moats like o1.38:18–45:31 · Guest disagreement 1/10 The Compute War: GPU Super-Rich Clusters vs GPU-Poor Practicality Swyx introduces the GPU smiling curve, comparing 100k-cluster mega-labs with lean customer-facing wrappers like Suno and Bolt.45:31–52:37 · Guest disagreement 2/10 The Multimodal War: Specialist Startups vs Integrated God Models Swyx and Alessio debate specialist multi-modal startups against integrated omni-models, bantering over Flux, Midjourney, and Sora.52:37–1:05:31 · Guest disagreement 2/10 The LLMOS & Agent War: Sandboxes, Memory Architecture, and Protocols Alessio evaluates PyPI metrics for LangChain and CrewAI while Swyx elaborates on why memory architectures need temporal decay and sleep consolidation.1:05:32–1:11:59 · Guest disagreement 1/10 The Shifting Evaluation Landscape: Benchmarks and Capabilities Tiering Swyx reviews benchmark degradation from MMLU and GPQA to SWE-bench verified, setting up the framework for mature versus emerging capabilities.1:12:01–1:22:48 · Guest disagreement 0/10 Editing Room Interlude: Capability Breakdown and the Price Collapse Swyx delivers an editing room monologue charting the 3-orders-of-magnitude price-intelligence collapse throughout 2024.1:22:48–1:30:37 · Guest disagreement 0/10 2024 Chronological AI Rewind: January to June Milestones Swyx continues his monologue rewind from January to June, touching on Perplexity, Devin's launch, Suno/Udio, Llama 3, and GPT-4o.1:30:39–1:43:11 · Guest disagreement 1/10 2024 Chronological AI Rewind: July to December & The Rise of Agents The hosts reunite to discuss SSI's single-product mission, o1's launch velocity, Canvas vs Google Docs, and how AI will establish workplace skill floors.1:43:11–1:50:52 · Guest disagreement 0/10 Top Latent Space Episodes of 2024 and Reflections Alessio and Swyx wrap up with reflections on their top 2024 episodes, the lack of a real AI winter, and community growth.0:03–3:14 · The hosts pushing back 0/10 Celebrating 100 Episodes and the AI Engineer Era Alessio and Swyx open their 100th episode by celebrating the validation of the AI Engineer movement and discussing its industry recognition via Gartner's hype curve.3:15–9:16 · The hosts pushing back 1/10 Research vs Production and the Big Scaling Debate at NeurIPS Swyx recounts the NeurIPS debates on pre-training walls, contrasting test-time compute with inference-time compute optimality alongside papers from Chinchilla to Noam Brown.9:17–18:18 · The hosts pushing back 2/10 Frontier Model Competition, Market Shifts, and Inference Economics Swyx challenges claims that open source is closing the gap with frontier reasoning models, dissecting market share shifts across OpenAI, Anthropic, and Gemini.18:18–26:53 · The hosts pushing back 1/10 Agent Frontiers, Systems ML, and Ilya's Keynote Insights The hosts discuss agent frontier challenges, Jeff Dean's systems ML approaches with AlphaChip, and Ilya Sutskever's historical scaling predictions.26:58–32:19 · The hosts pushing back 0/10 Long-Tail NeurIPS Discoveries: Steganography and Dataset Trends Swyx outlines long-tail research from NeurIPS, including DeepMind's steganography paper on agent collusion and the shift toward dataset tracks.32:22–38:18 · The hosts pushing back 1/10 The Four Wars of AI: Data Rights, Synthetic Data, and Reasoning Moats The hosts analyze the data wars, synthetic reasoning distillation via STaR methods, and the durability of reasoning moats like o1.38:18–45:31 · The hosts pushing back 1/10 The Compute War: GPU Super-Rich Clusters vs GPU-Poor Practicality Swyx introduces the GPU smiling curve, comparing 100k-cluster mega-labs with lean customer-facing wrappers like Suno and Bolt.45:31–52:37 · The hosts pushing back 2/10 The Multimodal War: Specialist Startups vs Integrated God Models Swyx and Alessio debate specialist multi-modal startups against integrated omni-models, bantering over Flux, Midjourney, and Sora.52:37–1:05:31 · The hosts pushing back 2/10 The LLMOS & Agent War: Sandboxes, Memory Architecture, and Protocols Alessio evaluates PyPI metrics for LangChain and CrewAI while Swyx elaborates on why memory architectures need temporal decay and sleep consolidation.1:05:32–1:11:59 · The hosts pushing back 1/10 The Shifting Evaluation Landscape: Benchmarks and Capabilities Tiering Swyx reviews benchmark degradation from MMLU and GPQA to SWE-bench verified, setting up the framework for mature versus emerging capabilities.1:12:01–1:22:48 · The hosts pushing back 0/10 Editing Room Interlude: Capability Breakdown and the Price Collapse Swyx delivers an editing room monologue charting the 3-orders-of-magnitude price-intelligence collapse throughout 2024.1:22:48–1:30:37 · The hosts pushing back 0/10 2024 Chronological AI Rewind: January to June Milestones Swyx continues his monologue rewind from January to June, touching on Perplexity, Devin's launch, Suno/Udio, Llama 3, and GPT-4o.1:30:39–1:43:11 · The hosts pushing back 1/10 2024 Chronological AI Rewind: July to December & The Rise of Agents The hosts reunite to discuss SSI's single-product mission, o1's launch velocity, Canvas vs Google Docs, and how AI will establish workplace skill floors.1:43:11–1:50:52 · The hosts pushing back 0/10 Top Latent Space Episodes of 2024 and Reflections Alessio and Swyx wrap up with reflections on their top 2024 episodes, the lack of a real AI winter, and community growth.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 100% · guest 0%0:00 · the hosts 100% · guest 0%3:00 · the hosts 100% · guest 0%3:00 · the hosts 100% · guest 0%6:00 · the hosts 100% · guest 0%6:00 · the hosts 100% · guest 0%9:00 · the hosts 100% · guest 0%9:00 · the hosts 100% · guest 0%12:00 · the hosts 99.8% · guest 0.2%12:00 · the hosts 99.8% · guest 0.2%15:00 · the hosts 99.9% · guest 0.1%15:00 · the hosts 99.9% · guest 0.1%18:00 · the hosts 99.3% · guest 0.7%18:00 · the hosts 99.3% · guest 0.7%21:00 · the hosts 100% · guest 0%21:00 · the hosts 100% · guest 0%24:00 · the hosts 100% · guest 0%24:00 · the hosts 100% · guest 0%27:00 · the hosts 100% · guest 0%27:00 · the hosts 100% · guest 0%30:00 · the hosts 100% · guest 0%30:00 · the hosts 100% · guest 0%33:00 · the hosts 100% · guest 0%33:00 · the hosts 100% · guest 0%36:00 · the hosts 99.9% · guest 0.1%36:00 · the hosts 99.9% · guest 0.1%39:00 · the hosts 100% · guest 0%39:00 · the hosts 100% · guest 0%42:00 · the hosts 100% · guest 0%42:00 · the hosts 100% · guest 0%45:00 · the hosts 98.8% · guest 1.2%45:00 · the hosts 98.8% · guest 1.2%48:00 · the hosts 99.9% · guest 0.1%48:00 · the hosts 99.9% · guest 0.1%51:00 · the hosts 100% · guest 0%51:00 · the hosts 100% · guest 0%54:00 · the hosts 100% · guest 0%54:00 · the hosts 100% · guest 0%57:00 · the hosts 100% · guest 0%57:00 · the hosts 100% · guest 0%1:00:00 · the hosts 100% · guest 0%1:00:00 · the hosts 100% · guest 0%1:03:00 · the hosts 99.9% · guest 0.1%1:03:00 · the hosts 99.9% · guest 0.1%1:06:00 · the hosts 100% · guest 0%1:06:00 · the hosts 100% · guest 0%1:09:00 · the hosts 99.9% · guest 0.1%1:09:00 · the hosts 99.9% · guest 0.1%1:12:00 · the hosts 100% · guest 0%1:12:00 · the hosts 100% · guest 0%1:15:00 · the hosts 100% · guest 0%1:15:00 · the hosts 100% · guest 0%1:18:00 · the hosts 100% · guest 0%1:18:00 · the hosts 100% · guest 0%1:21:00 · the hosts 100% · guest 0%1:21:00 · the hosts 100% · guest 0%1:24:00 · the hosts 100% · guest 0%1:24:00 · the hosts 100% · guest 0%1:27:00 · the hosts 100% · guest 0%1:27:00 · the hosts 100% · guest 0%1:30:00 · the hosts 100% · guest 0%1:30:00 · the hosts 100% · guest 0%1:33:00 · the hosts 100% · guest 0%1:33:00 · the hosts 100% · guest 0%1:36:00 · the hosts 100% · guest 0%1:36:00 · the hosts 100% · guest 0%1:39:00 · the hosts 99.9% · guest 0.1%1:39:00 · the hosts 99.9% · guest 0.1%1:42:00 · the hosts 99.8% · guest 0.2%1:42:00 · the hosts 99.8% · guest 0.2%1:45:00 · the hosts 100% · guest 0%1:45:00 · the hosts 100% · guest 0%1:48:00 · the hosts 99.6% · guest 0.4%1:48:00 · the hosts 99.6% · guest 0.4%1:51:00 · the hosts 0% · guest 0%1:51:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 48:53 Alessio defends Midjourney workflow over raw API models

Alessio sharply dismisses Swyx's 'skill issue' jab regarding Flux by framing usability and time efficiency as a Black Forest failure.

Hardest push from the hosts ▶ 15:10 Swyx rejects the open source convergence narrative

Swyx directly calls out saturation charts being near 100% as misleading evidence that open source models are catching up to o1.

Biggest teaching moment ▶ 1:01:15 Swyx breaks down memory vs vector database architecture

Swyx systematically clarifies why memory requires interaction history, decay rates, and sleep consolidation rather than standard RAG vector lookups.

The host holds their own ▶ 11:20 Swyx details the frontier model pricing shift and market share

Swyx demonstrates deep domain mastery by detailing exact Ramp and OpenRouter dataset metrics showing OpenAI's market share dropping from 95% to 50-75%.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Celebrating 100 Episodes and the AI Engineer Era 6000 Alessio and Swyx open their 100th episode by celebrating the validation of the AI Engineer movement and discussing its industry recognition via Gartner's hype curve.
Research vs Production and the Big Scaling Debate at NeurIPS 7111 Swyx recounts the NeurIPS debates on pre-training walls, contrasting test-time compute with inference-time compute optimality alongside papers from Chinchilla to Noam Brown.
Frontier Model Competition, Market Shifts, and Inference Economics 8122 Swyx challenges claims that open source is closing the gap with frontier reasoning models, dissecting market share shifts across OpenAI, Anthropic, and Gemini.
Agent Frontiers, Systems ML, and Ilya's Keynote Insights 7111 The hosts discuss agent frontier challenges, Jeff Dean's systems ML approaches with AlphaChip, and Ilya Sutskever's historical scaling predictions.
Long-Tail NeurIPS Discoveries: Steganography and Dataset Trends 7000 Swyx outlines long-tail research from NeurIPS, including DeepMind's steganography paper on agent collusion and the shift toward dataset tracks.
The Four Wars of AI: Data Rights, Synthetic Data, and Reasoning Moats 7111 The hosts analyze the data wars, synthetic reasoning distillation via STaR methods, and the durability of reasoning moats like o1.
The Compute War: GPU Super-Rich Clusters vs GPU-Poor Practicality 7111 Swyx introduces the GPU smiling curve, comparing 100k-cluster mega-labs with lean customer-facing wrappers like Suno and Bolt.
The Multimodal War: Specialist Startups vs Integrated God Models 7222 Swyx and Alessio debate specialist multi-modal startups against integrated omni-models, bantering over Flux, Midjourney, and Sora.
The LLMOS & Agent War: Sandboxes, Memory Architecture, and Protocols 8222 Alessio evaluates PyPI metrics for LangChain and CrewAI while Swyx elaborates on why memory architectures need temporal decay and sleep consolidation.
The Shifting Evaluation Landscape: Benchmarks and Capabilities Tiering 8111 Swyx reviews benchmark degradation from MMLU and GPQA to SWE-bench verified, setting up the framework for mature versus emerging capabilities.
Editing Room Interlude: Capability Breakdown and the Price Collapse 0000 Swyx delivers an editing room monologue charting the 3-orders-of-magnitude price-intelligence collapse throughout 2024.
2024 Chronological AI Rewind: January to June Milestones 0000 Swyx continues his monologue rewind from January to June, touching on Perplexity, Devin's launch, Suno/Udio, Llama 3, and GPT-4o.
2024 Chronological AI Rewind: July to December & The Rise of Agents 8111 The hosts reunite to discuss SSI's single-product mission, o1's launch velocity, Canvas vs Google Docs, and how AI will establish workplace skill floors.
Top Latent Space Episodes of 2024 and Reflections 6000 Alessio and Swyx wrap up with reflections on their top 2024 episodes, the lack of a real AI winter, and community growth.

Statements from this episode (29)

Prediction Not checkable as stated
Swyx: AI field will invert from research-heavy to engineering-heavy
“I think the AI world, the ML world is still very much research heavy. And that's as it should be because ML is very much in a research phase. But as we move this entire field into production, I think that ratio inverts into becoming more engineering heavy.”
Shawn Wang Jan 1, 2025 ▶ 3:26
Insight
Swyx: Inference compute optimality creates different scaling laws than Chinchilla
“Chinchilla paper is compute optimal training, but what is not stated in there is it's pre-trained compute optimal training. And once you start caring about inference, compute optimal training, you have a different scaling law and in a way that we did not know …”
Shawn Wang Jan 1, 2025 ▶ 8:32
Assertion Contradicted
Swyx: Gemini Flash accounts for 50% of OpenRouter requests
“Gemini Flash, according to Open Router, is now 50% of their Open Router requests.”
Shawn Wang Jan 1, 2025 ▶ 12:04
Assertion Partly supported
Swyx: OpenAI production market share dropped from 95% to 50–75%
“Basically over the course of 23, going into 24, OpenAI has gone from 95 market share to reasonably somewhere between 50 to 75 market share.”
Shawn Wang Jan 1, 2025 ▶ 12:20
Prediction Open · timeframe Jan 2028
Swyx predicts OpenAI will launch a $2,000 per month ChatGPT tier
“I think that 2000 dollars ChatGPT will come.”
Shawn Wang Jan 1, 2025 ▶ 18:02
Opinion
Swyx: AI VCs and content creators should spend $1,000 monthly on AI tools
“I think if your job is a, at least AI content creator, or VC, or, you know, someone who, whose job it is to stay on top of things, you should already be spending like a thousand dollars a month on, on stuff.”
Shawn Wang Jan 1, 2025 ▶ 18:18
Assertion Supported
Swyx: OpenHands is ranked number one on SWE-bench Full
“He started open hand is currently still number one on SweetBench full, which is the hardest one.”
Shawn Wang Jan 1, 2025 ▶ 19:45
Insight
Fanelli: Agents extracting undocumented business processes will unlock enterprise adoption
“The agents are, that most people are building, Are good at following instruction, but are not as good as like extracting them from you. Yeah. So I think that will be a big unlock.”
Alessio Fanelli Jan 1, 2025 ▶ 20:42
Opinion
Swyx: Model papers at NeurIPS are dead as research shifts to datasets
“The focus I saw is that model papers at NeurIPS are kind of dead. No one really presents models anymore. It's just data sets because it's all the grad students are working on.”
Shawn Wang Jan 1, 2025 ▶ 30:02
Assertion Not checkable as stated
Swyx: Frontier labs view new learning rate optimizers as unimportant
“Most people at the big labs who I asked about this say that it's cute, but it's not something that matters.”
Shawn Wang Jan 1, 2025 ▶ 32:08
Assertion Partly supported
Swyx: Claude Sonnet and Gemini Outperform o1-Preview in Coding
“Claude Sonnet so far is beating O-one on coding tasks without At least one preview without being a reasoning model and same for Gemini pro or Gemini two point O.”
Shawn Wang Jan 1, 2025 ▶ 37:15
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Shawn Wang Jan 1, 2025 ▶ 38:52
Prediction Not checkable as stated
Swyx: AI models hit a 2T parameter wall, won't reach 10T
“As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT 4.5 is not coming out. And Gemini two, like we don't have pro whatever we've hit that wall, whatever that wall is. Maybe I'll call it like the two trillion parameter wall. Like …”
Shawn Wang Jan 1, 2025 ▶ 40:13
Assertion Partly supported
Swyx: Suno grew from zero to $20M ARR running on Modal
“Suno ramp has rated as one of the top ranked fastest growing startups of the year. I think the last public number is like zero to twenty million this year in ARR and Suno runs on Moto. So Suno itself is not GPU rich, but they're just doing the training on, on …”
Shawn Wang Jan 1, 2025 ▶ 44:12
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Shawn Wang Jan 1, 2025 ▶ 44:32
Insight
Swyx: The 'GPU smiling curve' rewards infrastructure and user-facing apps over middle-tier models
“I kind of call this the GPU smiling curve where the edges do well, cause you're either close to the machines and you're like number one on the machines, or you're like close to the customers and you're number one on the customer side. And the people who are in…”
Shawn Wang Jan 1, 2025 ▶ 44:52
Assertion Supported
Swyx: HeyGen has reached $100 million in ARR
“Hey Jen, I think has reached a hundred million ARR.”
Shawn Wang Jan 1, 2025 ▶ 47:08
Assertion Supported
Swyx: Recraft V3 has overtaken Flux 1.1 as the top image model
“Recraft V three is now being Flux 1.1, which is very surprising because Flux and Black Forest Labs are the old stable diffusion crew who left Stability after the management issues. So ReCraft has come from nowhere to be the top image model.”
Shawn Wang Jan 1, 2025 ▶ 49:55
Prediction Not checkable as stated
Swyx: AI video generation needs five more years to become fully SOTA
“And so the researchers that I talked to are already starting to talk about that as the next frontier, but there's still maybe like five more years of video left. To actually be soda.”
Shawn Wang Jan 1, 2025 ▶ 51:14
Opinion
Swyx: DeepMind has roughly a four-year advantage over OpenAI in world modeling
“So like they have maybe four years advantage on world modeling that OpenAI does not have. Cause OpenAI basically only started Diffusion Transformers last year when they hired Build Peebles. So DeepMind has a bit of advantage here.”
Shawn Wang Jan 1, 2025 ▶ 51:53
Assertion Supported
Fanelli: LangChain and LlamaIndex usage grows on PyPI while CrewAI stays flat
“So if you look at, you know, like chain still growing. These are the last, last six months. Lama index still growing. What I've basically seen is like things that one, obviously these things have a commercial product. So there's like people buying this and sti…”
Alessio Fanelli Jan 1, 2025 ▶ 53:17
Prediction Not checkable as stated
Swyx: Diff mode will become the norm for AI code tools in 2025
“Canvas has incorporated the diff mode that both Anthropic and OpenAI and Fireworks has now shipped that I think is going to be the norm for next year, that everyone Need some kind of diff mode code interpreter thing.”
Shawn Wang Jan 1, 2025 ▶ 59:25
Assertion Supported
Swyx: SWE-bench resolution rates surged from 13% to ~50% in 2024
“Keep in mind, we started the year at 13%. And so now we're about 50 open hands is around there.”
Shawn Wang Jan 1, 2025 ▶ 1:08:29
Insight
Swyx: Frontier AI labs distinguish themselves by adopting new benchmarks
“The labs that are not that frontier will keep measuring themselves on last year's benchmarks. And then the labs that are actually frontier will tell you about benchmarks you've never heard of.”
Shawn Wang Jan 1, 2025 ▶ 1:08:52
Assertion Supported
Swyx: Iso-ELO LLM token costs dropped nearly 1,000x in 2024
“For the same amount of ELO, what you used to pay at the start of 24 you know, let's say, you know, 50, 40 to 50 dollars Per million tokens now is available approximately at, with Amazon Nova approximately at, I don't know, 0.075 dollars per token. So like seve…”
Shawn Wang Jan 1, 2025 ▶ 1:20:51
Assertion Not checkable as stated
Swyx: Apple Intelligence is the largest transformer rollout since Google's BERT
“It is the probably the largest scale rollout of transformers yet after Google rolled out BERT for search”
Shawn Wang Jan 1, 2025 ▶ 1:29:12
Prediction Not checkable as stated
Swyx: Always-on vision AI assistants will dominate desktop software by late 2025
“And like this time next year, I would be willing to bet that I would just have this running on my machine. And you know, I think That assistance always on that you can talk to with vision that sees what you're seeing. I think that is where at least one hour so…”
Shawn Wang Jan 1, 2025 ▶ 1:32:43
Prediction Not checkable as stated
Fanelli: 2025 will be the first year AI sets job skill floors
“And I think the skill floor more and more, I think, 20, 25 will be the first year where the AI sets the skill floor of a role.”
Alessio Fanelli Jan 1, 2025 ▶ 1:41:17
Assertion Supported
Fanelli: 133 AI funding rounds exceeded $100M in 2024
“There have been 133 funding rounds over a hundred million in AI this year.”
Alessio Fanelli Jan 1, 2025 ▶ 1:48:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.