Aug 2, 2024 · 1h 23m · latent-space

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)

Alessio Fanelli · 21m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this quarterly recap recorded in Singapore, Latent Space Podcast co-hosts Alessio Fanelli and Swix analyze the 'Four Wars of the AI Stack,' assessing frontier model competition, hardware efficiency, multimodal advancements, and the transition toward agentic LLM operating systems amid rapid economic depreciation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.8% of the talking time here. How this is scored →

The hosts as informed peer 5.9 Guest teaching 3.0 Guest disagreement 2.1 The hosts pushing back 2.3
05100:0020:0040:001:00:001:20:000:04–3:43 · The hosts as informed peer 5/10 Welcome to Singapore and Sovereign AI Discussion Alessio sets the stage in Singapore, explaining the Sovereign AI Summit context and institutional interest in state space models. Swix elaborates on AI engineering as an alternative to massive national compute clusters.3:44–16:41 · The hosts as informed peer 6/10 Frontier Model Wars: Claude 3.5 Sonnet and Llama 3.1 The co-hosts analyze Claude 3.5 Sonnet's benchmark dominance and Anthropic's monosemanticity research alongside Llama 3.1 405B's synthetic data roadmap. Alessio notes the shift toward distilling large models into smaller form factors.16:42–20:05 · The hosts as informed peer 5/10 Open Weights Dynamics: Mistral's Shifting Crown They discuss Mistral's release of Mistral Large 2 right after Llama 3.1 and evaluate how Mistral may have lost its open source crown to Meta due to license restrictions.20:05–33:51 · The hosts as informed peer 6/10 Hardware Moats, Inference Efficiency, and On-Device AI Alessio highlights FlashAttention-3 and market capitalization dynamics, while Swix breaks down Character.AI's hybrid local-global attention mechanisms and on-device runtimes like llamafile and Gemini Nano.33:51–44:11 · The hosts as informed peer 6/10 Data Quality Wars: Copyright Battles, Licensing, and AlphaProof Alessio points out Reddit's lucrative licensing deals and potential FTC antitrust scrutiny over data access. Swix discusses the IMO AlphaProof achievements and the concept of jagged intelligence.44:11–51:36 · The hosts as informed peer 5/10 Multimodal Wars: Real-Time Voice, Early Fusion, and Document Intelligence Swix details the real-time audio demo challenges at AI Engineer World's Fair and compares early-fusion architectures like Chameleon to late-fusion adapter approaches. Alessio references ColPali for PDF extraction.51:37–1:09:33 · The hosts as informed peer 7/10 LLM OS and the Proliferation of Agent Ecosystems Swix rebrands RAG/Ops to LLM OS and debates Alessio on the value of LLM ops infrastructure versus vertical capabilities. They further explore memory databases, agent coordination, and protocol standards.1:09:35–1:23:24 · The hosts as informed peer 7/10 The Winds of AI Winter: Depreciation Curves and Post-MMLU Benchmarks Swix presents his depreciation curve showing intelligence cost dropping an order of magnitude every four months, but Alessio stumps him by asking whether frontier models like GPT-Next will reset the curve upward. Alessio elaborates on selling services-as-software labor.0:04–3:43 · Guest teaching 2/10 Welcome to Singapore and Sovereign AI Discussion Alessio sets the stage in Singapore, explaining the Sovereign AI Summit context and institutional interest in state space models. Swix elaborates on AI engineering as an alternative to massive national compute clusters.3:44–16:41 · Guest teaching 3/10 Frontier Model Wars: Claude 3.5 Sonnet and Llama 3.1 The co-hosts analyze Claude 3.5 Sonnet's benchmark dominance and Anthropic's monosemanticity research alongside Llama 3.1 405B's synthetic data roadmap. Alessio notes the shift toward distilling large models into smaller form factors.16:42–20:05 · Guest teaching 3/10 Open Weights Dynamics: Mistral's Shifting Crown They discuss Mistral's release of Mistral Large 2 right after Llama 3.1 and evaluate how Mistral may have lost its open source crown to Meta due to license restrictions.20:05–33:51 · Guest teaching 3/10 Hardware Moats, Inference Efficiency, and On-Device AI Alessio highlights FlashAttention-3 and market capitalization dynamics, while Swix breaks down Character.AI's hybrid local-global attention mechanisms and on-device runtimes like llamafile and Gemini Nano.33:51–44:11 · Guest teaching 3/10 Data Quality Wars: Copyright Battles, Licensing, and AlphaProof Alessio points out Reddit's lucrative licensing deals and potential FTC antitrust scrutiny over data access. Swix discusses the IMO AlphaProof achievements and the concept of jagged intelligence.44:11–51:36 · Guest teaching 3/10 Multimodal Wars: Real-Time Voice, Early Fusion, and Document Intelligence Swix details the real-time audio demo challenges at AI Engineer World's Fair and compares early-fusion architectures like Chameleon to late-fusion adapter approaches. Alessio references ColPali for PDF extraction.51:37–1:09:33 · Guest teaching 4/10 LLM OS and the Proliferation of Agent Ecosystems Swix rebrands RAG/Ops to LLM OS and debates Alessio on the value of LLM ops infrastructure versus vertical capabilities. They further explore memory databases, agent coordination, and protocol standards.1:09:35–1:23:24 · Guest teaching 3/10 The Winds of AI Winter: Depreciation Curves and Post-MMLU Benchmarks Swix presents his depreciation curve showing intelligence cost dropping an order of magnitude every four months, but Alessio stumps him by asking whether frontier models like GPT-Next will reset the curve upward. Alessio elaborates on selling services-as-software labor.0:04–3:43 · Guest disagreement 1/10 Welcome to Singapore and Sovereign AI Discussion Alessio sets the stage in Singapore, explaining the Sovereign AI Summit context and institutional interest in state space models. Swix elaborates on AI engineering as an alternative to massive national compute clusters.3:44–16:41 · Guest disagreement 2/10 Frontier Model Wars: Claude 3.5 Sonnet and Llama 3.1 The co-hosts analyze Claude 3.5 Sonnet's benchmark dominance and Anthropic's monosemanticity research alongside Llama 3.1 405B's synthetic data roadmap. Alessio notes the shift toward distilling large models into smaller form factors.16:42–20:05 · Guest disagreement 2/10 Open Weights Dynamics: Mistral's Shifting Crown They discuss Mistral's release of Mistral Large 2 right after Llama 3.1 and evaluate how Mistral may have lost its open source crown to Meta due to license restrictions.20:05–33:51 · Guest disagreement 2/10 Hardware Moats, Inference Efficiency, and On-Device AI Alessio highlights FlashAttention-3 and market capitalization dynamics, while Swix breaks down Character.AI's hybrid local-global attention mechanisms and on-device runtimes like llamafile and Gemini Nano.33:51–44:11 · Guest disagreement 2/10 Data Quality Wars: Copyright Battles, Licensing, and AlphaProof Alessio points out Reddit's lucrative licensing deals and potential FTC antitrust scrutiny over data access. Swix discusses the IMO AlphaProof achievements and the concept of jagged intelligence.44:11–51:36 · Guest disagreement 1/10 Multimodal Wars: Real-Time Voice, Early Fusion, and Document Intelligence Swix details the real-time audio demo challenges at AI Engineer World's Fair and compares early-fusion architectures like Chameleon to late-fusion adapter approaches. Alessio references ColPali for PDF extraction.51:37–1:09:33 · Guest disagreement 4/10 LLM OS and the Proliferation of Agent Ecosystems Swix rebrands RAG/Ops to LLM OS and debates Alessio on the value of LLM ops infrastructure versus vertical capabilities. They further explore memory databases, agent coordination, and protocol standards.1:09:35–1:23:24 · Guest disagreement 3/10 The Winds of AI Winter: Depreciation Curves and Post-MMLU Benchmarks Swix presents his depreciation curve showing intelligence cost dropping an order of magnitude every four months, but Alessio stumps him by asking whether frontier models like GPT-Next will reset the curve upward. Alessio elaborates on selling services-as-software labor.0:04–3:43 · The hosts pushing back 1/10 Welcome to Singapore and Sovereign AI Discussion Alessio sets the stage in Singapore, explaining the Sovereign AI Summit context and institutional interest in state space models. Swix elaborates on AI engineering as an alternative to massive national compute clusters.3:44–16:41 · The hosts pushing back 2/10 Frontier Model Wars: Claude 3.5 Sonnet and Llama 3.1 The co-hosts analyze Claude 3.5 Sonnet's benchmark dominance and Anthropic's monosemanticity research alongside Llama 3.1 405B's synthetic data roadmap. Alessio notes the shift toward distilling large models into smaller form factors.16:42–20:05 · The hosts pushing back 2/10 Open Weights Dynamics: Mistral's Shifting Crown They discuss Mistral's release of Mistral Large 2 right after Llama 3.1 and evaluate how Mistral may have lost its open source crown to Meta due to license restrictions.20:05–33:51 · The hosts pushing back 2/10 Hardware Moats, Inference Efficiency, and On-Device AI Alessio highlights FlashAttention-3 and market capitalization dynamics, while Swix breaks down Character.AI's hybrid local-global attention mechanisms and on-device runtimes like llamafile and Gemini Nano.33:51–44:11 · The hosts pushing back 2/10 Data Quality Wars: Copyright Battles, Licensing, and AlphaProof Alessio points out Reddit's lucrative licensing deals and potential FTC antitrust scrutiny over data access. Swix discusses the IMO AlphaProof achievements and the concept of jagged intelligence.44:11–51:36 · The hosts pushing back 1/10 Multimodal Wars: Real-Time Voice, Early Fusion, and Document Intelligence Swix details the real-time audio demo challenges at AI Engineer World's Fair and compares early-fusion architectures like Chameleon to late-fusion adapter approaches. Alessio references ColPali for PDF extraction.51:37–1:09:33 · The hosts pushing back 4/10 LLM OS and the Proliferation of Agent Ecosystems Swix rebrands RAG/Ops to LLM OS and debates Alessio on the value of LLM ops infrastructure versus vertical capabilities. They further explore memory databases, agent coordination, and protocol standards.1:09:35–1:23:24 · The hosts pushing back 4/10 The Winds of AI Winter: Depreciation Curves and Post-MMLU Benchmarks Swix presents his depreciation curve showing intelligence cost dropping an order of magnitude every four months, but Alessio stumps him by asking whether frontier models like GPT-Next will reset the curve upward. Alessio elaborates on selling services-as-software labor.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 55.4% · guest 44.6%0:00 · the hosts 55.4% · guest 44.6%3:00 · the hosts 29.4% · guest 70.6%3:00 · the hosts 29.4% · guest 70.6%6:00 · the hosts 39.9% · guest 60.1%6:00 · the hosts 39.9% · guest 60.1%9:00 · the hosts 55.7% · guest 44.3%9:00 · the hosts 55.7% · guest 44.3%12:00 · the hosts 33.3% · guest 66.7%12:00 · the hosts 33.3% · guest 66.7%15:00 · the hosts 6.7% · guest 93.3%15:00 · the hosts 6.7% · guest 93.3%18:00 · the hosts 33% · guest 67%18:00 · the hosts 33% · guest 67%21:00 · the hosts 7.1% · guest 92.9%21:00 · the hosts 7.1% · guest 92.9%24:00 · the hosts 28.6% · guest 71.4%24:00 · the hosts 28.6% · guest 71.4%27:00 · the hosts 23.3% · guest 76.7%27:00 · the hosts 23.3% · guest 76.7%30:00 · the hosts 24.8% · guest 75.2%30:00 · the hosts 24.8% · guest 75.2%33:00 · the hosts 30% · guest 70%33:00 · the hosts 30% · guest 70%36:00 · the hosts 33.5% · guest 66.5%36:00 · the hosts 33.5% · guest 66.5%39:00 · the hosts 14.6% · guest 85.4%39:00 · the hosts 14.6% · guest 85.4%42:00 · the hosts 19.4% · guest 80.6%42:00 · the hosts 19.4% · guest 80.6%45:00 · the hosts 6.2% · guest 93.8%45:00 · the hosts 6.2% · guest 93.8%48:00 · the hosts 23.7% · guest 76.3%48:00 · the hosts 23.7% · guest 76.3%51:00 · the hosts 44.5% · guest 55.5%51:00 · the hosts 44.5% · guest 55.5%54:00 · the hosts 27.6% · guest 72.4%54:00 · the hosts 27.6% · guest 72.4%57:00 · the hosts 44.2% · guest 55.8%57:00 · the hosts 44.2% · guest 55.8%1:00:00 · the hosts 23.7% · guest 76.3%1:00:00 · the hosts 23.7% · guest 76.3%1:03:00 · the hosts 7.5% · guest 92.5%1:03:00 · the hosts 7.5% · guest 92.5%1:06:00 · the hosts 28% · guest 72%1:06:00 · the hosts 28% · guest 72%1:09:00 · the hosts 9.9% · guest 90.1%1:09:00 · the hosts 9.9% · guest 90.1%1:12:00 · the hosts 31.5% · guest 68.5%1:12:00 · the hosts 31.5% · guest 68.5%1:15:00 · the hosts 73.5% · guest 26.5%1:15:00 · the hosts 73.5% · guest 26.5%1:18:00 · the hosts 34.7% · guest 65.3%1:18:00 · the hosts 34.7% · guest 65.3%1:21:00 · the hosts 14.9% · guest 85.1%1:21:00 · the hosts 14.9% · guest 85.1%
Sharpest disagreement ▶ 54:10 Swix disagrees on ops startup value

Swix directly counters Alessio's assertion that capability plugins matter more than ops by arguing that ops remains his number one operational headache.

Hardest push from the hosts ▶ 1:11:06 Alessio challenges the continuous depreciation thesis

Alessio pushes back on Swix's downward cost curve by questioning whether upcoming frontier models will push the cost and intelligence ceiling back up.

Biggest teaching moment ▶ 47:38 Swix outlines early fusion versus adapter multimodality

Swix breaks down the architectural distinction between late-fusion adapter approaches like Llama 3 and deep native early-fusion models like Meta's Chameleon.

The host holds their own ▶ 1:17:39 Alessio breaks down the economics of AI labor replacement

Alessio provides concrete metrics on SOC alert costs ($35 human vs $6 Dropzone) to demonstrate why selling labor directly succeeds over selling productivity tooling.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome to Singapore and Sovereign AI Discussion 5211 Alessio sets the stage in Singapore, explaining the Sovereign AI Summit context and institutional interest in state space models. Swix elaborates on AI engineering as an alternative to massive national compute clusters.
Frontier Model Wars: Claude 3.5 Sonnet and Llama 3.1 6322 The co-hosts analyze Claude 3.5 Sonnet's benchmark dominance and Anthropic's monosemanticity research alongside Llama 3.1 405B's synthetic data roadmap. Alessio notes the shift toward distilling large models into smaller form factors.
Open Weights Dynamics: Mistral's Shifting Crown 5322 They discuss Mistral's release of Mistral Large 2 right after Llama 3.1 and evaluate how Mistral may have lost its open source crown to Meta due to license restrictions.
Hardware Moats, Inference Efficiency, and On-Device AI 6322 Alessio highlights FlashAttention-3 and market capitalization dynamics, while Swix breaks down Character.AI's hybrid local-global attention mechanisms and on-device runtimes like llamafile and Gemini Nano.
Data Quality Wars: Copyright Battles, Licensing, and AlphaProof 6322 Alessio points out Reddit's lucrative licensing deals and potential FTC antitrust scrutiny over data access. Swix discusses the IMO AlphaProof achievements and the concept of jagged intelligence.
Multimodal Wars: Real-Time Voice, Early Fusion, and Document Intelligence 5311 Swix details the real-time audio demo challenges at AI Engineer World's Fair and compares early-fusion architectures like Chameleon to late-fusion adapter approaches. Alessio references ColPali for PDF extraction.
LLM OS and the Proliferation of Agent Ecosystems 7444 Swix rebrands RAG/Ops to LLM OS and debates Alessio on the value of LLM ops infrastructure versus vertical capabilities. They further explore memory databases, agent coordination, and protocol standards.
The Winds of AI Winter: Depreciation Curves and Post-MMLU Benchmarks 7334 Swix presents his depreciation curve showing intelligence cost dropping an order of magnitude every four months, but Alessio stumps him by asking whether frontier models like GPT-Next will reset the curve upward. Alessio elaborates on selling services-as-software labor.

Statements from this episode (6)

Disclosure
Alessio Fanelli canceled ChatGPT subscription, switching to Claude for podcast workflows
“I canceled chat GBD a while ago. Really small podcaster run for Latent Space. It runs both on Claude and on OpenAI and Claude is definitely better most of the time.”
Alessio Fanelli Aug 2, 2024 ▶ 7:48
Assertion Supported
Reddit makes over 200 million dollars in AI data licensing deals
“Yeah, the, I guess the winner in all of this is Reddit, which is making over two hundred million just in data licensing to OpenAI and some of the other AI providers.”
Alessio Fanelli Aug 2, 2024 ▶ 36:09
Assertion Not checkable as stated
E2B scaled from 10,000 to one million cloud containers in four months
“I think they went in, like, four months from, like, 10 K to a million containers spun up on the cloud.”
Alessio Fanelli Aug 2, 2024 ▶ 53:05
Insight
Developers will trade 50 milliseconds of latency for higher model quality
“I think the biggest change in this market is like, Latency is actually not that important anymore. Like we lived in the past 10 years in a world where like 10, 15, 20 milliseconds made a big difference. I think today people will be happy to trade 50 millisecon…”
Alessio Fanelli Aug 2, 2024 ▶ 59:10
Disclosure
AI agent startups do not care about communicating with external agents
“We're investors in a bunch of agent companies. None of them really care about how to communicate with other agents. They're so focused internally”
Alessio Fanelli Aug 2, 2024 ▶ 1:07:15
Insight
Enterprise buyers want AI digital labor, not AI productivity tools
“Most companies that are buying AI tooling, they want the AI to do some sort of labor for them. And that's why the picks and shovels kind of disinterest maybe comes from a little bit. Most companies do not want to buy tools to build AI. They want the AI, and th…”
Alessio Fanelli Aug 2, 2024 ▶ 1:15:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.