Feb 13, 2025 · 22m · latent-space

smol agents are all you need

Aymeric (Emmerich) · 13m spoken Swyx (Marcos Swix) · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Hugging Face's Aymeric joins the Lightning Pod to discuss smolagents and the power of code-first agent architectures over standard JSON tool calling. The episode explores agent evaluation on the GAIA benchmark, secure sandbox environments, and the future transition toward multimodal computer-using agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.5 Guest teaching 5.7 Guest disagreement 0.7 The hosts pushing back 1.5
05100:0010:0020:000:02–2:31 · The hosts as informed peer 5/10 Defining AI Agents and the Agency Spectrum Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use.2:31–6:30 · The hosts as informed peer 6/10 Smolagents Philosophy: Code Agents vs. JSON Function Calling Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper.6:30–11:09 · The hosts as informed peer 5/10 Hugging Face Ecosystem, Sandbox Security, and Agent Course Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B.11:09–15:14 · The hosts as informed peer 6/10 Benchmarking General AI Assistants on the GAIA Leaderboard Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two.15:14–20:47 · The hosts as informed peer 6/10 The Frontier of Computer-Using Agents and GUI Automation Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure.20:48–22:16 · The hosts as informed peer 5/10 Model Limitations, 2025 Outlook, and Open Source Roadmap Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown.0:02–2:31 · Guest teaching 6/10 Defining AI Agents and the Agency Spectrum Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use.2:31–6:30 · Guest teaching 6/10 Smolagents Philosophy: Code Agents vs. JSON Function Calling Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper.6:30–11:09 · Guest teaching 5/10 Hugging Face Ecosystem, Sandbox Security, and Agent Course Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B.11:09–15:14 · Guest teaching 6/10 Benchmarking General AI Assistants on the GAIA Leaderboard Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two.15:14–20:47 · Guest teaching 6/10 The Frontier of Computer-Using Agents and GUI Automation Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure.20:48–22:16 · Guest teaching 5/10 Model Limitations, 2025 Outlook, and Open Source Roadmap Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown.0:02–2:31 · Guest disagreement 0/10 Defining AI Agents and the Agency Spectrum Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use.2:31–6:30 · Guest disagreement 1/10 Smolagents Philosophy: Code Agents vs. JSON Function Calling Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper.6:30–11:09 · Guest disagreement 0/10 Hugging Face Ecosystem, Sandbox Security, and Agent Course Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B.11:09–15:14 · Guest disagreement 1/10 Benchmarking General AI Assistants on the GAIA Leaderboard Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two.15:14–20:47 · Guest disagreement 2/10 The Frontier of Computer-Using Agents and GUI Automation Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure.20:48–22:16 · Guest disagreement 0/10 Model Limitations, 2025 Outlook, and Open Source Roadmap Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown.0:02–2:31 · The hosts pushing back 1/10 Defining AI Agents and the Agency Spectrum Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use.2:31–6:30 · The hosts pushing back 1/10 Smolagents Philosophy: Code Agents vs. JSON Function Calling Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper.6:30–11:09 · The hosts pushing back 1/10 Hugging Face Ecosystem, Sandbox Security, and Agent Course Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B.11:09–15:14 · The hosts pushing back 2/10 Benchmarking General AI Assistants on the GAIA Leaderboard Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two.15:14–20:47 · The hosts pushing back 3/10 The Frontier of Computer-Using Agents and GUI Automation Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure.20:48–22:16 · The hosts pushing back 1/10 Model Limitations, 2025 Outlook, and Open Source Roadmap Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 18:47 Aymeric defending GUI generality over computer-use terminology

Aymeric pushes back against Swyx's attempt to standardize on the CUA term, arguing that GUI automation represents a broader problem space than computer-only workflows.

Hardest push from the hosts ▶ 18:25 Swyx correcting the GUI label to CUA

Swyx interrupts the framing to insist on industry terminology adopted by Anthropic and OpenAI, citing internal discussions with Karina.

Biggest teaching moment ▶ 4:10 Aymeric demonstrating code agents superiority over JSON

Aymeric breaks down the architectural inefficiency of standard JSON tool calling, illustrating how code agents handle parallel loops and variable assignment natively.

The host holds their own ▶ 14:00 Alessio isolating GAIA benchmark performance divergence

Alessio demonstrates sharp analytical oversight by directly challenging the discrepancy between identical level two benchmark results and subsequent level three divergence.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Defining AI Agents and the Agency Spectrum 5601 Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use.
Smolagents Philosophy: Code Agents vs. JSON Function Calling 6611 Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper.
Hugging Face Ecosystem, Sandbox Security, and Agent Course 5501 Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B.
Benchmarking General AI Assistants on the GAIA Leaderboard 6612 Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two.
The Frontier of Computer-Using Agents and GUI Automation 6623 Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure.
Model Limitations, 2025 Outlook, and Open Source Roadmap 5501 Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown.

Statements from this episode (8)

Assertion Supported
Roucher: smolagents core agents.py file is under 1,000 lines
“The main file in it, the agents.py file that we have at the core of the library is under the 1000 lines of code.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 3:40
Opinion
Roucher: Code agents perform better than JSON tool calling
“I found this is highly suboptimal, and that's why I've designed the other code agent, which is really like the core opinion in the library is that code agents work better.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 4:42
Insight
Roucher: DeepSeek-R1 ranks slightly below OpenAI o1 on smolagents tasks
“I tried R one, but R one is a bit under O one with small agents. And I think this is also a matter of formatting. Like sometimes the model struggles to just output them, the code snippets in the correct way that we expect.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 14:25
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 16:12
Opinion
Roucher: Effective web browsing agents must use vision and direct GUI inputs
“Web browsing is designed for humans. So that means it's really visual. And so a web browsing agent, a good one, should use, in my opinion a vision model and perform actions with point and click and keyboard, basically.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 17:26
Disclosure
Roucher says Hugging Face plans to heavily prioritize GUI agents
“That's the next step for us at Hugging Face. We're going to push really hard on building GUI agents, so basically agents that can use any GUI.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 17:42
Prediction Not checkable as stated
Roucher: Visual AI models will likely jump the reasoning S-curve in 2025
“But as with text agents, we've really found that we made a jump on the S curve with reasoning models. I think it's going to be the same with the next visual models, basically better base models just allow you to jump over this S curve. And probably I think it'…”
Aymeric (Emmerich) Feb 13, 2025 ▶ 21:21
Disclosure
Roucher: Hugging Face will fine-tune DeepSeek-R1 for agent workflows
“Looking forward to this, and for instance we're going to fine-tune R-One on agentic stuff, so we'll, we'll get really good powerhouse models soon.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 22:05
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.