Feb 13, 2025 · 22m · latent-space
smol agents are all you need
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Hugging Face's Aymeric joins the Lightning Pod to discuss smolagents and the power of code-first agent architectures over standard JSON tool calling. The episode explores agent evaluation on the GAIA benchmark, secure sandbox environments, and the future transition toward multimodal computer-using agents.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Aymeric pushes back against Swyx's attempt to standardize on the CUA term, arguing that GUI automation represents a broader problem space than computer-only workflows.
Hardest push from the hosts ▶ 18:25 Swyx correcting the GUI label to CUASwyx interrupts the framing to insist on industry terminology adopted by Anthropic and OpenAI, citing internal discussions with Karina.
Biggest teaching moment ▶ 4:10 Aymeric demonstrating code agents superiority over JSONAymeric breaks down the architectural inefficiency of standard JSON tool calling, illustrating how code agents handle parallel loops and variable assignment natively.
The host holds their own ▶ 14:00 Alessio isolating GAIA benchmark performance divergenceAlessio demonstrates sharp analytical oversight by directly challenging the discrepancy between identical level two benchmark results and subsequent level three divergence.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Defining AI Agents and the Agency Spectrum | 5 | 6 | 0 | 1 | Aymeric defines agents across a continuum citing Harrison Chase's framework. Swyx complements this by referencing Lilian Weng's taxonomy of planning, memory, and tool use. | |
| Smolagents Philosophy: Code Agents vs. JSON Function Calling | 6 | 6 | 1 | 1 | Aymeric explains why code agents outperform traditional JSON tool calling for parallel loops and variable storage. Swyx demonstrates domain familiarity by citing their interview with Graham Neubig on the CodeAct paper. | |
| Hugging Face Ecosystem, Sandbox Security, and Agent Course | 5 | 5 | 0 | 1 | Aymeric discusses the evolution of Hugging Face agents and custom sandboxed interpreters. Swyx and Alessio engage on sandbox execution environments, noting their investor background with E2B. | |
| Benchmarking General AI Assistants on the GAIA Leaderboard | 6 | 6 | 1 | 2 | Aymeric outlines the GAIA benchmark methodology and leaderboard nuances. Alessio pushes into the data by asking why scores drop off significantly on level three despite matching level two. | |
| The Frontier of Computer-Using Agents and GUI Automation | 6 | 6 | 2 | 3 | Aymeric presents his timeline for solving GAIA and advocates for vision-based GUI automation. Swyx pushes back on terminology, preferring Anthropic and OpenAI's CUA designation, and cites Morph Labs' branching infrastructure. | |
| Model Limitations, 2025 Outlook, and Open Source Roadmap | 5 | 5 | 0 | 1 | Alessio asks whether near-term agent adoption is bottlenecked by UX or foundational model capability. Aymeric highlights the current limitations of vision-language models on screenshots versus raw markdown. |