Mar 11, 2025 · 25m · latent-space
The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI team members Romain Huet and Nikunj Handa join the Latent Space podcast to introduce their major agent platform releases, including the Responses API, Agents SDK, Web Search, File Search, and Computer Use tooling. They explain the architectural choices, migration pathways, and observability features designed to empower developers building autonomous multi-agent workflows.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.5% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
In a completely non-combative interview, the closest friction point is Nikunj candidly pushing back on OpenAI being one-size-fits-all, clarifying custom RAG is still superior for high-control use cases.
Hardest push from the hosts ▶ 2:21 Addressing chat completions sunset concernsSwyx directly voices developer anxiety regarding whether OpenAI is quietly sunsetting Chat Completions in favor of the new Responses API.
Biggest teaching moment ▶ 14:06 Using File Search as dynamic user memoryNikunj educates the hosts on an unexpected architectural pattern where developers use the file search vector store as user preference memory combined with web search in a single call.
The host holds their own ▶ 13:19 Swyx's advice on search hyperparametersSwyx demonstrates deep domain knowledge of open deep research implementations by advising OpenAI to implement score cutoffs rather than top-K retrieval to avoid unpredictably high token costs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Deep Dive into the Responses API and Migration Roadmap | 6 | 3 | 0 | 2 | Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence. | |
| Web Search Tool and GPT-4o Search Capabilities | 7 | 2 | 0 | 2 | Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions. | |
| Enhanced File Search and Managed RAG Infrastructure | 6 | 3 | 0 | 2 | Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks. | |
| Computer Use Agent Tooling and the Operator Model | 5 | 3 | 0 | 1 | The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps. | |
| Agents SDK Architecture, Observability, and RFT Integration | 7 | 2 | 0 | 1 | Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows. | |
| Conclusion and Final Thoughts | 1 | 0 | 0 | 0 | A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode. |