Mar 11, 2025 · 25m · latent-space

The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!

Nikunj Handa · 9m spoken Romain Huet · 6m spoken Shawn Wang · 4m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI team members Romain Huet and Nikunj Handa join the Latent Space podcast to introduce their major agent platform releases, including the Responses API, Agents SDK, Web Search, File Search, and Computer Use tooling. They explain the architectural choices, migration pathways, and observability features designed to empower developers building autonomous multi-agent workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.5% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 2.2 Guest disagreement 0.0 The hosts pushing back 1.3
05100:0010:0020:002:21–7:47 · The hosts as informed peer 6/10 Deep Dive into the Responses API and Migration Roadmap Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence.7:49–14:05 · The hosts as informed peer 7/10 Web Search Tool and GPT-4o Search Capabilities Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions.14:06–18:16 · The hosts as informed peer 6/10 Enhanced File Search and Managed RAG Infrastructure Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks.18:16–21:33 · The hosts as informed peer 5/10 Computer Use Agent Tooling and the Operator Model The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps.21:34–25:04 · The hosts as informed peer 7/10 Agents SDK Architecture, Observability, and RFT Integration Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows.25:05–25:24 · The hosts as informed peer 1/10 Conclusion and Final Thoughts A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode.2:21–7:47 · Guest teaching 3/10 Deep Dive into the Responses API and Migration Roadmap Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence.7:49–14:05 · Guest teaching 2/10 Web Search Tool and GPT-4o Search Capabilities Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions.14:06–18:16 · Guest teaching 3/10 Enhanced File Search and Managed RAG Infrastructure Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks.18:16–21:33 · Guest teaching 3/10 Computer Use Agent Tooling and the Operator Model The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps.21:34–25:04 · Guest teaching 2/10 Agents SDK Architecture, Observability, and RFT Integration Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows.25:05–25:24 · Guest teaching 0/10 Conclusion and Final Thoughts A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode.2:21–7:47 · Guest disagreement 0/10 Deep Dive into the Responses API and Migration Roadmap Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence.7:49–14:05 · Guest disagreement 0/10 Web Search Tool and GPT-4o Search Capabilities Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions.14:06–18:16 · Guest disagreement 0/10 Enhanced File Search and Managed RAG Infrastructure Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks.18:16–21:33 · Guest disagreement 0/10 Computer Use Agent Tooling and the Operator Model The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps.21:34–25:04 · Guest disagreement 0/10 Agents SDK Architecture, Observability, and RFT Integration Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows.25:05–25:24 · Guest disagreement 0/10 Conclusion and Final Thoughts A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode.2:21–7:47 · The hosts pushing back 2/10 Deep Dive into the Responses API and Migration Roadmap Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence.7:49–14:05 · The hosts pushing back 2/10 Web Search Tool and GPT-4o Search Capabilities Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions.14:06–18:16 · The hosts pushing back 2/10 Enhanced File Search and Managed RAG Infrastructure Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks.18:16–21:33 · The hosts pushing back 1/10 Computer Use Agent Tooling and the Operator Model The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps.21:34–25:04 · The hosts pushing back 1/10 Agents SDK Architecture, Observability, and RFT Integration Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows.25:05–25:24 · The hosts pushing back 0/10 Conclusion and Final Thoughts A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 30.4% · guest 69.6%0:00 · the hosts 30.4% · guest 69.6%3:00 · the hosts 23.1% · guest 76.9%3:00 · the hosts 23.1% · guest 76.9%6:00 · the hosts 45.5% · guest 54.5%6:00 · the hosts 45.5% · guest 54.5%9:00 · the hosts 37.1% · guest 62.9%9:00 · the hosts 37.1% · guest 62.9%12:00 · the hosts 31% · guest 69%12:00 · the hosts 31% · guest 69%15:00 · the hosts 28.7% · guest 71.3%15:00 · the hosts 28.7% · guest 71.3%18:00 · the hosts 25.3% · guest 74.7%18:00 · the hosts 25.3% · guest 74.7%21:00 · the hosts 8.9% · guest 91.1%21:00 · the hosts 8.9% · guest 91.1%24:00 · the hosts 25.6% · guest 74.4%24:00 · the hosts 25.6% · guest 74.4%
Sharpest disagreement ▶ 17:25 Managed service vs custom stack nuance

In a completely non-combative interview, the closest friction point is Nikunj candidly pushing back on OpenAI being one-size-fits-all, clarifying custom RAG is still superior for high-control use cases.

Hardest push from the hosts ▶ 2:21 Addressing chat completions sunset concerns

Swyx directly voices developer anxiety regarding whether OpenAI is quietly sunsetting Chat Completions in favor of the new Responses API.

Biggest teaching moment ▶ 14:06 Using File Search as dynamic user memory

Nikunj educates the hosts on an unexpected architectural pattern where developers use the file search vector store as user preference memory combined with web search in a single call.

The host holds their own ▶ 13:19 Swyx's advice on search hyperparameters

Swyx demonstrates deep domain knowledge of open deep research implementations by advising OpenAI to implement score cutoffs rather than top-K retrieval to avoid unpredictably high token costs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Deep Dive into the Responses API and Migration Roadmap 6302 Swyx probes into the developer pain points around deprecating chat completions vs assistants API and points out potential abuse of state storage. Romain and Nikunj explain the architectural unification of the responses API and free state persistence.
Web Search Tool and GPT-4o Search Capabilities 7202 Alessio and Swyx demonstrate deep technical context comparing fine-tuned search vs tool use, citations, and parameter tuning like similarity thresholds versus top-K retrieval. Nikunj and Romain welcome the product suggestions.
Enhanced File Search and Managed RAG Infrastructure 6302 Swyx asks whether developers should DIY their RAG stack or rely on OpenAI's managed file search tool. Nikunj acknowledges that developers wanting full control over chunking and retrieval strategies should still build custom stacks.
Computer Use Agent Tooling and the Operator Model 5301 The hosts discuss the fine-tuned Computer Use model behind Operator and benchmark evals like playing Pokemon. Romain explains the multi-turn action loop while Nikunj confirms preview model unification roadmaps.
Agents SDK Architecture, Observability, and RFT Integration 7201 Swyx and Alessio explore Swarm evolving into the Agents SDK and propose connecting execution traces to reinforcement fine-tuning (RFT). Nikunj validates that tracing directly powers evals and RFT workflows.
Conclusion and Final Thoughts 1000 A quick wrap-up and mutual thank-yous as the hosts conclude the lightning episode.

Statements from this episode (14)

Disclosure
OpenAI releases its Operator computer use tool to API developers
“And then we're also launching our computer use tool. So this is the tool behind the operator product in ChatGPT. So that's coming to developers today.”
Nikunj Handa Mar 11, 2025 ▶ 1:08
Disclosure
OpenAI launches Responses API to power future agentic products
“And to support all of these tools, we're going to have a new API. So, you know, we launched ChatCompletions, like, I think March, 23 or so. It's been a while. So, so we're looking for an update over here to support all the new things that the models can do. An…”
Nikunj Handa Mar 11, 2025 ▶ 1:17
Disclosure
OpenAI launches Agents SDK with built-in dashboard tracing
“Actually, the last thing we're launching is the agent's SDK. We launched this thing called swarm last year, where You know, it was an experimental SDK for people to do multi-agent orchestration and stuff like that. It was supposed to be like educational, exper…”
Nikunj Handa Mar 11, 2025 ▶ 1:45
Disclosure
Huet: OpenAI will continue to maintain and support Chat Completions API
“Chat completion is definitely, like, here to stay. You know, it's a bare metal API we've had for quite some time, lots of tools built around it, so we want to make sure that it's maintained and people can confidently keep on building on it.”
Romain Huet Mar 11, 2025 ▶ 2:42
Assertion Supported
Swyx: OpenAI Assistants API target sunset is H1 2026
“And assistance API we've has a target sunset date of first half of 26.”
Shawn Wang Mar 11, 2025 ▶ 3:27
Disclosure
Handa: Responses API will support all Chat Completions and Assistants features
“So, so the responses API is going to support everything that the chat, it's at launch going to support everything that chat completion supports. And then over time, it's going to support everything that assistance supports.”
Nikunj Handa Mar 11, 2025 ▶ 5:23
Assertion Supported
Handa: OpenAI Responses API stores conversation state for 30 days free
“Yeah, it's free. We store your state for 30 days. You can turn it off. But yeah, it's free.”
Nikunj Handa Mar 11, 2025 ▶ 6:44
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”
Alessio Fanelli Mar 11, 2025 ▶ 8:12
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Nikunj Handa Mar 11, 2025 ▶ 9:26
Insight
Handa: Vector stores require metadata filtering above 5,000 records
“Metadata filtering was like the main thing people were asking us for a while, and that's the one I'm super excited about. I mean, it's just so critical. Once you're like vector store size goes over, you know, more than like, you know, five, 10,000 records, you…”
Nikunj Handa Mar 11, 2025 ▶ 16:15
Insight
OpenAI's Handa: Current computer use models are at the GPT-1 or GPT-2 stage
“The cool thing about computer use is that we're just so, so early. It's like the GPT-II of computer use or maybe GPT-I of computer use right now.”
Nikunj Handa Mar 11, 2025 ▶ 18:35
Disclosure
Handa: OpenAI plans to merge preview models into core mainline models
“I think in the early days, research teams that open AI like operate with like fine-tuned models. And then once the thing gets like more stable, we sort of merge it into the mainline. So that's definitely the vision, like going out of preview as we get more com…”
Nikunj Handa Mar 11, 2025 ▶ 20:50
Disclosure
OpenAI Agents SDK supports any Chat Completions-compatible API provider
“We also, like, made this pretty flexible, so you can pick any API from any provider that supports the chat completions API format. So it supports responses by default, but you can, like, easily plug it into Anyone that uses the Chat Completions API.”
Nikunj Handa Mar 11, 2025 ▶ 22:23
Disclosure
OpenAI plans to connect agent traces to evals and reinforcement fine-tuning
“Like you got to tie the traces to the evals product so that you can generate good evals. Once you have good evals and graders and tasks, You can use that to do reinforcement fine tuning and you know, lots of details to be figured out over here, but that's the …”
Nikunj Handa Mar 11, 2025 ▶ 24:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.