Oct 4, 2024 · 2h 9m · latent-space

Building AGI in Real Time (OpenAI Dev Day 2024)

Sam Altman · 21m spoken Kevin Weil · 15m spoken Shawn Wang · 13m spoken Alistair Pullen · 10m spoken Olivier Godement · 10m spoken Romain Huet · 6m spoken Ilan Biggio · 5m spoken Michelle Pokrass · 5m spoken Simon Willison · 4m spoken Alessio Fanelli · 4m spoken NotebookLM Host 2 · 3m spoken NotebookLM Host 1 · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

This comprehensive dispatch from OpenAI Dev Day 2024 explores OpenAI's major technical breakthroughs through in-depth interviews and live demonstrations. The coverage details the WebSocket-powered Realtime API, o1 reasoning models, model distillation, vision fine-tuning, and OpenAI's strategic roadmap toward autonomous agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 15.2% of the talking time here. How this is scored →

The hosts as informed peer 4.3 Guest teaching 2.7 Guest disagreement 0.6 The hosts pushing back 1.7
05100:0020:0040:001:00:001:20:001:40:002:00:001:23–9:09 · The hosts as informed peer 0/10 NotebookLM Deep Dive: OpenAI Dev Day 2024 Recap Synthetic AI recap segment generated by NotebookLM. Since this is an automated monologue/dialogue recap without the human podcast hosts, all host interaction scores are zeroed out.9:09–19:32 · The hosts as informed peer 5/10 Live Voice Demo: Ordering Strawberries with Realtime API Swyx asks detailed technical questions regarding session loops, Twilio integration, and function calling. Ilan educates him on how Realtime API handles speech-to-text without audio output for reliable UI command control.19:32–38:06 · The hosts as informed peer 6/10 Interview: Olivier Godement on OpenAI Platform and Product Strategy The hosts press Olivier on OpenAI becoming an AWS-style AI cloud, prompt caching lifetimes, and directly challenge OpenAI's aggressive refusal policies on voice models.38:06–48:36 · The hosts as informed peer 5/10 Interview: Romain Huet on Live Demos and Developer Tooling Collaborative discussion with Romain Huet covering the live demos, differences between o1-preview and o1-mini for coding, and client-side WebRTC integration.48:36–1:06:46 · The hosts as informed peer 6/10 Interview: Michelle Pokras and Simon Willison on API Architecture A technical peer exchange where Simon and Swyx discuss WebSocket proxies, OAuth API key security, and token accounting, while Michelle explains forthcoming chat completion audio features.1:06:46–1:22:01 · The hosts as informed peer 7/10 Interview: Alistair Pullen on Custom Reasoning and Model Fine-Tuning Swyx demonstrates deep domain knowledge of SWE-bench rankings, pointing out that Cosine's fine-tuned GPT-4o outperformed standard o1, validating Ali's choice to hide reasoning traces.1:22:01–2:09:08 · The hosts as informed peer 1/10 Keynote Fireside: Sam Altman and Kevin Weil on AGI, Agents, and Safety Dev Day fireside keynote where Kevin Weil and audience members question Sam Altman on AGI definitions, agent safety bottlenecks, voice singing restrictions, and alignment strategy.1:23–9:09 · Guest teaching 0/10 NotebookLM Deep Dive: OpenAI Dev Day 2024 Recap Synthetic AI recap segment generated by NotebookLM. Since this is an automated monologue/dialogue recap without the human podcast hosts, all host interaction scores are zeroed out.9:09–19:32 · Guest teaching 4/10 Live Voice Demo: Ordering Strawberries with Realtime API Swyx asks detailed technical questions regarding session loops, Twilio integration, and function calling. Ilan educates him on how Realtime API handles speech-to-text without audio output for reliable UI command control.19:32–38:06 · Guest teaching 3/10 Interview: Olivier Godement on OpenAI Platform and Product Strategy The hosts press Olivier on OpenAI becoming an AWS-style AI cloud, prompt caching lifetimes, and directly challenge OpenAI's aggressive refusal policies on voice models.38:06–48:36 · Guest teaching 2/10 Interview: Romain Huet on Live Demos and Developer Tooling Collaborative discussion with Romain Huet covering the live demos, differences between o1-preview and o1-mini for coding, and client-side WebRTC integration.48:36–1:06:46 · Guest teaching 3/10 Interview: Michelle Pokras and Simon Willison on API Architecture A technical peer exchange where Simon and Swyx discuss WebSocket proxies, OAuth API key security, and token accounting, while Michelle explains forthcoming chat completion audio features.1:06:46–1:22:01 · Guest teaching 3/10 Interview: Alistair Pullen on Custom Reasoning and Model Fine-Tuning Swyx demonstrates deep domain knowledge of SWE-bench rankings, pointing out that Cosine's fine-tuned GPT-4o outperformed standard o1, validating Ali's choice to hide reasoning traces.1:22:01–2:09:08 · Guest teaching 4/10 Keynote Fireside: Sam Altman and Kevin Weil on AGI, Agents, and Safety Dev Day fireside keynote where Kevin Weil and audience members question Sam Altman on AGI definitions, agent safety bottlenecks, voice singing restrictions, and alignment strategy.1:23–9:09 · Guest disagreement 0/10 NotebookLM Deep Dive: OpenAI Dev Day 2024 Recap Synthetic AI recap segment generated by NotebookLM. Since this is an automated monologue/dialogue recap without the human podcast hosts, all host interaction scores are zeroed out.9:09–19:32 · Guest disagreement 1/10 Live Voice Demo: Ordering Strawberries with Realtime API Swyx asks detailed technical questions regarding session loops, Twilio integration, and function calling. Ilan educates him on how Realtime API handles speech-to-text without audio output for reliable UI command control.19:32–38:06 · Guest disagreement 1/10 Interview: Olivier Godement on OpenAI Platform and Product Strategy The hosts press Olivier on OpenAI becoming an AWS-style AI cloud, prompt caching lifetimes, and directly challenge OpenAI's aggressive refusal policies on voice models.38:06–48:36 · Guest disagreement 0/10 Interview: Romain Huet on Live Demos and Developer Tooling Collaborative discussion with Romain Huet covering the live demos, differences between o1-preview and o1-mini for coding, and client-side WebRTC integration.48:36–1:06:46 · Guest disagreement 0/10 Interview: Michelle Pokras and Simon Willison on API Architecture A technical peer exchange where Simon and Swyx discuss WebSocket proxies, OAuth API key security, and token accounting, while Michelle explains forthcoming chat completion audio features.1:06:46–1:22:01 · Guest disagreement 1/10 Interview: Alistair Pullen on Custom Reasoning and Model Fine-Tuning Swyx demonstrates deep domain knowledge of SWE-bench rankings, pointing out that Cosine's fine-tuned GPT-4o outperformed standard o1, validating Ali's choice to hide reasoning traces.1:22:01–2:09:08 · Guest disagreement 1/10 Keynote Fireside: Sam Altman and Kevin Weil on AGI, Agents, and Safety Dev Day fireside keynote where Kevin Weil and audience members question Sam Altman on AGI definitions, agent safety bottlenecks, voice singing restrictions, and alignment strategy.1:23–9:09 · The hosts pushing back 0/10 NotebookLM Deep Dive: OpenAI Dev Day 2024 Recap Synthetic AI recap segment generated by NotebookLM. Since this is an automated monologue/dialogue recap without the human podcast hosts, all host interaction scores are zeroed out.9:09–19:32 · The hosts pushing back 1/10 Live Voice Demo: Ordering Strawberries with Realtime API Swyx asks detailed technical questions regarding session loops, Twilio integration, and function calling. Ilan educates him on how Realtime API handles speech-to-text without audio output for reliable UI command control.19:32–38:06 · The hosts pushing back 4/10 Interview: Olivier Godement on OpenAI Platform and Product Strategy The hosts press Olivier on OpenAI becoming an AWS-style AI cloud, prompt caching lifetimes, and directly challenge OpenAI's aggressive refusal policies on voice models.38:06–48:36 · The hosts pushing back 1/10 Interview: Romain Huet on Live Demos and Developer Tooling Collaborative discussion with Romain Huet covering the live demos, differences between o1-preview and o1-mini for coding, and client-side WebRTC integration.48:36–1:06:46 · The hosts pushing back 2/10 Interview: Michelle Pokras and Simon Willison on API Architecture A technical peer exchange where Simon and Swyx discuss WebSocket proxies, OAuth API key security, and token accounting, while Michelle explains forthcoming chat completion audio features.1:06:46–1:22:01 · The hosts pushing back 2/10 Interview: Alistair Pullen on Custom Reasoning and Model Fine-Tuning Swyx demonstrates deep domain knowledge of SWE-bench rankings, pointing out that Cosine's fine-tuned GPT-4o outperformed standard o1, validating Ali's choice to hide reasoning traces.1:22:01–2:09:08 · The hosts pushing back 2/10 Keynote Fireside: Sam Altman and Kevin Weil on AGI, Agents, and Safety Dev Day fireside keynote where Kevin Weil and audience members question Sam Altman on AGI definitions, agent safety bottlenecks, voice singing restrictions, and alignment strategy.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 14.4% · guest 85.6%9:00 · the hosts 14.4% · guest 85.6%12:00 · the hosts 29.9% · guest 70.1%12:00 · the hosts 29.9% · guest 70.1%15:00 · the hosts 26.5% · guest 73.5%15:00 · the hosts 26.5% · guest 73.5%18:00 · the hosts 34.6% · guest 65.4%18:00 · the hosts 34.6% · guest 65.4%21:00 · the hosts 41.2% · guest 58.8%21:00 · the hosts 41.2% · guest 58.8%24:00 · the hosts 36.4% · guest 63.6%24:00 · the hosts 36.4% · guest 63.6%27:00 · the hosts 31.4% · guest 68.6%27:00 · the hosts 31.4% · guest 68.6%30:00 · the hosts 40% · guest 60%30:00 · the hosts 40% · guest 60%33:00 · the hosts 26.7% · guest 73.3%33:00 · the hosts 26.7% · guest 73.3%36:00 · the hosts 21.7% · guest 78.3%36:00 · the hosts 21.7% · guest 78.3%39:00 · the hosts 35.5% · guest 64.5%39:00 · the hosts 35.5% · guest 64.5%42:00 · the hosts 23% · guest 77%42:00 · the hosts 23% · guest 77%45:00 · the hosts 18.1% · guest 81.9%45:00 · the hosts 18.1% · guest 81.9%48:00 · the hosts 30.6% · guest 69.4%48:00 · the hosts 30.6% · guest 69.4%51:00 · the hosts 22.8% · guest 77.2%51:00 · the hosts 22.8% · guest 77.2%54:00 · the hosts 22.6% · guest 77.4%54:00 · the hosts 22.6% · guest 77.4%57:00 · the hosts 37.7% · guest 62.3%57:00 · the hosts 37.7% · guest 62.3%1:00:00 · the hosts 21.1% · guest 78.9%1:00:00 · the hosts 21.1% · guest 78.9%1:03:00 · the hosts 29.3% · guest 70.7%1:03:00 · the hosts 29.3% · guest 70.7%1:06:00 · the hosts 19.5% · guest 80.5%1:06:00 · the hosts 19.5% · guest 80.5%1:09:00 · the hosts 21.4% · guest 78.6%1:09:00 · the hosts 21.4% · guest 78.6%1:12:00 · the hosts 37.1% · guest 62.9%1:12:00 · the hosts 37.1% · guest 62.9%1:15:00 · the hosts 12% · guest 88%1:15:00 · the hosts 12% · guest 88%1:18:00 · the hosts 12.4% · guest 87.6%1:18:00 · the hosts 12.4% · guest 87.6%1:21:00 · the hosts 4.7% · guest 95.3%1:21:00 · the hosts 4.7% · guest 95.3%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:57:00 · the hosts 0% · guest 100%1:57:00 · the hosts 0% · guest 100%2:00:00 · the hosts 0% · guest 100%2:00:00 · the hosts 0% · guest 100%2:03:00 · the hosts 0% · guest 100%2:03:00 · the hosts 0% · guest 100%2:06:00 · the hosts 0% · guest 100%2:06:00 · the hosts 0% · guest 100%2:09:00 · the hosts 0% · guest 100%2:09:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:42:44 Sam Altman defends conservative safety defaults against user pushback

Sam acknowledges user frustration regarding over-refusal in voice mode, firmly defending OpenAI's stance to launch conservatively to understand real-world harms.

Hardest push from the hosts ▶ 34:23 Swyx challenges Olivier on strict voice moderation and refusals

Swyx directly confronts Olivier on why voice mode over-refuses and questions whether developers will get a controllable safety knob.

Biggest teaching moment ▶ 16:43 Ilan explains Realtime API's headless text and UI command capabilities

Ilan educates Swyx on how Realtime API can operate without audio outputs to create ultra-reliable speech-to-function UI command architectures.

The host holds their own ▶ 1:12:18 Swyx points out Genie's SWE-bench score superior to o1

Swyx demonstrates technical command by pointing out that Cosine's fine-tuned GPT-4o beat OpenAI's new o1 reasoning model on SWE-bench Verified.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
NotebookLM Deep Dive: OpenAI Dev Day 2024 Recap 0000 Synthetic AI recap segment generated by NotebookLM. Since this is an automated monologue/dialogue recap without the human podcast hosts, all host interaction scores are zeroed out.
Live Voice Demo: Ordering Strawberries with Realtime API 5411 Swyx asks detailed technical questions regarding session loops, Twilio integration, and function calling. Ilan educates him on how Realtime API handles speech-to-text without audio output for reliable UI command control.
Interview: Olivier Godement on OpenAI Platform and Product Strategy 6314 The hosts press Olivier on OpenAI becoming an AWS-style AI cloud, prompt caching lifetimes, and directly challenge OpenAI's aggressive refusal policies on voice models.
Interview: Romain Huet on Live Demos and Developer Tooling 5201 Collaborative discussion with Romain Huet covering the live demos, differences between o1-preview and o1-mini for coding, and client-side WebRTC integration.
Interview: Michelle Pokras and Simon Willison on API Architecture 6302 A technical peer exchange where Simon and Swyx discuss WebSocket proxies, OAuth API key security, and token accounting, while Michelle explains forthcoming chat completion audio features.
Interview: Alistair Pullen on Custom Reasoning and Model Fine-Tuning 7312 Swyx demonstrates deep domain knowledge of SWE-bench rankings, pointing out that Cosine's fine-tuned GPT-4o outperformed standard o1, validating Ali's choice to hide reasoning traces.
Keynote Fireside: Sam Altman and Kevin Weil on AGI, Agents, and Safety 1412 Dev Day fireside keynote where Kevin Weil and audience members question Sam Altman on AGI definitions, agent safety bottlenecks, voice singing restrictions, and alignment strategy.

Statements from this episode (23)

Prediction Held up
Biggio: Real-time video/vision is likely the next Realtime API feature
“To use ChatGPT's voice mode as an example, like we've demoed the video, right? Like real-time image, right? So I'm not actually sure what timelines are, but I would expect, if I had to guess, that like that is probably the next thing that we're going to be mak…”
Ilan Biggio Oct 4, 2024 ▶ 16:15
Insight
Biggio: Forcing function calls over voice yields reliable UI control
“Speech to text is really interesting because you can prevent, you can prevent responses like audio responses and force function calls. And so you can do stuff like UI control that is like super, super reliable... If you like cut out the audio outputs and make …”
Ilan Biggio Oct 4, 2024 ▶ 16:56
Assertion Not checkable as stated
Godement: Vision fine-tuning shows higher performance uplift than text
“We've been alpha testing, like, the vision fine tuning, like, for several weeks at that point. We are seeing, like, even higher performance uplift compared to text fine tuning.”
Olivier Godement Oct 4, 2024 ▶ 25:34
Prediction Didn’t hold up
Godement: Developers will rely on continuous, automated fine-tuning within years
“The vision we have is, fast forward a couple of years, I think, like, most developers will essentially, like, have an automated, continuous, fine-tuned model. The more, like, you use the model, the more data you pass to the mobile provider, like, the model is …”
Olivier Godement Oct 4, 2024 ▶ 26:46
Disclosure
Godement: OpenAI will only build tools closest to the model
“There is no freaking way that OpenAI can build everything. Like, there is just too much to build, frankly. And so, my philosophy is, essentially, we'll focus on, like, the tools which are, like, the closest to the model itself. So that's why you see us, like, …”
Olivier Godement Oct 4, 2024 ▶ 31:05
Disclosure
Godement: OpenAI plans API knobs for developers to adjust safety thresholds
“And so I think the direction where we'll go here is that, basically, there will always be, like, you know, a set of behavior that will, you know, just, like, forbid, frankly, because they're illegal against our terms of services, but then there will be, like, …”
Olivier Godement Oct 4, 2024 ▶ 34:59
Insight
Godement: Founders should build apps slightly too difficult for current models
“As a developer, as a founder, you basically want to build an app, which is a bit too difficult for the model today, right? Like what you think is right. It's like sort of working, sometimes not working. And that way, you know, that basically gives us like a go…”
Olivier Godement Oct 4, 2024 ▶ 37:14
Assertion Supported
Huet: Over 3 million developers build on OpenAI
“I'm sure we talked about this before, but there's now more than three million developers building on OpenAI, so it's pretty exciting to see all of that energy into creating new things.”
Romain Huet Oct 4, 2024 ▶ 39:27
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Shawn Wang Oct 4, 2024 ▶ 48:14
Opinion
Pokras: Vision fine-tuning is the most underrated release for bespoke OCR
“Vision fine-tuning is so underrated. For the past, like, two months, whenever I talk to founders, they tell me this is the thing they need most. A lot of people are doing, like, OCR on, on very bespoke formats, like government documents, and vision fine-tuning…”
Michelle Pokrass Oct 4, 2024 ▶ 56:21
Disclosure
Pokras: OpenAI will ship raw audio in Chat Completions API
“We're actually going to be shipping audio capabilities in chat completions. So this is like the lowest level capability. So you supply in audio and you can get back raw audio and it works at the request response layer.”
Michelle Pokrass Oct 4, 2024 ▶ 59:41
Opinion
Pokras: OpenAI Assistants API requires too many initial API requests
“Some of the things that are good in the assistance API is hosted tools. People really like posted tools and especially RAG. And then some things that are, you know, less intuitive is just how many API requests you need to get going with the Assistant's API.”
Michelle Pokrass Oct 4, 2024 ▶ 1:05:05
Disclosure
Pullen: Replacing Genie's reasoning traces with o1 traces improves performance
“Even now we've started Replacing some of the reasoning traces in our Genie model with reasoning traces generated by O-one, or at least in tandem with O-one, and we've already started seeing improvements in performance from that point.”
Alistair Pullen Oct 4, 2024 ▶ 1:09:59
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04
Opinion
Pullen: SWE-bench is a poor proxy for real-world AI coding competence
“I know Sweebench is, like, the most commonly talked about thing, and honestly, it's a very, it's an amazing project, but one of the things we've learned the most from actually shipping this product to users is, it's a pretty bad proxy at telling us how compete…”
Alistair Pullen Oct 4, 2024 ▶ 1:16:10
Assertion Not checkable as stated
Altman: OpenAI reached Level 2 AGI with o1
“I think we clearly got to level two, or we clearly got to level two with O-one.”
Sam Altman Oct 4, 2024 ▶ 1:24:50
Assertion Not checkable as stated
Altman: o1 is OpenAI's most aligned model ever by a lot
“And O-one is obviously our most capable model ever, but it's also our most aligned model ever by a lot.”
Sam Altman Oct 4, 2024 ▶ 1:33:35
Disclosure
Weil: OpenAI o1 will support function calling by end of 2024
“I'm really excited to see things like system prompts, And structured outputs, and function calling, make it into a one, we will be there by the end of the year.”
Kevin Weil Oct 4, 2024 ▶ 1:48:29
Assertion Supported
Weil: ChatGPT supports over 200 million weekly active users
“As we, you know, we support over two hundred million people every week on ChatGPT.”
Kevin Weil Oct 4, 2024 ▶ 1:51:49
Assertion Not checkable as stated
Weil: OpenAI customer support team is 20% expected size thanks to AI
“There are things that get closer to that, I mean, there, like, customer service, we have bots internally that do what's fun about answering external questions and fielding internal people's questions on Slack and so on, and our customer success, our customer s…”
Kevin Weil Oct 4, 2024 ▶ 1:57:42
Prediction Not checkable as stated
Altman: Infinite AI context windows will happen within a decade
“That obviously takes some research breakthroughs, but I assume that infinite context will happen at some point. At some point, it's like, less than a decade.”
Sam Altman Oct 4, 2024 ▶ 2:05:18
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Sam Altman Oct 4, 2024 ▶ 2:05:30
Prediction Not checkable as stated
Altman: AI will dynamically render custom real-time interfaces for any request
“At some point in not that many years in the future, you'll walk up to a piece of glass, you will say whatever you want they will have, like, there will be incredible reasoning models, agents connected to everything, there will be a video model streaming back t…”
Sam Altman Oct 4, 2024 ▶ 2:08:17
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.