Apr 25, 2024 · 31m · no-priors

No Priors Ep.61 | OpenAI's Sora Leaders Aditya Ramesh, Tim Brooks and Bill Peebles

Tim Brooks · 9m spoken Bill Peebles · 7m spoken Sarah Guo · 4m spoken Aditya Ramesh · 3m spoken Elad Gil · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, OpenAI's Sora leadership team—Aditya Ramesh, Tim Brooks, and Bill Peebles—discusses the diffusion transformer architecture, emergent physics simulation, safety guardrails, and creative potential of their generative video model. They explain how scaling video generation marks a foundational milestone on the direct path toward Artificial General Intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.7% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 5.7 Guest disagreement 0.3 The hosts pushing back 1.2
05100:0010:0020:0030:002:14–7:08 · The hosts as informed peer 4/10 Safety Roadmap, Creator Feedback, and Emerging Creative Workflows Elad draws thoughtful historical parallels to Pixar's early shorts and computer graphics evolution. The guests explain their deliberate rollout strategy, red teaming, and early artist feedback like the Airhead short and Bling Zoo.7:08–10:55 · The hosts as informed peer 5/10 Video as World Models and Diffusion Transformer Architecture Sarah prompts Tim to break down diffusion transformers and asks Bill about empirical scaling laws. Tim clearly articulates how iteratively removing noise combines with transformer compute scaling.10:56–15:26 · The hosts as informed peer 5/10 Latent Space-Time Patches and First-Principles Architecture Sarah and Elad engage deeply on tokenization paradigms and compare Sora's design to end-to-end deep learning in self-driving cars. Bill and Tim explain how 3D spacetime cubes broke away from legacy fixed-crop image extension hacks.15:27–20:06 · The hosts as informed peer 5/10 Aesthetic Steering, Model Personalization, and Education Sarah shares her personal experience synthesizing stories for her kids and Elad asks about avatar applications. Aditya notes that Sora's aesthetic is emergent and prompt-steered rather than hand-tuned.20:07–25:02 · The hosts as informed peer 5/10 Safety Challenges, Misinformation, and Physical Coherence Limitations Elad pushes on the liability question using Photoshop as a precedent for user responsibility. Bill transparently outlines current physical coherence failures, such as soccer balls vaporizing during interactions.25:03–31:21 · The hosts as informed peer 5/10 The Bitter Lesson, Scaling Laws, and Video as the GPT-1 Moment Sarah asks about approximate human physics versus exact simulation, and Tim and Bill explain how the Bitter Lesson dictates scaling raw data prediction over engineered approximations.2:14–7:08 · Guest teaching 4/10 Safety Roadmap, Creator Feedback, and Emerging Creative Workflows Elad draws thoughtful historical parallels to Pixar's early shorts and computer graphics evolution. The guests explain their deliberate rollout strategy, red teaming, and early artist feedback like the Airhead short and Bling Zoo.7:08–10:55 · Guest teaching 6/10 Video as World Models and Diffusion Transformer Architecture Sarah prompts Tim to break down diffusion transformers and asks Bill about empirical scaling laws. Tim clearly articulates how iteratively removing noise combines with transformer compute scaling.10:56–15:26 · Guest teaching 6/10 Latent Space-Time Patches and First-Principles Architecture Sarah and Elad engage deeply on tokenization paradigms and compare Sora's design to end-to-end deep learning in self-driving cars. Bill and Tim explain how 3D spacetime cubes broke away from legacy fixed-crop image extension hacks.15:27–20:06 · Guest teaching 5/10 Aesthetic Steering, Model Personalization, and Education Sarah shares her personal experience synthesizing stories for her kids and Elad asks about avatar applications. Aditya notes that Sora's aesthetic is emergent and prompt-steered rather than hand-tuned.20:07–25:02 · Guest teaching 6/10 Safety Challenges, Misinformation, and Physical Coherence Limitations Elad pushes on the liability question using Photoshop as a precedent for user responsibility. Bill transparently outlines current physical coherence failures, such as soccer balls vaporizing during interactions.25:03–31:21 · Guest teaching 7/10 The Bitter Lesson, Scaling Laws, and Video as the GPT-1 Moment Sarah asks about approximate human physics versus exact simulation, and Tim and Bill explain how the Bitter Lesson dictates scaling raw data prediction over engineered approximations.2:14–7:08 · Guest disagreement 0/10 Safety Roadmap, Creator Feedback, and Emerging Creative Workflows Elad draws thoughtful historical parallels to Pixar's early shorts and computer graphics evolution. The guests explain their deliberate rollout strategy, red teaming, and early artist feedback like the Airhead short and Bling Zoo.7:08–10:55 · Guest disagreement 0/10 Video as World Models and Diffusion Transformer Architecture Sarah prompts Tim to break down diffusion transformers and asks Bill about empirical scaling laws. Tim clearly articulates how iteratively removing noise combines with transformer compute scaling.10:56–15:26 · Guest disagreement 0/10 Latent Space-Time Patches and First-Principles Architecture Sarah and Elad engage deeply on tokenization paradigms and compare Sora's design to end-to-end deep learning in self-driving cars. Bill and Tim explain how 3D spacetime cubes broke away from legacy fixed-crop image extension hacks.15:27–20:06 · Guest disagreement 0/10 Aesthetic Steering, Model Personalization, and Education Sarah shares her personal experience synthesizing stories for her kids and Elad asks about avatar applications. Aditya notes that Sora's aesthetic is emergent and prompt-steered rather than hand-tuned.20:07–25:02 · Guest disagreement 1/10 Safety Challenges, Misinformation, and Physical Coherence Limitations Elad pushes on the liability question using Photoshop as a precedent for user responsibility. Bill transparently outlines current physical coherence failures, such as soccer balls vaporizing during interactions.25:03–31:21 · Guest disagreement 1/10 The Bitter Lesson, Scaling Laws, and Video as the GPT-1 Moment Sarah asks about approximate human physics versus exact simulation, and Tim and Bill explain how the Bitter Lesson dictates scaling raw data prediction over engineered approximations.2:14–7:08 · The hosts pushing back 1/10 Safety Roadmap, Creator Feedback, and Emerging Creative Workflows Elad draws thoughtful historical parallels to Pixar's early shorts and computer graphics evolution. The guests explain their deliberate rollout strategy, red teaming, and early artist feedback like the Airhead short and Bling Zoo.7:08–10:55 · The hosts pushing back 1/10 Video as World Models and Diffusion Transformer Architecture Sarah prompts Tim to break down diffusion transformers and asks Bill about empirical scaling laws. Tim clearly articulates how iteratively removing noise combines with transformer compute scaling.10:56–15:26 · The hosts pushing back 1/10 Latent Space-Time Patches and First-Principles Architecture Sarah and Elad engage deeply on tokenization paradigms and compare Sora's design to end-to-end deep learning in self-driving cars. Bill and Tim explain how 3D spacetime cubes broke away from legacy fixed-crop image extension hacks.15:27–20:06 · The hosts pushing back 1/10 Aesthetic Steering, Model Personalization, and Education Sarah shares her personal experience synthesizing stories for her kids and Elad asks about avatar applications. Aditya notes that Sora's aesthetic is emergent and prompt-steered rather than hand-tuned.20:07–25:02 · The hosts pushing back 2/10 Safety Challenges, Misinformation, and Physical Coherence Limitations Elad pushes on the liability question using Photoshop as a precedent for user responsibility. Bill transparently outlines current physical coherence failures, such as soccer balls vaporizing during interactions.25:03–31:21 · The hosts pushing back 1/10 The Bitter Lesson, Scaling Laws, and Video as the GPT-1 Moment Sarah asks about approximate human physics versus exact simulation, and Tim and Bill explain how the Bitter Lesson dictates scaling raw data prediction over engineered approximations.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 36.8% · guest 63.2%0:00 · the hosts 36.8% · guest 63.2%3:00 · the hosts 21.1% · guest 78.9%3:00 · the hosts 21.1% · guest 78.9%6:00 · the hosts 39.5% · guest 60.5%6:00 · the hosts 39.5% · guest 60.5%9:00 · the hosts 18.1% · guest 81.9%9:00 · the hosts 18.1% · guest 81.9%12:00 · the hosts 6.5% · guest 93.5%12:00 · the hosts 6.5% · guest 93.5%15:00 · the hosts 54.3% · guest 45.7%15:00 · the hosts 54.3% · guest 45.7%18:00 · the hosts 26.1% · guest 73.9%18:00 · the hosts 26.1% · guest 73.9%21:00 · the hosts 29.9% · guest 70.1%21:00 · the hosts 29.9% · guest 70.1%24:00 · the hosts 16.5% · guest 83.5%24:00 · the hosts 16.5% · guest 83.5%27:00 · the hosts 27.5% · guest 72.5%27:00 · the hosts 27.5% · guest 72.5%30:00 · the hosts 30.1% · guest 69.9%30:00 · the hosts 30.1% · guest 69.9%
Sharpest disagreement ▶ 27:25 Bill reframes human physics intuition as a deficiency

Bill politely rejects the premise that models should merely match approximate human physics intuition, arguing that human fidelity is deficient and will be superseded by scaled neural world models.

Hardest push from the hosts ▶ 21:18 Elad pushes back on AI platform liability using Photoshop precedent

Elad challenges the notion of developer liability for misuse by citing historical precedents like Photoshop, where toolmakers are not held responsible for malicious user edits.

Biggest teaching moment ▶ 11:32 Bill details the shift from fixed-resolution crops to 3D spacetime tokens

Bill systematically educates the hosts on how previous visual models threw away diverse internet data via rigid crops, whereas 3D spacetime patches unlock broad LLM-style generalist capabilities.

The host holds their own ▶ 15:14 Elad draws technical analogy to end-to-end deep learning in AVs

Elad demonstrates sharp domain insight by connecting Tim's first-principles architecture explanation to the recent paradigm shift from heuristic systems to end-to-end learning in autonomous driving.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Safety Roadmap, Creator Feedback, and Emerging Creative Workflows 4401 Elad draws thoughtful historical parallels to Pixar's early shorts and computer graphics evolution. The guests explain their deliberate rollout strategy, red teaming, and early artist feedback like the Airhead short and Bling Zoo.
Video as World Models and Diffusion Transformer Architecture 5601 Sarah prompts Tim to break down diffusion transformers and asks Bill about empirical scaling laws. Tim clearly articulates how iteratively removing noise combines with transformer compute scaling.
Latent Space-Time Patches and First-Principles Architecture 5601 Sarah and Elad engage deeply on tokenization paradigms and compare Sora's design to end-to-end deep learning in self-driving cars. Bill and Tim explain how 3D spacetime cubes broke away from legacy fixed-crop image extension hacks.
Aesthetic Steering, Model Personalization, and Education 5501 Sarah shares her personal experience synthesizing stories for her kids and Elad asks about avatar applications. Aditya notes that Sora's aesthetic is emergent and prompt-steered rather than hand-tuned.
Safety Challenges, Misinformation, and Physical Coherence Limitations 5612 Elad pushes on the liability question using Photoshop as a precedent for user responsibility. Bill transparently outlines current physical coherence failures, such as soccer balls vaporizing during interactions.
The Bitter Lesson, Scaling Laws, and Video as the GPT-1 Moment 5711 Sarah asks about approximate human physics versus exact simulation, and Tim and Bill explain how the Bitter Lesson dictates scaling raw data prediction over engineered approximations.

Statements from this episode (14)

Opinion
Peebles: Generative video models are on the critical path to AGI
“Yeah, we absolutely believe models like Sora are really on the critical pathway to AGI.”
Bill Peebles Apr 25, 2024 ▶ 1:07
Prediction Not checkable as stated
Peebles: Sora will evolve into interactive world simulators with autonomous humans
“And so looking forward, as we continue to scale up models like Sora, we think we're going to be able to build these, like, world simulators, where essentially, you know, anybody can interact with them. I, as a human, can have my own simulator running, and I ca…”
Bill Peebles Apr 25, 2024 ▶ 1:52
Disclosure
Brooks: OpenAI has no immediate timeline to launch a Sora product
“We don't currently have immediate plans or even a timeline for creating a product.”
Tim Brooks Apr 25, 2024 ▶ 2:37
Prediction Not checkable as stated
Brooks: Generative video will birth entirely new media formats within years
“I do think that, yeah, maybe over the next couple years we'll see people starting to make, like, more and more films, but I think people will also find completely new ways to use these models that are just different from the current media that we're used to.”
Tim Brooks Apr 25, 2024 ▶ 6:34
Prediction Not checkable as stated
Peebles: Training on raw video is essential for future embodied robotics
“So you learn so much about the physical world just from training on raw video that we really believe that it's going to be essential for things like physical embodiment moving forward.”
Bill Peebles Apr 25, 2024 ▶ 8:20
Assertion Not checkable as stated
Peebles: Sora is the first visual model with LLM-like breadth
“And so this is really the first generative model of visual content that has breadth in a way that language models have breadth.”
Bill Peebles Apr 25, 2024 ▶ 12:59
Disclosure
Brooks: OpenAI built Sora from scratch to achieve minute-long HD video
“We started from scratch, and we started with the question of, how are we going to do a minute of HD footage? And that was our goal. And when you have that goal, we knew that we couldn't just extend an image generator. We knew that in order to do a minute of HD…”
Tim Brooks Apr 25, 2024 ▶ 13:58
Insight
Brooks: AI teams should build for three-year end-states, not incremental features
“There is this pressure to do things fast because AI is so fast, and the fastest thing to do is, oh, let's take what's working now and let's kind of, like, add on something to it, and that probably is, as you're saying, more general than just image to video but…”
Tim Brooks Apr 25, 2024 ▶ 14:43
Disclosure
Ramesh: OpenAI did not spend much effort tuning Sora's visual aesthetic
“Well, to be honest, we didn't spend a ton of effort on it for Sora.”
Aditya Ramesh Apr 25, 2024 ▶ 15:52
Opinion
Brooks: Sora represents the 'GPT-1' stage of generative visual models
“I think where we are in the trajectory of Sora right now is like, this is the GPT-I of these, this new paradigm of visual models, and that we're really looking at the fundamental research into making these way better, making it a way better engine that can pow…”
Tim Brooks Apr 25, 2024 ▶ 19:42
Assertion Not checkable as stated
Peebles: Generating long videos with Sora currently takes several minutes
“It's not instant and you have to wait at least like a few minutes for like these really long videos that we're generating.”
Bill Peebles Apr 25, 2024 ▶ 23:15
Assertion Not checkable as stated
Brooks: Sora learned 3D spatial understanding purely from raw 2D video
“We didn't explicitly bake three D information into it whatsoever. We just trained it on video data and it learned about three D because three D exists in those videos.”
Tim Brooks Apr 25, 2024 ▶ 25:24
Prediction Not checkable as stated
Peebles: Video models will eventually surpass humans as physical world models
“We're optimistic that Sora will, you know, supersede that kind of capability and will, you know, in the long run enable it to be More intelligent one day than humans as world models.”
Bill Peebles Apr 25, 2024 ▶ 27:46
Insight
Brooks: The best way to scale intelligence is simple data prediction
“What works really well as you increase scale is just predict data. And that's what we do with text. We just predict text. And that's exactly what we're doing with visual data with Sora, which is we're not making some complicated trying to figure out some new t…”
Tim Brooks Apr 25, 2024 ▶ 28:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.