Jiaming Song

Former Chief Scientist, Luma AI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistacademicexecutiveLinkedIn ↗tsong.me ↗

Jiaming Song is an artificial intelligence researcher best known as a co-creator of Denoising Diffusion Implicit Models (DDIM). He previously led foundation model research as Chief Scientist at Luma AI, directing work on systems including the Dream Machine video model.

16statements → 9claims → 3claims resolved → 67%fully supported → 3.69/5average certainty → 2.31/5average debate potential →

2 supported 1 partly supported 0 contradicted 6 not checkable as stated how the 9 claims stand · each chip opens the sources

1 prediction · 8 assertions · 6 insights · 1 disclosure · every statement was checked. The prediction and assertions are the 9 claims: statements the public record can support or contradict. 3 are resolved, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Jiaming argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Song: Dream Machine Simulates Complex Metallic Reflections From a Single Photo
“With Dream Machine, we just took another coffee machine that we have in the office, and then just do image to video, and we can actually see a lot of the light reflections On top of the metal surface of the coffee machine being simulated by the model.”
Jiaming Song Sep 9, 2024 ▶ 17:35 Luma's Dream Machine and Reasoning in Video Models

How they sound: speaking style how? →

251 words/min while actually speaking · 28 um and uh per 1k words

No argument clarity score for Jiaming Song: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 4,252 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Jiaming Song said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Song: Compute Scale Replaces Years of Explicit Graphics and Physics Engineering
“What's really surprising to me is how a large scale of compute is Mostly all you need to capture a lot of the, like, intricate effects that people spend years to develop in, like, graphics and, you know, physics simulation community”
Jiaming Song Sep 9, 2024 ▶ 5:21 Luma's Dream Machine and Reasoning in Video Models
Disclosure
Song: Dream Machine Learned Physical Simulation and Lighting Without 3D Priors
“But what's really surprising about Dream Machine is that we almost did nothing with regards to these NERF datasets or have zero, three D priors into how the model works. But the model just by itself learns to kind of uncover these interesting physical aspects …”
Jiaming Song Sep 9, 2024 ▶ 12:28 Luma's Dream Machine and Reasoning in Video Models
Assertion Partly supported
Song: Dream Machine Spontaneously Generates Consistent Cinematic Cuts Without Prompting
“Even though we haven't really asked the model to do anything about it, the model is able to reason about the second shot being a cut of the first shot that happens to have the same things in the first shot. So this is curse definitely some kind of a non-trivia…”
Jiaming Song Sep 9, 2024 ▶ 21:41 Luma's Dream Machine and Reasoning in Video Models
Insight
Song: Strict Physics Simulation Cannot Solve Fictional or Artistic Video Generation
“If we were trying to reason about the world via the traditional kind of physics techniques, it is very difficult to kind of reason about things inside a totally fictional world. So I think this is something that more happens like inside our dreams rather than …”
Jiaming Song Sep 9, 2024 ▶ 25:42 Luma's Dream Machine and Reasoning in Video Models
Insight
Song: Video AI Models Learn Psychological Causality Beyond Basic Physics
“And this is possibly caused by this eye being very unnaturally looking, and this is some kind of a cause and effect that is, like, even harder to reason, like, strictly in the physics, but more delving into how just human psychology works. So I think the causa…”
Jiaming Song Sep 9, 2024 ▶ 22:41 Luma's Dream Machine and Reasoning in Video Models
Prediction Not checkable as stated
Song: AI Video Models Will Evolve Into 4D Multi-Angle World Simulators
“I don't think it's a much of a stretch to say maybe we can get from like videos to four D. So that's being able to do world simulators meaning, meaning that you might be able to simulate multiple angles at the same time.”
Jiaming Song Sep 9, 2024 ▶ 27:32 Luma's Dream Machine and Reasoning in Video Models
Insight
Song: 3D AI Models Face Severe Data Scalability Bottleneck Compared to 2D
“Three D data has this scalability issue, and you If you compare that with like images, everyone can, you know, use their cell phone to take a photo or, you know, take a video. Whereas if you try to do the same with three D is very difficult. You either have to…”
Jiaming Song Sep 9, 2024 ▶ 1:55 Luma's Dream Machine and Reasoning in Video Models
Insight
Song: 2D Image Models Cannot Reason About Camera Physics or Motion
“The limitations of images was that it wasn't able to reason about how You know, the camera works in the world because it only has like a relatively independent shots of different objects.”
Jiaming Song Sep 9, 2024 ▶ 4:24 Luma's Dream Machine and Reasoning in Video Models
Assertion Not checkable as stated
Song: Video Models Unexpectedly Learn 3D Spatial Reasoning From Raw Video
“It turns out from these type of videos, we are able to show that a video model is able to reason about three D quite well, which in some sense is unexpected before in, in the community.”
Jiaming Song Sep 9, 2024 ▶ 5:04 Luma's Dream Machine and Reasoning in Video Models
Assertion Not checkable as stated
Song: Dream Machine Video Outputs Enable Consistent 3D Scene Reconstruction
“For example, in this case, what we do is we literally took one of the videos from the last side and we put this video into our three D reconstruction pipeline. And it turns out that it is able to reconstruct a three D scene at this direction quite reasonably w…”
Jiaming Song Sep 9, 2024 ▶ 6:01 Luma's Dream Machine and Reasoning in Video Models
Assertion Not checkable as stated
Song: Dream Machine Surpasses Multi-View Image Models in Detail and Resolution
“So I think that tells me that Dream Machine is definitely able to reason about better than any of the models that we've, you know, worked with before, and this is very much unlike, like, even the models that you try to obtain by fine-tuning on the, you know, l…”
Jiaming Song Sep 9, 2024 ▶ 7:01 Luma's Dream Machine and Reasoning in Video Models
Insight
Song: NeRF and Gaussian Splatting Suffer Major Real-World Capture Limitations
“When you try to develop, like, deploy these techniques in, in, in the wild, there are many issues that comes with this imperfect capture, like, capturing system that comes along. Like, people, when they're trying to capture an object, will oftentimes not captu…”
Jiaming Song Sep 9, 2024 ▶ 8:00 Luma's Dream Machine and Reasoning in Video Models
Assertion Not checkable as stated
Song: Dream Machine Synthesizes 3D-Consistent Video From Single NeRF Frames
“We put the first frame of an image into Dream Machine, and then Dream Machine will give us the output as a video. So as you can see here, like, the three-d consistency of the generated video looks quite amazing”
Jiaming Song Sep 9, 2024 ▶ 9:02 Luma's Dream Machine and Reasoning in Video Models
Assertion Supported
Song: Dream Machine Simulates Complex Metallic Reflections From a Single Photo
“With Dream Machine, we just took another coffee machine that we have in the office, and then just do image to video, and we can actually see a lot of the light reflections On top of the metal surface of the coffee machine being simulated by the model.”
Jiaming Song Sep 9, 2024 ▶ 17:35 Luma's Dream Machine and Reasoning in Video Models
Assertion Supported
Song: Dream Machine Reasons About Artistic and Non-Physical Video Scenes
“So I think the other interesting thing that we showed a slightly earlier, but also want to kind of reemphasize here is how dream machine is able to reason about the non-physical world as well. Even in cases where it's in entirely a scene of art, it is able to,…”
Jiaming Song Sep 9, 2024 ▶ 25:16 Luma's Dream Machine and Reasoning in Video Models
Assertion Not checkable as stated
Song: Current Text-to-Video AI Models Are Only at a Version Zero Stage
“Currently. No, we are only at a very, very you know, basic stage at, you know, being able to like generate videos from text and images at this stage. I will say this is more like a, as you said, research preview or version zero of the model.”
Jiaming Song Sep 9, 2024 ▶ 28:48 Luma's Dream Machine and Reasoning in Video Models

Appearances (1)

EpisodeDateSpeaking time
Luma's Dream Machine and Reasoning in Video Models Sep 9, 2024 23m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.