Bill Peebles, co-lead of OpenAI's Sora, discusses why OpenAI is investing heavily in generative video research as part of its AGI mission.
Prediction Not checkable as stated
Peebles: Sora will evolve into interactive world simulators with autonomous humans
“And so looking forward, as we continue to scale up models like Sora, we think we're going to be able to build these, like, world simulators, where essentially, you know, anybody can interact with them. I, as a human, can have my own simulator running, and I ca…”
Prediction Not checkable as stated
Peebles: Video models will eventually surpass humans as physical world models
“We're optimistic that Sora will, you know, supersede that kind of capability and will, you know, in the long run enable it to be More intelligent one day than humans as world models.”
Prediction Not checkable as stated
Peebles: Training AI on raw video is essential for robotics
“There's so much you learn from video, which you don't necessarily get from other modalities, which companies like OpenAI have invested a lot in the past, like language, you know, like the minutia of like how arms and joints move through space, you know, again,…”
Prediction Not checkable as stated
Peebles: Training on raw video is essential for future embodied robotics
“So you learn so much about the physical world just from training on raw video that we really believe that it's going to be essential for
things like physical embodiment moving forward.”
Assertion Not checkable as stated
Peebles: Sora is the first visual model with LLM-like breadth
“And so this is really the first generative model of visual content that has breadth in a way that language models have breadth.”
Assertion Not checkable as stated
Peebles: Generating long videos with Sora currently takes several minutes
“It's not instant and you have to wait at least like a few minutes for like these really long videos that we're generating.”