The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Robin Rombach no published score: only 1 usable exchange on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
1exchanges match
1on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So from the image to the video, the audio, and then eventually the real world with robotics, And a real world model, because if you can make the image, you, and you can train the model, that means by default, you understand the world. In order to make a video of the world, you have to understand the world, yeah? And the objects in it?

A I think that's, yeah, I think that's like a really good, like, way to think about it. It's like, it's like an intuitive way, uh, to, Interact with the world, right? Like, I would say there's, like, these, like, complementary forms of intelligence, ultimately. There's, like, intuitive intelligence, and then there's, like, a deep reasoning layer. Now, ultimately, you need, for, like, a kind of, like, complete form, you need both, um, and you need them to interact, and I think, like, we've been approaching it more from, like, the intuitive side. Um, images is, like, a very natural way to approach this whole field, because it's not as computationally intensive as, let's say, video, right? But now, yeah, I think, like, we're combining it, it's converging into, like, a multimodal model, and, uh, yeah, we see, like, exactly, like, pre-training on videos gives, like, implicit understanding of the physics of interactions with the real world, and then you can get stuff like action prediction, like robotics out of the same model.

AI assessment note: “pre-training on videos gives, like, implicit understanding of the physics of interactions”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 460 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.