The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jakob Uszkoreit no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Are there other big research areas that you're excited about right now, or areas where you see enormous progress being made?

A So in terms of foundations, I think different flavors of elasticity are, are really interesting. So You could actually claim that a lot of these questions boil down to the question that I just, to basically this problem that I just, uh, described, right? That compute is, in a certain sense, very crudely allocated. But you can look at different incarnations of this problem. So another one would be, why don't we have models that in an elegant way manage to consume, say, visual sensor output of different resolutions, different sampling rates, different durations, right? Right now, it's actually quite tricky to have, other than maybe recurrent architectures, a model that takes videos of different lengths, different image resolutions, or ultimately different densities, if you wish, and different sizes, and, and really elegantly adjusts compute to what you really want to know about this, or the, how, how difficult it really has to be to generate the representations that you, that you need in order to do whatever you want to do. And here, again, an example that makes this, I think, pretty clear is you can take a video You can scale it up. You can frame interpolate with trivial algorithms, and then run it again. And if the problem you're trying to solve, conditioned on that video is the same, then I wouldn't want more computers to use. But right now, that's what's going to happen. You'…

AI assessment note: “in terms of foundations, I think different flavors of elasticity are, are really interesting.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So one thing that you've been working on for the last few years is Inceptive, which is really starting to focus on how can you apply machine learning and different aspects of software to biology? Could you share a little bit about the company, how you got interested in bio and what you view as some of the interesting problems there?

A Yeah, so basically I've always been interested in bio and know nothing about it. And that's a conundrum because it's difficult to learn a lot about biology when you're not in school and I didn't want to go back to school. But at the same time, it always felt like something where there is a lot of headroom in terms of efficiency and actually also where maybe even Alternative approaches, at least if what you are interested in is really solving acute problems, where there's maybe a dire need for alternative approaches. Alternative to, basically, biology to science that is trying to develop a complete conceptual understanding of how life works. I don't have very high hopes for humanity to develop that conceptual understanding to the level that we would need it in order to do all the interventions we want to do. We don't really have great tools in our toolbox, or we didn't have them until somewhat recently as, as alternatives to understanding how it works, and then basically based on that understanding, fixing it, if it needs fixing. And I think now we have an alternative that's an extremely good match, and that's deep learning at scale. We're really, we can, Potentially, to a pretty large extent, if not entirely, whatever this even means, work around the following two problems. Number one is, we don't know all the stuff that's going on in life, right? So we still just don't even ha…

AI assessment note: “I've always been interested in bio and know nothing about it.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Where do you think people should invest that sort of optimism going forward? Like what are the big areas that people need to work on to increase the performance of these systems or add memory or do other things that you feel are, if you, if you were to sort of paint the roadmap going ahead in terms of making these really valuable performance systems, what would you focus on?

A I mean, I think there's one thing that still boggles my mind in terms of just from first principles that it can't be optimal. And that is that if you think about it, The way you today scale the compute that's invested in a given problem, right? Let's say the problem is what's the response to a prompt in some large language model. Then ultimately, the way you scale that compute depends on the prompt and how much, how long that is. The longer the prompt, the more compute you get. And it depends on, and there's of course many different screws to tweak here, the length of the response. There are many very hard problems where the response is incredibly short. And you can, in many cases, actually formulate those problems very, very succinctly. So you're not going to be using a lot of compute, even though the problem we know is really, really difficult. Say, I don't know, prime factor, prime factorization. A problem like that, simply stated, big potential impact, and right, right now, there's no knob that you can easily tweak as a user, but also really, there's no knob that the architecture can tweak itself when it comes to then basically deciding, oh, this is hard. I actually need to use more compute for this. And ironically, and this comes back to a question that many people ask, I think around, does it make any sense to train on generated data? Because information theory, family in…

AI assessment note: “there's no knob that the architecture can tweak itself when it comes to... more compute”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q So how do you think about human augmentation in the context of all this stuff? You know, how bullish are you on human augmentation and what forms do you think it'll take in the near term?

A I'm very bullish on human augmentation in the very long term, but it's one that I don't see intuitively. I think looking at our brains, even just physically, they seem to be very focused, and this is not surprising, on our IO. And why would there somewhere in there be some kind of computational capacity that if we just boosted our IO by a few orders of magnitude could still cope? Why would evolution put that there? I don't know why. And so, yes, you could argue, you know, maybe to do long-term planning tasks and so on and so forth, but sure, right, let's found it a lifetime. So, right, it's just not so clear whether there would have been any evolutionary pressures to really make our capacity there Much bigger than, say, some multiplier, basically time on our I.O. capacity.

AI assessment note: “I'm very bullish on human augmentation in the very long term”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q be other architectures that have equally interesting or perhaps more interesting properties at scale, but there's sort of two impediments. Number one is people just aren't throwing a lot of money and compute at it. And two is the underlying accelerator architecture actually fits so well. It is dramatically less performant to do other architectures and therefore we may never actually test them. Do you think that's a true statement?

A I think that the big question is, does it matter? It would be really interesting to evaluate, especially if we can make them simpler, to evaluate combinations of different hardware and then models or architectures that fit like gloves to those. And I feel at the moment, given where GPUs came from, they weren't built for this, right? Why would it be that they are anywhere near optimal? If at least they were engineered for this purpose, and lots of people basically banged their head against walls until they had this kind of somewhat optimized, but that's not how the basic architecture came to be. And so you can talk a lot about and reason a lot about, and I think that some of that is true, the generality of basically really fast scalable matrix multipliers and how that just does everything in scientific computing really well.

AI assessment note: “I think that the big question is, does it matter?”

Partly raw tape D 3 · C 3 · P 3 · Cm 3 3.00

Q When you think about how we get progress from here, usually people think of Software as driving the hardware, right? Do you think we get accelerators designed for the large scale transformer architectures we already have or new hardware designs? Like it's chicken or egg a little bit here.

A It's chicken and egg. And if you look at the newest accelerator designs, they are taking this into account to a significant extent, actually, increasingly. So There are a couple of interesting examples. We had a computer vision architecture that really was just an MLP called Mixer, and while it wasn't significantly better, it also wasn't significantly worse than the vision transformers, right? And I think that already goes to show it's not that difficult, and especially if you simplify on the way, it might really be a possibility. I will say one other thing, aside from efficiency, but just really raw efficiency in terms of its fit, the architectures fit to the accelerator hardware. The other main contributor, I think, to the success of this architecture was optimism and hope. So suddenly you were in a situation where, for whatever reason, a bunch of things that people tried with this started to work, and then more started to work, and that's not coincidence. It's really just because ultimately the human cycles invested to getting all these things, all these diverse things to work, are ultimately fueled by suspension of disbelief. And AKA hope or whatever you want to call it. And, and that really, I mean, the community became so energized so quickly and then just try everything under the sun. And because the prior was just a different one. The prior now was, oh, look, we have th…

AI assessment note: “if you look at the newest accelerator designs, they are taking this into account”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.