The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Aidan Gomez no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q uh, increase in GPUs are going to get us first and foremost, and then we'll talk about whether the right way to scale these models is with just throwing more compute. Uh, cause I know you have a nuanced take on that. Uh, but like, first and foremost, like, If you go from 16,000 GPUs at the state of the art to a 100,000, what do you think that delivers?

A Uh, definitely delivers a bigger and better model. You have more compute. We know that scaling up improves things. Um, there's questions around saturation and whether continuing to scale up is justified, whether there's going to be enough gains from that strategy to justify the increase in cost. My personal perspective is that, you know, building a massive model, it's not actually useful for the world if it's too big to be consumed, if it's too expensive to actually deploy. Um, and so for Cohere, we've been very focused on building the right size of model. Um, but if your question is what is more compute unlock, it will be a better model. Objectively, we know that scaling leads to more capability, a smarter model that's more reliable. Um, and so that's the, the output.

AI assessment note: “definitely delivers a bigger and better model. You have more compute.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q The models are as good, right? We're still like at work.

A I don't think those two things are in conflict. I don't think that really, why not? Um, well, because I think that, um, we will never see mass unemployment of humans. I think that this technology is going to unlock more opportunities. It will let us do more as opposed to Scaling back what we do. Um, humanity is very supply side constrained, not demand side. We always, we want more. We want better. We want to be, um, healthier. We want to do more. We want to have things be cheaper. Um, and so we have all this demand and we're trying to keep up with our own society's demand and this technology, it's, it's true promise is in bringing productivity. And letting us do more. Now you can zoom in and you can like pick a specific field and you can say this field might be automated by AI. And I think that's true. And, you know, we should be thinking about retraining and shifting certain skill sets over to other, uh, new domains, like retraining people. But in general at the macro scale, I think this technology will create much more opportunity, uh, than it will take away.

AI assessment note: “I don't think those two things are in conflict. [...] we will never see mass unemployment”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q that they've found them able to sort of come up with things that are outside of their training set. And there have been some papers that say, okay, actually they don't really have any emergent properties or emergent behaviors. And as someone who wrote the paper that kicked this all off, what do you think about that? Can LLMs have Any emergent behavior or discoveries that they weren't trained on.

A I think that they can, um, what's the right word? I think they can interpolate between skills, and so if they've seen how to do A, and they've seen how to do B, they can get kind of the average of A and B, um, but they don't just go completely beyond anything that they've seen. Um, I've never seen a model behave in a, um, totally unexplainable way. Um, they're really good interpolators. If you show them different domains, they can blend domains quite well, and, um, but yeah, I, I've heard the same thing about emergent behaviors, and, um, I, I think the research is really inconclusive there. Uh, there's not a lot of compelling evidence that says we're gonna have some Uh, total step change or capability takeoff. Um, even in the, the, like the latest state of the art research, a lot of it's about synthetic data and models teaching themselves. And so, um, self-improvement is this notion of can a model actually teach itself without human intervention. This is now a huge part of model building. It's a big part of how we create data at Cohere. Um, and before this started to become mainstream and actually part of the production process of creating these models, people were saying self-improvement, these things are just going to take off. They're going to become superhuman overnight and we won't be able to control it. Um, well, it turns out that doesn't actually happen.

AI assessment note: “they can interpolate between skills... but they don't just go completely beyond”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q This is your expertise. You're like a, you know, almost like a narrow neural net train for that one specific purpose. And now we're giving it over to AI. And they'd look at the, just the billions being invested in this technology and say, well, what am I really getting for that? If this is effectively doing some of these things that humans are quite good at to begin with.

A I would counter and say that, um, risk to supply chains is many trillions of dollars. I would say that lawyers are extremely, extremely expensive and you don't want them combing through your documents, no matter how efficient you think they are. And same thing with doctors. We, we really want them spending time with patients, not, um, combing through hundreds of notes, uh, and filling out, filling out forms afterwards. Um, I would say those are. Maybe this stuff feels banal. Maybe productivity feels, um, boring compared to some of the hype of AI, but it is the value. This is what we're trying to build for. Um, and so I, I would push back quite firmly on that.

AI assessment note: “risk to supply chains is many trillions of dollars. I would say that lawyers are extremely”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q can add knowledge to a model. But as the model got smart, the models got smarter and smarter, became less easy for people to add supplemental. Knowledge to them, which points to like sort of running out of available data to make these, uh, AI models smarter. So how does like AI generated data actually solve that problem? And where is synthetic data being used to make these models better?

A Yeah, um, yeah, so I think the example you gave is, is a good one, like, where it's getting harder and harder to get the data that incrementally improves the model, and it's important to note that that's because the model is getting so much better, um, and so before we could just grab anyone off the street and they could teach them all something, and then that signal started to go away, and so we had to go to Undergrad students in bio to teach the model about bio. And then we had to go to master's students and then PhDs. Um, and we're kind of at that level where we're currently, um, hiring PhDs to teach the model in their specific domain. Um, but then after PhDs, where do you go? Right? Like you, I guess professors, uh, what about after that? Um, so I, I think, um, the models are catching up with the state of knowledge across a bunch of different fields. Um, I would say that synthetic data Probably doesn't get us out of that, that issue. I actually, I don't know if synthetic data outside of easily, um, verifiable domains like math, it's hard to use synthetic data to drive outcomes. So we'll be able to do it in.

AI assessment note: “synthetic data Probably doesn't get us out of that, that issue.”

Partly raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q How far away do you think we are from having reliable assistance? Like a lot of people looked at, uh, OpenAI's O-One reasoning model, and they're like, oh, this is just kind of like a step toward assistive AI. Um, what do you think?

A I think the notion of using reasoning or letting the model have an inner monologue to work through problems, think through them, um, make mistakes, but then realize that catch mistakes and correct them. I think that's a crucial piece in improving not only the, the accuracy, um, or robustness or usefulness of the model. Um, but also the. The trust in the model because you're able to inspect how it arrived at its conclusions, how it decided to do what it did. You actually trust it much more. It's explicitly written out. Um, and so I, I think we've all known these sorts of tools, uh, would need to emerge. Um, and yeah, I think it is a big step towards dramatically more reliable assistance, ones that you can trust and work with and give feedback to. Um, yeah, I think it's really exciting.

AI assessment note: “I think it is a big step towards dramatically more reliable assistance”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.