Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q have like come out of nowhere in the last, you know, sort of like five, six months in, in particular, like for, from your perspective that, um, you know, is somewhat obvious idea that was sort of, um, you know, it was a matter of time until, uh, it was going to be implemented and, um, You know, it's good, but not completely groundbreaking. Or is that a major development?
A Well, it's been worked on for a very long time. So we've known it was coming for, um, for years, for years, for the past three years. Um, and it is kind of obvious. It is kind of obvious because if you think of like the pre reasoning world, um, The input space to a language model is everything. It's all of language. Uh, and so you can ask it very simple questions like one plus one or extremely complex ones like, um, go cure cancer. Like those are two strings or requests that you can ask it, and you really don't expect it to spend the same amount of energy and time on those two different problems. Um, One, it should respond immediately. The other one, it should probably, uh, you know, it might take years of thinking and trying to accomplish. Um, but we didn't have that reality before reasoning. Um, we had an input and then an immediate response. And so both of those got the same energy and effort put into them. So it had to come at some point, this notion of different amounts of energy or time being spent on problems, test time, compute, um, I think the effectiveness of it was surprising. Like it was really quite incredible to see how much gets unlocked. Like these models actually can with very little supervision, uh, very little data from humans saying, this is how you think through problems, sort of figure it out for themselves. Um, so that's been incredible to watch. I think …
AI assessment note: “It is kind of obvious because if you think of like the pre reasoning world”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Alright, so that's a model layer. Um, on top of that, uh, do you have an API layer? I mean, it seems that part of the business is to provide the models to companies like Oracle or Notion. Um, is that the model themselves? Is it a separate product?
A No, we provide the whole Platform. So serving of the models, optimization on different hardware, like AMD, um, we're up and running on Cerebris, Grok, of course, Nvidia. Um, so all of that we provide and we can deploy anywhere. And that's been quite unique. I think one thing that's different about agents and AI compared to other SaaS is that usually you're trying to do something that a human is doing in the organization. And To do that, you need the con, you need the same context that human has. And so that means you need to give very broad visibility, like our north platform to be fully useful needs to see all of your internal communications, all of your emails, all of your documents, your customer records, your et cetera, et cetera, et cetera. And that is a huge security risk, very unique one compared to like CRM software or HR software, that type of thing. Um, and so our security posturing, the fact that we don't say, Hey, send your data over to us, like hit our API, trust us. We're socked to the fact that we don't say that. And instead we say, we're going to ship our models directly on your hardware, whether it's in your VPC on a cloud or, you know, for regulated industries in your data center, that's been a huge unlock. People get comfortable plugging in much more. And so they can actually do more with the product.
AI assessment note: “No, we provide the whole Platform. So serving of the models, optimization on different hardware”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q is that because, uh, you know, it's the gift that keeps on giving, and the more, uh, you know, uh, data and compute you feed it, the better it performs, and we haven't reached the, you know, that, that moment when things get less exciting? Or is there something in research where, you know, people are, I don't know, for whatever reason, not as productive in terms of new ideas?
A I don't think people are not productive in terms of new ideas, but I, I do think it's, it's sort of a, um, You know, uh, a reinforcing loop or like a self-fulfilling prophecy where the community got super excited about the transformer. They built so much infrastructure specialized to the transformer. And so it's like we dug ourselves into this. Well, um, Like we now have chips that are being optimized explicitly to that architecture. And so to move architecture, it requires so much effort, energy, lift to rewrite everything and start from scratch. Um, that new architecture needs to present something extraordinarily compelling, like a very good reason to move. Um, and we just haven't found that architecture yet.
AI assessment note: “I don't think people are not productive in terms of new ideas, but”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q what, what, uh, what I would hear, uh, and tell me that's not correct. Uh, and this year, it seems to be, uh, people seem to be viewing synthetic data as, uh, you know, something that's completely ready for prime time and, um, uh, that they are increasingly using in just everything AI. So one, is that, is that fair or not? And two, is that fair or what happened?
A Yeah, there was a period where there were a lot of people saying, I forget the word of like the snake eating the, the snake, the, uh, or a Boris or whatever, like a human centipede of, uh, of data. Um, and yeah, I, I think it just got like decisively proven wrong. Um, Synthetic data is incredibly effective. It's now the majority of the data that we train on, um, for creating something like command A. And in many instances, it's actually more useful to the model than human data. Um, the most obvious example of that is stylistically. So humans, if you ask them to respond to a question, you know, we're lazy. Uh, if I ask what's, uh, you know, nine plus two, um, a human is just gonna say 11. Um, but what the human actually wants to see is, Aiden, that's a great question. That's so interesting.
AI assessment note: “Synthetic data is incredibly effective. It's now the majority of the data that we train on”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q um, telecom and, and so on and so forth is, is, is part of the idea, um, that you can train the model working with these partners, but like ultimately, obviously their data is their data. Uh, and if you want to generalize what you're doing to other financial customers, then you need to create synthetic data that looks like that data. Is that, is that part of the idea?
A Yeah, that's exactly correct. So sometimes our customers, um, either can't or don't want to train a model on their data either. Like they actually don't want to train on their data. Either they don't have permission to, or, you know, um, they're just not comfortable with it. And so in that case, synthetic is the only option. Um, but given a few examples of the ground truth data, synthetic data works extremely well. You can create huge quantities of fake data that is very, very faithful, um, to the ground truth. Uh, and so that's being used all over the place for us. Um, we're lucky in that most customers do trust us because of our deployment model. So we can deploy completely privately, like on premise, um, we can air gap it. Um, so it's just way more secure than some of the other fine tuning options that are out there. But if they still don't have comfort, we have this option that we can develop a bunch of synthetic data and show uplift in performance without having to actually train on real user data or real patient data.
AI assessment note: “Yeah, that's exactly correct. So sometimes our customers, um, either can't or don't want”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So going back to the story, the eight of you got together, and then how long was that process of writing the paper?
A Super, super fast. Um, so probably about a month in, we decided to consolidate and all work on, uh, the Transformer together. Um, and it just became A mad dash towards the, the NeurIPS conference deadline. So NeurIPS is like this, the biggest AI conference, uh, for academics where you submit your papers to. And so we were just all out sprinting. Um, and it was a lot of, honestly, it was a lot of like throwing shit at the wall and seeing what sticks. So many different things were tried. So many little bugs kind of hackily patched. Um, Like one example is pure attention architecture. Um, the model can't tell the difference between positions of the elements in the same way that an LSTM could, because it could consume each one, one by one. Um, And so then Noam just came up with this idea. I remember the day I was sitting next to him and he was talking to me about it, um, of just throwing these sinusoids, uh, into the embeddings and having that represent the position and it's stuck. And, uh, I think we've moved on a little bit, but shockingly, we're still quite close to that, that strategy today. Um, so that's, that's how stuff worked. It was just Fixing bugs one by one as quickly as possible. And then whatever we were left with at the last moment, that's what we submitted, uh, to NeurIps. And so one of the big shocks is how over the past eight years, how little things have changed.…
AI assessment note: “Super, super fast. Um, so probably about a month in, we decided to consolidate”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Interesting. Do you see, uh, enterprise demand for multimodal or is that more anticipation of what may come?
A Totally. No, there's, there's lots of demand. Some of the use cases, I mean, there's all the OCR type stuff. Um, but there's also like multimodal is Essential for understanding enterprise data, like PDF documents, where there's graphs and this type of thing, or understanding slide decks. A lot of the modalities that enterprises work in are visual. So it's sort of table stakes. The other thing that vision is crucial for is computer use, right? Like the ability to control a computer. It's a real, like it's a GUI. It's like a visual experience. Um, trying to use a computer or even like surf the web, looking at HTML, very, you could try doing it if you want. It's a horrible experience, uh, for language models too. Um, and so the ability to see is, is essential to being able to navigate it.
AI assessment note: “No, there's, there's lots of demand. Some of the use cases”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What are agents not Quite ready for what would you advise customers to not do?
A Well, there's all the like the sensitive use cases like medicine and these sorts of places, um, lots in finance, um, for those you want a human in the loop. You don't want to just hand it over to an agent and let the model, um, go crazy. Um, so there's places where it's not ready Just because we need good oversight, and it may never be ready, right? There's a, there's a large swath of things that we always want a human in the loop for. In terms of like technological limitations, where the model is not yet smart enough to accomplish something, I think one of the things that's said quite often now, and is a really good example of like how far the bar has been raised, is people are saying like, you know, these models haven't discovered new science. They haven't like, you know, created, like solve some like millennium problem or something like that. That is somewhere that the models are not yet super helpful. If you're a postdoc, it might be useful to your productivity of like reading papers and preparing talks and that type of thing. Um, but actually helping you discover a new compound I'm not sure how useful it is yet, but I'm very confident it will be useful very soon.
AI assessment note: “there's places where it's not ready Just because we need good oversight”
Answered raw tape
D 4 · C 5 · P 5 · Cm 4 4.55
Q So after the Transformers paper, what was your journey to starting Cohere?
A So I bounced around Google for a while. Uh, I went back to Toronto. I started working at Google for Jeff Hinton there. I met my co-founders, Nick and Ivan. Um, and then I started my PhD in England. Uh, and I was still working at Google. I was flying back and forth, uh, between London and Berlin because in Berlin, like London is, it was sort of exclusively deep mind, which was another arm of Google, which did AI research. And so brain didn't really exist there, but one of the, uh, transformer coauthors, Jakob, he had opened a brain office in Berlin. And so I would bounce back and forth to go, go see him. Um, and we were doing work with, um, Jeff Dean on scaling up infrastructure. So training on networks of TPU pods. So instead of having one supercomputer, you chain together a bunch of supercomputers and you can train something dramatically larger. And that's where we started to see the first instances of, uh, Scaling up large language models.
AI assessment note: “So I bounced around Google for a while. Uh, I went back to Toronto.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 5 4.45
Q So to get started, I'd love for you to, uh, tell the story of the Transformers paper. You're famously one of the eight co-authors of it, and of course, Transformers and attention is all you need is, uh, the seminal paper, uh, that led to everything that we're experiencing today in generative AI. What was your journey to it? How did you become part of the eight?
A I was a student at the University of Toronto, um, And I had been working on deep learning because a lot of the early work in the field was obviously done by, uh, Jeff Hinton and others at the school. Um, so I'd been getting quite close to it and I was reading up on papers and I just kept seeing Google brain repeatedly across these papers again and again and again, researchers from, from brain. And so eventually I, I ended up reaching out to them, uh, reaching out to some of the researchers there and saying, Hey, listen, I read your paper. I have this idea of how to extend it. And one thing led to another, and I got an offer to join.
AI assessment note: “I ended up reaching out to them... and I got an offer to join.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q contextual, and he was talking about FAIR, uh, at the time, and it sounded like it was a little bit like that as well. Do you, do you think that was a moment in time when like all those labs were authorized To, or, you know, allowed, uh, to, to operate that way and given free reign to explore anything. Is that still true today? Has it, has it changed?
A I don't know. I've, I've been out of Google for long enough that I'm not sure how the culture has shifted. I would say that the, the economic relevance of this work is very different to when I was, uh, when I was interning eight years ago. Um, so I imagine things would have to shift out of necessity. Um, Especially because of the product implications, the amount of resources that are being thrown at these projects and these models, it's much more consolidated. I would imagine and expect certainly the way we run things at Cohere, it's much more like a product organization. Like you have very clear work streams. Uh, there's scope to experiment and try to find new alpha. Um, but towards the ends of the product, right? So it's, it's focused and more narrow. Uh, but back then. It was very greenfield. Uh, and so you could work on whatever you were excited about. Um, and yeah, I do think that is like a crucial component of successful research organizations. Um, and it worked. It produced incredible technology.
AI assessment note: “it's focused and more narrow. Uh, but back then. It was very greenfield.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q the most side, uh, this concept of, uh, you know, AI continuing to be on some kind of exponential curve towards, uh, whether you call it human level intelligence or superhuman, is that something that, um, you see progress or effectively we've made a lot of progress, but like nobody really knows when that's gonna stop or whether it's gonna accelerate or. What's your sort of expert take on it?
A Listen, the models are going to continue to get better. They're going to continue to get better, and they're going to do some incredible things. Um, and there will be specialist models that emerge to help in things like, um, in pharma, right? The creation of new drugs, uh, in material sciences, um, for advanced materials. Um, The models will be able to be incredibly helpful to us. The definition of ASI and AGI is so hazy and ill-defined. It's hard for me to like give a concrete answer to what you're, what you're asking aside from it, it must be a continuum and not a discrete bit flip where suddenly it's ASI. Um, or AGI. I think we already have AGI to, to a large extent. Like, if you have, you have a choice between, like, you right now, you have some symptoms, and the only option is, uh, you know, Aiden prescribes me drugs, or Aiden's model, Cohere's model command prescribes me drugs. The logical and actually correct answer is to have the model do. I, I promise you it knows more than I do. Um, now it shouldn't be prescribing anyone drugs, but, um, It's smarter than me at that thing. Uh, so it is, and many other things, right? Like most things actually.
AI assessment note: “it must be a continuum and not a discrete bit flip where suddenly it's ASI”