The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Joelle Pineau argument clarity score 4.4/5 from 40 exchanges on raw tape · average scores: directness 4.6 · coherence 4.9 · precision 4 · compression 4 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
40exchanges match
40on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q You said about synthetic data, and that also being a very important segment to consider. Do you get model degradation when you get this kind of reinforcing loop of models learning on synthetic data, which creates more data for synthetic, and it actually degrades, or does it improve?

A It really depends how you're generating your synthetic data. So in some domains, if you think like images, languages, like LLMs talking to each other at some point, you definitely get the degradation and that degradation is due to essentially like a loss of diversity of your data. So, you know, you can make an analogy, you know, you, you take a bunch of people, put them on an island and let them reproduce, you know, at some point the genetic diversity is gonna keep shrinking. Um, and so you get a reasonably similar phenomenon with, uh, with models. Because you're not injecting diversity into the data. So for their domains where lack of diversity means you get a collapse of distribution. There's other domains where, um, you don't need diversity. If you think of like, you know, playing chess, playing go, these kinds of games, we know exactly how to generate board configurations. And so we can generate tons of synthetic data. Not endless because it's a closed world, but still tons of synthetic data and through that learn for a long time. Then there's domains that are sort of in between. If I think of coding, we can generate synthetic code. You take normal code and we know how to inject diversity into the code. Like I can take a couple of repositories, mix and match, apply an LLM to transform it. And so there's a way to generate synthetic data. The language is predictable enough, a…

AI assessment note: “It really depends how you're generating your synthetic data.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Does progression happen in kind of a linear fashion, or does it happen in step functions like AlphaGo, like a DeepSeq, which depending, depending on kind of what you believe, suggests a lot of efficiency in terms of model improvement. Is it step function or is it linear?

A I tend to decompose different ingredients that lead to progress. You know, people often talk about like the algorithms, the data, the compute. I think in general, Compute and data have a more linear effect on progress. You build more compute. You run bigger models. You can typically get better performance. You feed in more data. It's not just quantity. You need to worry about quality and diversity as well, but roughly it's more linearish with respect to the data. The algorithms are the ones that have the nonlinear effect, and so you can explore lots of ideas, and then something like the transformer comes along and just changes the paradigm. And it's not just a transformer, you know. On the optimization side, suddenly we hit upon atom, which is a technique to, to do the optimization of your model, changes in the paradigm. Reasoning, suddenly we start thinking about how to put that in the loop reasoning, and it changed the paradigm. So those ideas tend to have a nonlinear effect. The challenge with these algorithmic ideas, though, is that actually It may take a long time to prove themselves out. So like the paper can be sitting out there. There's thousands of papers coming out. The idea is sitting out there and we may not think to try it with the right data at the right scale with the right combination of hyperparameters. And so you don't notice that effect for a while. So it's h…

AI assessment note: “Compute and data have a more linear effect... algorithms... have the nonlinear effect”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Does progression happen in kind of a linear fashion, or does it happen in step functions like AlphaGo, like a DeepSeq, which depending, depending on kind of what you believe, suggests a lot of efficiency in terms of model improvement. Is it step function or is it linear?

A I tend to decompose different ingredients that lead to progress. You know, people often talk about like the algorithms, the data, the compute. I think in general, Compute and data have a more linear effect on progress. You build more compute. You run bigger models. You can typically get better performance. You feed in more data. It's not just quantity. You need to worry about quality and diversity as well, but roughly it's more linearish with respect to the data. The algorithms are the ones that have the nonlinear effect, and so you can explore lots of ideas, and then something like the transformer comes along and just changes the paradigm. And it's not just a transformer, you know. On the optimization side, suddenly we hit upon atom, which is a technique to, to do the optimization of your model, changes in the paradigm. Reasoning, suddenly we start thinking about how to put that in the loop reasoning, and it changed the paradigm. So those ideas tend to have a nonlinear effect. The challenge with these algorithmic ideas, though, is that actually It may take a long time to prove themselves out. So like the paper can be sitting out there. There's thousands of papers coming out. The idea is sitting out there and we may not think to try it with the right data at the right scale with the right combination of hyperparameters. And so you don't notice that effect for a while. So it's h…

AI assessment note: “Compute and data have a more linear effect... algorithms are the ones that have the nonlinear effect”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q You said about synthetic data, and that also being a very important segment to consider. Do you get model degradation when you get this kind of reinforcing loop of models learning on synthetic data, which creates more data for synthetic, and it actually degrades, or does it improve?

A It really depends how you're generating your synthetic data. So in some domains, if you think like images, languages, like LLMs talking to each other at some point, you definitely get the degradation and that degradation is due to essentially like a loss of diversity of your data. So, you know, you can make an analogy, you know, you, you take a bunch of people, put them on an island and let them reproduce, you know, at some point the genetic diversity is gonna keep shrinking. Um, and so you get a reasonably similar phenomenon with, uh, with models. Because you're not injecting diversity into the data. So for their domains where lack of diversity means you get a collapse of distribution. There's other domains where, um, you don't need diversity. If you think of like, you know, playing chess, playing go, these kinds of games, we know exactly how to generate board configurations. And so we can generate tons of synthetic data. Not endless because it's a closed world, but still tons of synthetic data and through that learn for a long time. Then there's domains that are sort of in between. If I think of coding, we can generate synthetic code. You take normal code and we know how to inject diversity into the code. Like I can take a couple of repositories, mix and match, apply an LLM to transform it. And so there's a way to generate synthetic data. The language is predictable enough, a…

AI assessment note: “It really depends how you're generating your synthetic data.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What are the potential vulnerabilities, though, in an agent world?

A In terms of agents, you know, we, we worry a lot about hallucinations in LLMs. The parallel in agents is impersonation. So agents that come along and are essentially impersonating entities which they don't legitimately represent, and in doing so, taking actions on the behalf of these entities where they don't legitimately represent, whether it's, uh, infiltrating, you know, banking systems and, and, and so on. And so, I, I do think we have to be quite lucid about this, develop standards towards that, develop ways to, to test for that in a very rigorous way. There's ways to reduce that risk drastically. You run your agent, you know, completely cut off from the web. You're reducing your risk exposure significantly. Um, but then you lose access to some information. So depending on, depending on your use case, depending on what you actually need, there's different solutions that may be appropriate.

AI assessment note: “The parallel in agents is impersonation.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Great. Why is data becoming more expansive?

A Data comes in, in different forms. And, um, on the one hand, you know, the, the notion, the days of like having data labelers who can say this is a cat and this is a dog or somewhat over like the easy test the AI can do. So we're getting in a space where we need more specialized tasks. So, you know, imagine you're building AI for enterprise. There's a particular business logic. You need to make sure that you're catching the errors. You're gonna need someone with, like, deeper understanding of the tools, so that's more expensive talent to come in and actually prepare the data. Um, there's also a lot of data that's synthetic data, so when you're building agents, you need to build environments, and to build environments, you need some pretty creative folks who are going to build you, like, synthetic simulators. Um, we've seen this on the robot side for many years, people building robot simulators. Now you're building AI for enterprise, so you need to think of like, how are you going to simulate these work processes in a reasonably realistic way that the AI can train on that? And so that generation of environments and benchmarks and dynamic domains, uh, can, can be pretty expensive too.

AI assessment note: “so that's more expensive talent to come in and actually prepare the data.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally get that. Final one. What are you most excited for? You don't like the existential risk. I don't like, like, the doomsday planning. When you think about the positivity that can come, what are you most excited for when you look forward to the next three to five years?

A I do think some of the work in terms of, um, AI for, um, a scientific discovery is going to be pretty fascinated, pretty fascinating to see. Um, just in terms of the doors, it's gonna open up the ability to explore, you know, combinatorial space of, of solutions. So I'm curious about that. Um, And then I'm super curious to see, you know, how can we actually make our models more efficient? There's larger and larger and larger models. No one wants to run these models. You know, I spent a lot of my career building, um, open source models. And I'll give you one example. You know, we were in, in the frenzy of large language models, and I pulled the stats on, you know, most downloaded models of last month. We had a model like Roberta from 2019, small language model, was getting twenty million downloads a month. People want efficient models that they can use, that they can run, so I'm also super keen to see what we're going to be able to do at the scale that runs on like one or two GPUs.

AI assessment note: “AI for, um, a scientific discovery is going to be pretty fascinated, pretty fascinating”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you think prompts and the way that we interact today with prompts and with chat largely is like the enduring interface for human engagement with AI?

A It's awfully limited, and, you know, prompts can mean a few different things, but, but the idea of, like, typing in a box, that to me is very limited, and we're gonna break out of that box already. We're seeing a lot of cases where voice is a lot more natural as an interface. I do expect we'll see, you know, gesture, eye gaze, these kinds of much more multimodal ways to interact with the, with the AI rather than just stick in that, in, in, in that prompt box. But language is incredibly powerful. So if you think of prompt as being more language as a way to express ideas and communicate with a machine, that's a powerful paradigm. I mean, as humans, we, so much of our communication is based on language. I don't think we're going to move away from that because it encodes information. You know, language words are symbols that encode so much information so efficiently. And so I don't think we're close to getting away from that.

AI assessment note: “the idea of, like, typing in a box, that to me is very limited”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally, totally get you. Um, what's the biggest challenge about capital efficient AI today? I know that sounds strange, but when you look at the economics, so to speak, what's the biggest challenge?

A Um, there's a lot of challenge today, I think, in terms of the economics of AI. I think one of the biggest challenges, the fact that it's very hard to have predictability, right? Everyone wants to know when are we going to hit the breakthrough? Everyone wants to know how many GPUs do I actually need? Everyone wants to know, like, what's the return I can expect? There's just a lot of uncertainty built into the system. A lot of that is because There's a lot we don't know about this technology, and so that means we have to take in quite a bit of risk when you're building out, whether you're building out your data center, whether you're building out your workforce, whether you're trying to figure out, you know, how much data to, to curate, and so that makes it difficult for a lot of people. People want answers, and this is a world where we don't have that level of predictability compared to other industries.

AI assessment note: “one of the biggest challenges, the fact that it's very hard to have predictability”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm sorry, how, how does that actually shape out then?

A I think you have to identify very concretely the types of work that you are delivering, but I think we're starting to see, like, you know, Hollywood quality productions being made in a matter of hours. We're seeing, you know, to take a super concrete case, like machine translation, if humans are doing the machine translation compared to machines doing it, you go from hours to seconds on long form text, multi-page documents. And so for a lot of work, It's not like AI can do all of the work. Humans still need to ask the right question. They need to verify the information. They need to shape the tasks. But once the task is well-defined, the product are clear, like all the design considerations are fed into the prompt, you press the button and you've got an answer in seconds for something that used to take sometimes weeks and months.

AI assessment note: “you press the button and you've got an answer in seconds”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q When we look at the cost curve for RL, you said you've been working on it for 20 years. Have we seen that dramatically come down? Will we see it continue to dramatically come down, or is it a case of it is just a fundamentally expensive method of training?

A It's coming down, especially in domains where we have good reward functions. So the place where most people started hearing about RL is around the AlphaGo time, you know, the game of Go, which was sort of one of the goals for AI. Many people thought we were, at that time, we were still a decade away. From being able to have machines play go at the level of humans, and out comes a team from DeepMind, you know, goes off, plays against the world champion, and shows that that RL can basically do it. Um, and so I would say, you know, in cases where we clearly know what's the goal, we can write down precisely the reward function, we're good. We can make a ton of progress, so that's why you're seeing progress in mathematics, um, Very well-defined reasoning tasks, aims, these kinds of things. RL, to shape the behavior of models, to get them to be social creatures, that we have no idea how to do. I mean, I don't know if you have children, but like shaping their behaviors, you know, the number of times you can repeat the same thing, and still they do something else. And so there's something there. You don't know how to write that out mathematically. And that's where I think we're still in for, for some hard work.

AI assessment note: “It's coming down, especially in domains where we have good reward functions.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally get that. On the team building side, obviously, Canada has great talent. You mentioned, obviously, some in London as well. What have been your biggest lessons, observations on team building in this, like, talent frenzy that we're in also? How, how do you analyze that?

A One of the things that's, that's important when you're, you're building a team for AI, I do think you need people who have vision, who have like a sense of like, what can we create? Just because we're in a space where there's so much innovation that is still needed. So you need an ingredient of vision that can be one, two, three people who bring that, that ingredient of, of vision. You need people who have amazing execution muscle. Like, they don't care that it's their idea. They care that if the team agrees on an idea, they are just gonna push this and get it done. They're gonna build a system. They're gonna run the experiments. They just have that technical rigor to execute. And then you need people who kind of, like, keep the team together, who have, like, the sense of, like, who needs what to operate well and who are that social glue. You know, humans are still social beings, and that social glue in a team matters a lot. Where I've seen it fail is to have just sort of one type of person inside the team. Um, I don't think it, it becomes that productive to, to put a bunch of AI superstars all together in a room without the execution machine, without the social glue. I don't think you get necessarily the same results. So I'm, I'm a big believer in building teams with, with diverse complementary.

AI assessment note: “Where I've seen it fail is to have just sort of one type of person”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Will we have, what will the developer world look like in 10 years when that is the case?

A Well, if I carry my analogy further, I don't know if it's a reassuring scenario, because if we look at where we are today in terms of image generation, there's just, like, the volume of image getting generated is huge. What matters now is sort of, you know, picking the, picking the quality out of the volume, and so if I fast forward 10 years on code generation, when we have the ability, To generate a ton of code to do a ton of different things. We're going to need some selection mechanism to decide what code we actually want when there's actually value. And so that's going to come. There's still going to be some sort of editorial design choice. Someone needs to decide, like, of all the code we can generate, what's the code we want to generate? What do we need to be running in terms of our digital world?

AI assessment note: “There's still going to be some sort of editorial design choice.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q on the show suggest that evals are, to put it delicately, bullshit, uh, and that they don't actually mean anything anymore, and like, you know, humanity's last test, like, what does that really even mean? Uh, and we have these new tests that come up, What is this? And leaderboards, what is this? Is that fair, or do you think they actually serve a very effective utility to the ecosystem?

A I do think they, they are, um, really good indicators. So I think you do need to take evaluation seriously in terms of knowledge, but you shouldn't take them seriously in terms of the ultimate goals. So, you know, evaluation, and there's lots of different benchmarks and so on. You have to decide, like, what type of models or model are you building? What's the characteristics of your system? And then think of evaluations as, like, unit test for the performance of your system. I mean, software engineers will know what that is, right? Like, you run through that evaluation, and that gives you, like, a signal of how the system is doing in a particular dimension. But as we're building systems that are more and more general, do very specific tasks, you don't optimize for these, right? Like, I mean, we build AI systems that go into enterprise. None of our clients ask about, like, are you able to win the math Olympiad with this model? That's not what they care about. They care about bringing value to their business. Now, we're curious to know how well we do on math problems because it can be predictive of behavior on other things, but you don't obsess over specific benchmarks. You kind of look at the ROI in terms of what you're trying to build.

AI assessment note: “I do think they, they are, um, really good indicators.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Have we over invested in RL based methods at the expense of maybe, like, more scalable alternatives?

A Oh, I'm still super bullish on RL in that, like, the concept itself is so fundamental. You know, this idea of training through a system of rewards, of indicating what's valuable and what's not valuable through numerical values, like, that is so fundamental. It's not going away. Now, you know, where we're maybe getting a little bit ahead is thinking that just RL out of the box is gonna give us AGI. That part, a lot less so. You know, if you look at the curve of progress, RL is terribly inefficient, and so the amount of signal you need to get in order to really shape the behavior of a model is far from where we are today, and so we'll need to figure out how to, to really deal with this, with this learning efficiency problem.

AI assessment note: “I'm still super bullish on RL in that, like, the concept itself is so fundamental.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Security is a topic that we quite often glaze over, especially when investing in kind of application layer AI tools. What does no one know about AI security that people should know?

A With respect to AI security, I think there's a new front that's opening up with, um, the development of agents, and frankly, there's a lot we don't know yet in terms of the vulnerability of these systems. With LLMs, we're starting to get a better understanding. We've had, you know, quite a bit of red teaming exercise and jailbreaking and so on, and so people have identified different risk vectors, prompt injections, things like that, which are Vectors for malicious actors to interfere with the system. With AI agents, we haven't seen that. And one of the features of computer security in general is often, you know, it's a bit of a cat and mouse game, quite frankly. Like, there's a lot of ingenuity in terms of breaking into systems, and then you need a lot of ingenuity in terms of building defenses. And so, we just have to stay very active in that sense.

AI assessment note: “there's a new front that's opening up with, um, the development of agents”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally get that. On the team building side, obviously, Canada has great talent. You mentioned, obviously, some in London as well. What have been your biggest lessons, observations on team building in this, like, talent frenzy that we're in also? How, how do you analyze that?

A One of the things that's, that's important when you're, you're building a team for AI, I do think you need people who have vision, who have like a sense of like, what can we create? Just because we're in a space where there's so much innovation that is still needed. So you need an ingredient of vision that can be one, two, three people who bring that, that ingredient of, of vision. You need people who have amazing execution muscle. Like, they don't care that it's their idea. They care that if the team agrees on an idea, they are just gonna push this and get it done. They're gonna build a system. They're gonna run the experiments. They just have that technical rigor to execute. And then you need people who kind of, like, keep the team together, who have, like, the sense of, like, who needs what to operate well and who are that social glue. You know, humans are still social beings, and that social glue in a team matters a lot. Where I've seen it fail is to have just sort of one type of person inside the team. Um, I don't think it, it becomes that productive to, to put a bunch of AI superstars all together in a room without the execution machine, without the social glue. I don't think you get necessarily the same results. So I'm, I'm a big believer in building teams with, with diverse complementary.

AI assessment note: “I don't think it, it becomes that productive to, to put a bunch of AI superstars”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Will we have, what will the developer world look like in 10 years when that is the case?

A Well, if I carry my analogy further, I don't know if it's a reassuring scenario, because if we look at where we are today in terms of image generation, there's just, like, the volume of image getting generated is huge. What matters now is sort of, you know, picking the, picking the quality out of the volume, and so if I fast forward 10 years on code generation, when we have the ability, To generate a ton of code to do a ton of different things. We're going to need some selection mechanism to decide what code we actually want when there's actually value. And so that's going to come. There's still going to be some sort of editorial design choice. Someone needs to decide, like, of all the code we can generate, what's the code we want to generate? What do we need to be running in terms of our digital world?

AI assessment note: “if I fast forward 10 years on code generation... We're going to need some selection mechanism”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q there about kind of access for enterprises. Enterprises have money, and that's a great luxury in a lot of cases. Research institutes, universities often don't. With the kind of bubble-like tendencies, people with money are able to afford to compute the talent. Are we seeing this kind of lack of access or democratization for great institutions that are educational maybe, who now can't afford to compete in this new world?

A Certainly a lot of universities have a lot less resources than, than companies today. That's not completely new. Um, when I joined Meta in 2017, um, one of the reasons I did that is because I, I could already see the, the disparity in terms of access to compute. And I was really curious to see how you could do research with a lot more, uh, compute, but you know, there's still amazing research that's being done in, in universities. You go to the major international conferences, Nureps, ICML, and others, and often the best paper awards are actually won by researchers out of universities. There's a lot of good ideas that you need to test out at small scale. Um, and in a university, you have a lot more freedom to pick pretty risky ideas at a small scale, but still, you know, no one's asking you to, to, to, To justify your, your, your research in, in ways that, that often happens in companies. So I think they have, they play different roles in the ecosystem, and what's actually especially good is talent flows between, between them. You know, university students come in, do internships, take jobs at companies. We've also seen a movement of people coming out of these large companies, going back to university, teaching, sharing with the next generation what they've learned.

AI assessment note: “Certainly a lot of universities have a lot less resources than, than companies today.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q When we look at the cost curve for RL, you said you've been working on it for 20 years. Have we seen that dramatically come down? Will we see it continue to dramatically come down, or is it a case of it is just a fundamentally expensive method of training?

A It's coming down, especially in domains where we have good reward functions. So the place where most people started hearing about RL is around the AlphaGo time, you know, the game of Go, which was sort of one of the goals for AI. Many people thought we were, at that time, we were still a decade away. From being able to have machines play go at the level of humans, and out comes a team from DeepMind, you know, goes off, plays against the world champion, and shows that that RL can basically do it. Um, and so I would say, you know, in cases where we clearly know what's the goal, we can write down precisely the reward function, we're good. We can make a ton of progress, so that's why you're seeing progress in mathematics, um, Very well-defined reasoning tasks, aims, these kinds of things. RL, to shape the behavior of models, to get them to be social creatures, that we have no idea how to do. I mean, I don't know if you have children, but like shaping their behaviors, you know, the number of times you can repeat the same thing, and still they do something else. And so there's something there. You don't know how to write that out mathematically. And that's where I think we're still in for, for some hard work.

AI assessment note: “It's coming down, especially in domains where we have good reward functions.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally, totally get you. Um, what's the biggest challenge about capital efficient AI today? I know that sounds strange, but when you look at the economics, so to speak, what's the biggest challenge?

A Um, there's a lot of challenge today, I think, in terms of the economics of AI. I think one of the biggest challenges, the fact that it's very hard to have predictability, right? Everyone wants to know when are we going to hit the breakthrough? Everyone wants to know how many GPUs do I actually need? Everyone wants to know, like, what's the return I can expect? There's just a lot of uncertainty built into the system. A lot of that is because There's a lot we don't know about this technology, and so that means we have to take in quite a bit of risk when you're building out, whether you're building out your data center, whether you're building out your workforce, whether you're trying to figure out, you know, how much data to, to curate, and so that makes it difficult for a lot of people. People want answers, and this is a world where we don't have that level of predictability compared to other industries.

AI assessment note: “one of the biggest challenges, the fact that it's very hard to have predictability”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q To now also building product. Is there ever this, like, inherent conflict between intellectually interesting research with the need to productize and monetize, and how do you think about that?

A I mean, one of the reasons I'm really excited to be joining Cohere actually is because like we're at a stage where AI is really starting to be useful, maybe not as useful as people think it is, but we are there. Um, and by working on AI that's going into enterprise, I feel we're going to get such an interesting signal of what works and what doesn't work. You know, we keep on talking about, you know, AGI and AI for the masses and so on, but Actually, like when you need to sell AI to a business, you get a real signal of what works, what doesn't work. Um, and that's what I'm most curious to see. Um, and you know, we've been using these academic benchmarks for many years. You get some signal, but it's not the same as, as getting this to do productive work. Um, so I'm curious to, I'm curious to learn out of that, you know, we're going to get new types of data. We're going to get, I think, a lot of insights That are then going to drive the research ideas. Um, I think that's, that's the other thing to think through when you have a large space of ideas to explore. Getting that feedback signal from the real world is super useful to guide you through that search of ideas.

AI assessment note: “insights That are then going to drive the research ideas.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm sorry, how, how does that actually shape out then?

A I think you have to identify very concretely the types of work that you are delivering, but I think we're starting to see, like, you know, Hollywood quality productions being made in a matter of hours. We're seeing, you know, to take a super concrete case, like machine translation, if humans are doing the machine translation compared to machines doing it, you go from hours to seconds on long form text, multi-page documents. And so for a lot of work, It's not like AI can do all of the work. Humans still need to ask the right question. They need to verify the information. They need to shape the tasks. But once the task is well-defined, the product are clear, like all the design considerations are fed into the prompt, you press the button and you've got an answer in seconds for something that used to take sometimes weeks and months.

AI assessment note: “to take a super concrete case, like machine translation... you go from hours to seconds”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What are the potential vulnerabilities, though, in an agent world?

A In terms of agents, you know, we, we worry a lot about hallucinations in LLMs. The parallel in agents is impersonation. So agents that come along and are essentially impersonating entities which they don't legitimately represent, and in doing so, taking actions on the behalf of these entities where they don't legitimately represent, whether it's, uh, infiltrating, you know, banking systems and, and, and so on. And so, I, I do think we have to be quite lucid about this, develop standards towards that, develop ways to, to test for that in a very rigorous way. There's ways to reduce that risk drastically. You run your agent, you know, completely cut off from the web. You're reducing your risk exposure significantly. Um, but then you lose access to some information. So depending on, depending on your use case, depending on what you actually need, there's different solutions that may be appropriate.

AI assessment note: “The parallel in agents is impersonation. So agents that come along and are essentially impersonating”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Great. Why is data becoming more expansive?

A Data comes in, in different forms. And, um, on the one hand, you know, the, the notion, the days of like having data labelers who can say this is a cat and this is a dog or somewhat over like the easy test the AI can do. So we're getting in a space where we need more specialized tasks. So, you know, imagine you're building AI for enterprise. There's a particular business logic. You need to make sure that you're catching the errors. You're gonna need someone with, like, deeper understanding of the tools, so that's more expensive talent to come in and actually prepare the data. Um, there's also a lot of data that's synthetic data, so when you're building agents, you need to build environments, and to build environments, you need some pretty creative folks who are going to build you, like, synthetic simulators. Um, we've seen this on the robot side for many years, people building robot simulators. Now you're building AI for enterprise, so you need to think of like, how are you going to simulate these work processes in a reasonably realistic way that the AI can train on that? And so that generation of environments and benchmarks and dynamic domains, uh, can, can be pretty expensive too.

AI assessment note: “we need more specialized tasks... that's more expensive talent to come in and actually prepare”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you think prompts and the way that we interact today with prompts and with chat largely is like the enduring interface for human engagement with AI?

A It's awfully limited, and, you know, prompts can mean a few different things, but, but the idea of, like, typing in a box, that to me is very limited, and we're gonna break out of that box already. We're seeing a lot of cases where voice is a lot more natural as an interface. I do expect we'll see, you know, gesture, eye gaze, these kinds of much more multimodal ways to interact with the, with the AI rather than just stick in that, in, in, in that prompt box. But language is incredibly powerful. So if you think of prompt as being more language as a way to express ideas and communicate with a machine, that's a powerful paradigm. I mean, as humans, we, so much of our communication is based on language. I don't think we're going to move away from that because it encodes information. You know, language words are symbols that encode so much information so efficiently. And so I don't think we're close to getting away from that.

AI assessment note: “the idea of, like, typing in a box, that to me is very limited”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q on the show suggest that evals are, to put it delicately, bullshit, uh, and that they don't actually mean anything anymore, and like, you know, humanity's last test, like, what does that really even mean? Uh, and we have these new tests that come up, What is this? And leaderboards, what is this? Is that fair, or do you think they actually serve a very effective utility to the ecosystem?

A I do think they, they are, um, really good indicators. So I think you do need to take evaluation seriously in terms of knowledge, but you shouldn't take them seriously in terms of the ultimate goals. So, you know, evaluation, and there's lots of different benchmarks and so on. You have to decide, like, what type of models or model are you building? What's the characteristics of your system? And then think of evaluations as, like, unit test for the performance of your system. I mean, software engineers will know what that is, right? Like, you run through that evaluation, and that gives you, like, a signal of how the system is doing in a particular dimension. But as we're building systems that are more and more general, do very specific tasks, you don't optimize for these, right? Like, I mean, we build AI systems that go into enterprise. None of our clients ask about, like, are you able to win the math Olympiad with this model? That's not what they care about. They care about bringing value to their business. Now, we're curious to know how well we do on math problems because it can be predictive of behavior on other things, but you don't obsess over specific benchmarks. You kind of look at the ROI in terms of what you're trying to build.

AI assessment note: “I do think they, they are, um, really good indicators.”

Answered raw tape D 4 · C 5 · P 5 · Cm 4 4.55

Q It's funny, this conversation has changed a lot of kind of previously held assumptions for me. When you think about what you did believe that you now have changed your mind on, what's most prescient?

A I'm a scientist that is happy to be proven wrong anytime, as long as there's new evidence. I'm genuinely curious to know. Other scientists are much more like holding on to very, very strong conviction. I have weak conviction, but very strong respect for the scientific method and, and, and rigor, experimental rigor, you know, theoretical rigor as well. Um, so there's a ton of things. I mean, I, ah, I mean, I used to be quite skeptical that neural networks were necessarily the ultimate solution to machine learning. I'd seen enough cycles of neural networks kind of peaking and, and, and then, um, being less, less useful, and I used to think every time you change the scale of the data, you know, you go from hundreds of examples to thousands, thousands, to hundreds of thousands, to millions of examples. Every time you change the size paradigm that neural networks were the first thing we tried, Because they're a universal function approximator. And then something else comes out that was better. And that was true for the previous generations. You know, some of you may remember SVMs as like being better than the neural networks in, in early 2000. And I, I seem to be quite wrong on this one. Like neural nets seem to be here to stay. Um, and the ability to do back propagation and gradient descent and all that seems to be a really powerful way to learn.

AI assessment note: “I used to be quite skeptical that neural networks were necessarily the ultimate solution”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Have we over invested in RL based methods at the expense of maybe, like, more scalable alternatives?

A Oh, I'm still super bullish on RL in that, like, the concept itself is so fundamental. You know, this idea of training through a system of rewards, of indicating what's valuable and what's not valuable through numerical values, like, that is so fundamental. It's not going away. Now, you know, where we're maybe getting a little bit ahead is thinking that just RL out of the box is gonna give us AGI. That part, a lot less so. You know, if you look at the curve of progress, RL is terribly inefficient, and so the amount of signal you need to get in order to really shape the behavior of a model is far from where we are today, and so we'll need to figure out how to, to really deal with this, with this learning efficiency problem.

AI assessment note: “where we're maybe getting a little bit ahead is thinking that just RL”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Security is a topic that we quite often glaze over, especially when investing in kind of application layer AI tools. What does no one know about AI security that people should know?

A With respect to AI security, I think there's a new front that's opening up with, um, the development of agents, and frankly, there's a lot we don't know yet in terms of the vulnerability of these systems. With LLMs, we're starting to get a better understanding. We've had, you know, quite a bit of red teaming exercise and jailbreaking and so on, and so people have identified different risk vectors, prompt injections, things like that, which are Vectors for malicious actors to interfere with the system. With AI agents, we haven't seen that. And one of the features of computer security in general is often, you know, it's a bit of a cat and mouse game, quite frankly. Like, there's a lot of ingenuity in terms of breaking into systems, and then you need a lot of ingenuity in terms of building defenses. And so, we just have to stay very active in that sense.

AI assessment note: “there's a new front that's opening up with, um, the development of agents”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.