The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jonathan Siddharth argument clarity score 4.4/5 from 37 exchanges on raw tape · average scores: directness 4.7 · coherence 4.6 · precision 4.3 · compression 3.9 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
37exchanges match
37on raw tape
2redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Why would you need a custom model? When you look at a lot of the customers that you mentioned there, for the ones where you have FDs who go in and build custom models, what is the reasoning around that? And is that a temporary moment in time? Or is that a permanent requirement from them for a certain reason?

A I think it's a permanent requirement. I'll give you an example. Um, let's pick an insurance company, right? Uh, and for insurance companies, two really important problems they have to solve are underwriting and claims processing. Let's pick underwriting, for example. So with underwriting, the problem statement is you might get multiple types of unstructured medical data. Uh, it could be Somebody taking a picture of their medical history on a, on their smartphone, or it could be some OCR data from somebody's medical history or data in PDFs, et cetera. And a human has to look at that person's medical information and then decide, is this person high risk, medium risk, or low risk? What medical conditions do they have? Do they have cardiovascular? Do they have renal? And how do you price insurance for somebody like this? Do you even take them on as a, As a client, if you're an insurance company, right? Now, this is a problem that, um, an LLM can solve really well with a human in the loop system. Now, you may not need a trillion parameter world model to do a task like this. Um, in fact, uh, there's lots of research that shows a smaller language model will actually be faster and more accurate at a task like this, um, than a giant world model. And the insurance company also may not want their data to go back to like a frontier, uh, model. So oftentimes in these cases, uh, what we woul…

AI assessment note: “I think it's a permanent requirement. I'll give you an example.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q the UBI and we're going to sit and write poetry and I'm like, I think that might be a little bit challenging. Um, When technology is not the moat, what is the moat? I had the founder of base, 44 on, um, and he said, 99% of code in the next year will be written by AI. Technology is no longer the moat. What is the moat in that world?

A I think one moat will be data-driven feedback loops. Um, for example, uh, one reason Google had such a great lead in search for a while, um, was these data-driven feedback loops that come from people using your product and generating data that gives you the algorithm developer A high quality gradient for which direction to step in, right? So PageRank, the importance of PageRank was known. The recipe for ranking search results was well known among Google, Yahoo, Microsoft, and a few others. Obviously people move around these companies all the time, but the advantage Google had was because everybody preferred Google and liked, liked working, liked that search engine. Um, the, You saw a much more representative set of queries. You had data from clickstream, from the clickstream of what results people were clicking on. That helps your algorithms improve at a much faster rate. I think data-driven feedback loops will be key for all types of enterprise applications also. Um, today, OpenAI and ChatGPT has a good data-driven Feedback loop. In enterprises, again, I think it's wide open, like whoever is deploying the right custom fine-tuned models and agents for specific workflows or roles or functions or companies, if you get in first and solve a customer's problem really well, you start getting that flywheel going where you will discover first where the models don't work well. And you w…

AI assessment note: “I think one moat will be data-driven feedback loops.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q trillion dollars that you said, it's ours. And if we don't, We operate in maybe a slightly larger software technology budget world, but by no means a world that we can have the valuations and the money that we have going in. When you look at that, are there any areas truly today where you're like, we have seen the full transition from human labor budgets to AI technology budgets?

A I think the transfer is pretty high in areas like, uh, customer support, um, copywriting, um, SEO, like some of these marketing related, uh, areas. As you would expect, the transfer is faster in these low risk to fail areas, like when it's, when it's relatively easy. Uh, but I encourage your listeners to Look up GDP Val, which is this paper by OpenAI, where they measured the impact of today's AI models in automating all types of economically valuable work. It's a lovely piece of research. I encourage everybody to read it. Where they did the study, where they looked at, I think, like, Uh, nine verticals and 44 occupations. They took like a very diverse sampling of different types of knowledge work. Everything from financial services to real estate to, um, uh, healthcare to, to law. And they took very specific occupations. And in those specific occupations, they took real tasks where like a real deliverable has to be produced. Imagine an engineer, like a civil engineer Creating a blueprint for a building they're about to build, or somebody who's at the set of a movie studio coming up with a schedule for how you organize your crews, like real work, right? And for coding, you can imagine like a real world software engineering project. Um, and they saw that today's models were quite good at achieving parity with, uh, the best human experts in that field. Uh, we were roughly at like,…

AI assessment note: “I think the transfer is pretty high in areas like, uh, customer support”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Do you agree then with the notion that if you believe in AGI, you cannot be investing in SaaS apps?

A SaaS as we know it, I think is over. I feel like quite a few SaaS apps, um, were built at a time when, um, software was relatively hard to build, uh, and complex to build. Imagine if you were building some customer support, um, software, some customer support bot. To build a company like that, you would have had to hire some Stanford PhDs in NLP. You'll collect data for six months. You'll have like, you'll use like a support vector machine or a neural network that'll kind of sort of work, and then you'll deploy it and you'll grind away for a while. There is a significant amount of capital that needs to be invested to get an app like that to work well. So it made sense for many companies to not bother doing that if that's not their core business. Let me just use like some third party Now, many of these AI applications are incredibly easy to build on top of these LLMs. So I feel like most companies will start building custom software super easily. I mean, we help companies build some of these custom apps. And the bar to create many of these apps significantly comes down. So that's one risk. Risk number one, companies do it themselves. Risk number two is that the, you, you get sonic boomed by the foundation model companies.

AI assessment note: “SaaS as we know it, I think is over.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I could talk to you all day. I do want to move into a quickfire answer. I say a short statement. You give me your immediate thoughts. What's one widely held belief about AI? That you think is wrong?

A I don't think we'll see rapid takeoff. I think we'll see incremental continuous improvement in AI. I actually think this is good for the world because if What we believe happens, which is all types of digital knowledge work gets automated. I think humanity needs time to prepare its workforce. I think we could use the extra time to upskill humans, to rethink education, to make sure there isn't massive job displacement. Um, I also think in the steady continuous improvement in AI models, There'll be value realized every step of the way. Unlike self-driving cars, I feel like people have this wrong model for AI that comes from self-driving cars where you get it, 99% of the way accurate, and you, if you can't solve the last one percent, they're not useful. AGI is not like that. I think when we automate the job of an underwriter or a claims processor or, or a CEO, there's incremental value that's unlocked for every percentage improvement in the models becoming more reliable. So I believe in slow and steady takeoff, and that's actually going to be great for the world.

AI assessment note: “I don't think we'll see rapid takeoff.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q How is that different? That's so interesting. So in the transition from chatbots to agents, how does the data required change with that transition?

A When you are training a chatbot, uh, you'd usually do a lot of, uh, SFT and RLHF. Uh, SFT where you're giving the model input prompts and output completions. You kind of teach the model to imitate experts. Uh, with RLHF, like you're basically, uh, teaching the model to produce responses that a human would tend to prefer. Uh, RLHF is used to train what's called a reward model. And then the model is trying to produce completions that, uh, give it a high reward, right? With agents, and let's define an agent. I mean, different people define agents in different ways. I would define an agent as, um, something that's capable of taking action, um, in the real world or in the physical world. Uh, something that's executing a multi-step workflow, calling different functions, Um, uh, the agent could be, um, operating a computer, uh, or making backend API calls to actually do stuff, right? Um, you might have an agent to file your taxes. You might have, uh, an agent to prepare your monthly financials. Um, so to train an agent, um, you would also want to teach the model how to do tool use. So you teach the model how to call other functions, how to use other applications to be more leveraged. Today, the dominant paradigm is reinforcement learning. Um, and oftentimes these agents are trained through reinforcement learning, where you'd build what's called an RL environment, which is like a mini …

AI assessment note: “you'd build what's called an RL environment, which is like a mini world model”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What does that mean? First mile schlep and last mile schlep?

A When I say first mile schlep, I mean, um, for that underwriting co-pilot example that I gave you for that insurance company, I painted a pretty rosy picture of like how you take this model, you fine tune it on your proprietary underwriting data. In the real world, it doesn't work that way. First mile schlep is, let's say I'm talking to the CEO of this insurance company or the CTO of this insurance company. They'll say, our data is a mess. It's in silos. It's super fragmented. Some of the data is in spreadsheets. Some of the data is in a file that Bob has, and Bob doesn't work here anymore, right? The data kind of is all over the place. You first have to acquire the data, convert the unstructured data into structured data into a format to fine tune LLMs. You might want to set up good infrastructure for evals. You'd want to create good evals for the models or agents. Um, you might want to build a workflow designed for partial autonomy. So this human underwriter that's about to, um, use this model to evaluate, um, these medical histories, you might want to build a cursor like interface for them so that they can work alongside the AI to do their job. Also training the humans in these new workflows. Um, you, uh, want to make sure you're collecting data the right way. Um, you'd, ah, the way, for example, we do deployments is, ah, like a tandem system where you'd have a human, ah, and…

AI assessment note: “When I say first mile schlep, I mean”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Why would you need a custom model? When you look at a lot of the customers that you mentioned there, for the ones where you have FDs who go in and build custom models, what is the reasoning around that? And is that a temporary moment in time? Or is that a permanent requirement from them for a certain reason?

A I think it's a permanent requirement. I'll give you an example. Um, let's pick an insurance company, right? Uh, and for insurance companies, two really important problems they have to solve are underwriting and claims processing. Let's pick underwriting, for example. So with underwriting, the problem statement is you might get multiple types of unstructured medical data. Uh, it could be Somebody taking a picture of their medical history on a, on their smartphone, or it could be some OCR data from somebody's medical history or data in PDFs, et cetera. And a human has to look at that person's medical information and then decide, is this person high risk, medium risk, or low risk? What medical conditions do they have? Do they have cardiovascular? Do they have renal? And how do you price insurance for somebody like this? Do you even take them on as a, As a client, if you're an insurance company, right? Now, this is a problem that, um, an LLM can solve really well with a human in the loop system. Now, you may not need a trillion parameter world model to do a task like this. Um, in fact, uh, there's lots of research that shows a smaller language model will actually be faster and more accurate at a task like this, um, than a giant world model. And the insurance company also may not want their data to go back to like a frontier, uh, model. So oftentimes in these cases, uh, what we woul…

AI assessment note: “I think it's a permanent requirement. I'll give you an example.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What happens in that world? If AI automates all types of knowledge work, What happens then?

A Three things that will happen. Uh, first, I think we are all going to be We'll all have the potential to be a hundred X more productive. Today, I'm able to run one company. Elon can run maybe like five companies, but in a world where I'm a hundred X more productive, maybe I'm able to run a hundred companies. Elon maybe runs 600 companies. Um, I think every human will just be so much more leveraged. Um, the nature of a job itself could change where there could be Today we are accustomed to the idea of like one person doing one job, um, but people could be doing multiple jobs at the same time. People could be running different companies at the same time. The second implication I think is it's going to be wonderful for entrepreneurship. So you, Harry, I think are going to be very happy because today, if I think of like for a lot of ideas, founders are intelligence constrained. Um, I think of being capital constrained as a form of being intelligence constrained. For example, if you pick a, a therapist who wants to start a mental health startup, Today, that founder would have to raise at least a few hundred K, if not a few million, to recruit some software engineers, maybe a marketing person or a growth person, maybe a product manager. But in a future where AGI exists, this person will recruit a marketing GPT, a software engineer GPT, a PM GPT, and get off the ground for a lot less …

AI assessment note: “Three things that will happen. Uh, first, I think we are all going to be”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q the UBI and we're going to sit and write poetry and I'm like, I think that might be a little bit challenging. Um, When technology is not the moat, what is the moat? I had the founder of base, 44 on, um, and he said, 99% of code in the next year will be written by AI. Technology is no longer the moat. What is the moat in that world?

A I think one moat will be data-driven feedback loops. Um, for example, uh, one reason Google had such a great lead in search for a while, um, was these data-driven feedback loops that come from people using your product and generating data that gives you the algorithm developer A high quality gradient for which direction to step in, right? So PageRank, the importance of PageRank was known. The recipe for ranking search results was well known among Google, Yahoo, Microsoft, and a few others. Obviously people move around these companies all the time, but the advantage Google had was because everybody preferred Google and liked, liked working, liked that search engine. Um, the, You saw a much more representative set of queries. You had data from clickstream, from the clickstream of what results people were clicking on. That helps your algorithms improve at a much faster rate. I think data-driven feedback loops will be key for all types of enterprise applications also. Um, today, OpenAI and ChatGPT has a good data-driven Feedback loop. In enterprises, again, I think it's wide open, like whoever is deploying the right custom fine-tuned models and agents for specific workflows or roles or functions or companies, if you get in first and solve a customer's problem really well, you start getting that flywheel going where you will discover first where the models don't work well. And you w…

AI assessment note: “I think one moat will be data-driven feedback loops.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Does every role need to go through that pathway of cursor for X before it goes to full autonomy, or are there some roles like customer service where it just goes to full autonomy?

A I think for some roles, uh, where the, where you can see that the models are quite good at matching humans, I think, uh, we don't need that intermediate step. They, there are certain roles where, um, by virtue of how the models are trained, where they are trained with, they're pre-trained with tokens on the internet, and then of course with, uh, talent from, um, research accelerators like Turing that's fine tuning the models, but there are certain types of roles where the tokens from the internet give it sufficient intelligence to do the job well. Customer support is an example, but if you pick Other roles, like for example, if you picked the role of, um, let's say an AI researcher, or you picked the role of a lawyer specializing in, um, venture capital financing, uh, it's possible there's not enough of those tokens on the internet. So the models will be relatively weak there out of the box. And also the way, I don't know, the way a Wilson Sincini does financing might look different from the way A Cooley does it. Maybe they have their own way. So you might want to fine tune them on your own proprietary data, distill the proprietary intelligence of humans working there. Um, so, so I would, so, so for those things, like you may need to do some fine tuning, the models may not work very well out of the box.

AI assessment note: “I think for some roles... we don't need that intermediate step.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Do you agree then with the notion that if you believe in AGI, you cannot be investing in SaaS apps?

A SaaS as we know it, I think is over. I feel like quite a few SaaS apps, um, were built at a time when, um, software was relatively hard to build, uh, and complex to build. Imagine if you were building some customer support, um, software, some customer support bot. To build a company like that, you would have had to hire some Stanford PhDs in NLP. You'll collect data for six months. You'll have like, you'll use like a support vector machine or a neural network that'll kind of sort of work, and then you'll deploy it and you'll grind away for a while. There is a significant amount of capital that needs to be invested to get an app like that to work well. So it made sense for many companies to not bother doing that if that's not their core business. Let me just use like some third party Now, many of these AI applications are incredibly easy to build on top of these LLMs. So I feel like most companies will start building custom software super easily. I mean, we help companies build some of these custom apps. And the bar to create many of these apps significantly comes down. So that's one risk. Risk number one, companies do it themselves. Risk number two is that the, you, you get sonic boomed by the foundation model companies.

AI assessment note: “SaaS as we know it, I think is over.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What did you believe that you now no longer believe?

A I used to believe that, um, to build a Enduring, valuable company. You hire a strong, ah, exec team and operate with a lot of leverage. Basically hire strong people, hire great people and get out of the way. I used to believe that. Now I believe you hire great people and work really closely with them and their directs and their directs and their directs and Get as close to the ground as you can, where ground truth usually exists with the customers. The next step to customers is the engineers writing code and the salespeople talking to your customers. So now I believe in being, basically I used to like, for lack of a better word, follow the org chart a little bit. And this was also part of one of my learnings from Elon's biography is that he was so hands-on, like he would be like, Walking the factory floor and asking an engineer why this door in the model three has three bolts instead of, ah, instead of maybe two, right? And it is a different way to operate, like where you are in the details of the most important things that matter, completely working in like a flat structure and just operating as close to the, to the ground truth as you can. Um, and generally, like, I feel like in the early days of starting Turing, like, I, now that I think about it, I may have had a, um, subconscious desire to be liked. I think I must have. Now I don't. Like the, now I just think about just do…

AI assessment note: “I used to believe that... Now I believe you hire great people and work”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What happens in that world? If AI automates all types of knowledge work, What happens then?

A Three things that will happen. Uh, first, I think we are all going to be We'll all have the potential to be a hundred X more productive. Today, I'm able to run one company. Elon can run maybe like five companies, but in a world where I'm a hundred X more productive, maybe I'm able to run a hundred companies. Elon maybe runs 600 companies. Um, I think every human will just be so much more leveraged. Um, the nature of a job itself could change where there could be Today we are accustomed to the idea of like one person doing one job, um, but people could be doing multiple jobs at the same time. People could be running different companies at the same time. The second implication I think is it's going to be wonderful for entrepreneurship. So you, Harry, I think are going to be very happy because today, if I think of like for a lot of ideas, founders are intelligence constrained. Um, I think of being capital constrained as a form of being intelligence constrained. For example, if you pick a, a therapist who wants to start a mental health startup, Today, that founder would have to raise at least a few hundred K, if not a few million, to recruit some software engineers, maybe a marketing person or a growth person, maybe a product manager. But in a future where AGI exists, this person will recruit a marketing GPT, a software engineer GPT, a PM GPT, and get off the ground for a lot less …

AI assessment note: “Three things that will happen. Uh, first, I think we are all going to be”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q of cooling period, which everyone kind of suggests and thinks that we're going to go through in the next six to 18 months, which is, as I said, it doesn't hit the revenues that we said it would, and kind of the AI bubble kind of deflates slowly. To what extent do you think that's possible, or we'll see this continuing gradual increase as we kind of touched on there?

A I don't see an AI bubble. Like, I feel like these models are incredibly powerful today. Like, GPT-V is like, fucking awesome. I don't know what people were talking about when they're talking about, you know, I know there was some chatter. I think we've just gotten used to magic, and it's like the, some of the, uh, I feel like, A, these models are incredibly powerful today, and they're the worst they'll ever be. They're only going to keep, keep improving. Um, and I say that about, um, the Gemini Pro models, the Grock models, the Claude models, like these models are amazing. And there's a very significant model capability overhang. By that, what I mean is the models are capable of X, but what we are getting out of the models is X minus Delta. Um, there is, uh, with the right agentic scaffold, uh, around these models, In terms of the right system prompts, the right user prompts, giving the models access to the right context, teaching the models how to acquire additional context, teaching the models how to use the right internal tools. There is significant amount of capability that can be unlocked with today's models. For example, you, Harry, I imagine when you do an interview with, uh, with somebody, One of the things you probably do is you apply your secret sauce to pull out the right clips from the interviews, like what to highlight, what are the catchphrases, what will sort of …

AI assessment note: “I don't see an AI bubble. Like, I feel like these models are incredibly powerful”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q around that because I'll get in trouble for it. Um, I think we grossly overestimate the intelligence of the general population, and I know that sounds incredibly arrogant, but most people say, actually, a lot of people just don't want to work And are not at the level of recruiting GPT's assistants. Do you not worry that it will widen the chasm between those that have and those that haven't?

A I'm an optimist, and I think the opposite will happen because what we're really doing when we're training super intelligence is you'll be basically training intelligence as an API. And what's the alternative to that? It's like hiring a human to provide you with that intelligence. That human is quite expensive, right? And if anything, that creates an even broader gap between the haves and the have nots. Whereas for 20 dollars a month, If you have, if you had access to the smartest experts in coding, in STEM, in sales, in marketing, I feel like more people will be able to start companies and produce active, ah, like valuable work. I also think, I believe that when we have access to super intelligence, I firmly believe this. We are not all going to chill out on a beach somewhere and contemplate what do we do next? I think we're going to I mean, we humans, like we are tool builders. We are problem solvers. We'll solve problems at higher and higher levels of abstraction. And I feel like in a world where we have AGI, we'll just solve much more exciting problems. Maybe we'll cure diseases, reverse aging, maybe go to the stars. There's all sorts of fun things we'll do. I don't think we'll be bored.

AI assessment note: “I'm an optimist, and I think the opposite will happen”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What did you believe that you now no longer believe?

A I used to believe that, um, to build a Enduring, valuable company. You hire a strong, ah, exec team and operate with a lot of leverage. Basically hire strong people, hire great people and get out of the way. I used to believe that. Now I believe you hire great people and work really closely with them and their directs and their directs and their directs and Get as close to the ground as you can, where ground truth usually exists with the customers. The next step to customers is the engineers writing code and the salespeople talking to your customers. So now I believe in being, basically I used to like, for lack of a better word, follow the org chart a little bit. And this was also part of one of my learnings from Elon's biography is that he was so hands-on, like he would be like, Walking the factory floor and asking an engineer why this door in the model three has three bolts instead of, ah, instead of maybe two, right? And it is a different way to operate, like where you are in the details of the most important things that matter, completely working in like a flat structure and just operating as close to the, to the ground truth as you can. Um, and generally, like, I feel like in the early days of starting Turing, like, I, now that I think about it, I may have had a, um, subconscious desire to be liked. I think I must have. Now I don't. Like the, now I just think about just do…

AI assessment note: “I used to believe that... Now I believe you hire great people and work really closely”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q around that because I'll get in trouble for it. Um, I think we grossly overestimate the intelligence of the general population, and I know that sounds incredibly arrogant, but most people say, actually, a lot of people just don't want to work And are not at the level of recruiting GPT's assistants. Do you not worry that it will widen the chasm between those that have and those that haven't?

A I'm an optimist, and I think the opposite will happen because what we're really doing when we're training super intelligence is you'll be basically training intelligence as an API. And what's the alternative to that? It's like hiring a human to provide you with that intelligence. That human is quite expensive, right? And if anything, that creates an even broader gap between the haves and the have nots. Whereas for 20 dollars a month, If you have, if you had access to the smartest experts in coding, in STEM, in sales, in marketing, I feel like more people will be able to start companies and produce active, ah, like valuable work. I also think, I believe that when we have access to super intelligence, I firmly believe this. We are not all going to chill out on a beach somewhere and contemplate what do we do next? I think we're going to I mean, we humans, like we are tool builders. We are problem solvers. We'll solve problems at higher and higher levels of abstraction. And I feel like in a world where we have AGI, we'll just solve much more exciting problems. Maybe we'll cure diseases, reverse aging, maybe go to the stars. There's all sorts of fun things we'll do. I don't think we'll be bored.

AI assessment note: “I'm an optimist, and I think the opposite will happen”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q of cooling period, which everyone kind of suggests and thinks that we're going to go through in the next six to 18 months, which is, as I said, it doesn't hit the revenues that we said it would, and kind of the AI bubble kind of deflates slowly. To what extent do you think that's possible, or we'll see this continuing gradual increase as we kind of touched on there?

A I don't see an AI bubble. Like, I feel like these models are incredibly powerful today. Like, GPT-V is like, fucking awesome. I don't know what people were talking about when they're talking about, you know, I know there was some chatter. I think we've just gotten used to magic, and it's like the, some of the, uh, I feel like, A, these models are incredibly powerful today, and they're the worst they'll ever be. They're only going to keep, keep improving. Um, and I say that about, um, the Gemini Pro models, the Grock models, the Claude models, like these models are amazing. And there's a very significant model capability overhang. By that, what I mean is the models are capable of X, but what we are getting out of the models is X minus Delta. Um, there is, uh, with the right agentic scaffold, uh, around these models, In terms of the right system prompts, the right user prompts, giving the models access to the right context, teaching the models how to acquire additional context, teaching the models how to use the right internal tools. There is significant amount of capability that can be unlocked with today's models. For example, you, Harry, I imagine when you do an interview with, uh, with somebody, One of the things you probably do is you apply your secret sauce to pull out the right clips from the interviews, like what to highlight, what are the catchphrases, what will sort of …

AI assessment note: “I don't see an AI bubble. Like, I feel like these models are incredibly powerful”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q If you were to invest in companies in your space, where would you invest?

A Probably in robotics or embodied AI. The vertical stuff, like, I mean, we are scaling up pretty massively in generating data for different verticals. So I don't, I don't see that as like a big white space. But I think everybody is relatively early with robotics. Um, and robotics is such a vast realm that there could be interesting things to do there. One way I see the space, uh, Harry is, um, Again, think of like a, like these three dimensions. The first dimension being the type of intelligence that you're baking into the models. That could be in coding, in STEM, in functional expertise like sales, marketing, software engineering, or vertical expertise like healthcare, legal, finance, et cetera. So the first dimension is the type of intelligence. I do a cross product of that with, um, the modality. So audio, video, image, computer use. So that's multimodality. That's the second dimension. The third dimension is multilinguality, like different languages, right? And, um, the fourth dimension is different learning paradigms, like imitation learning, reinforcement learning, pre-training as well, which is unsupervised learning. And all of those may require different platforms to be built. Like we've had to adapt our platform For imitation learning, for reinforcement learning, for multimodality. Um, so I feel like in this matrix, there's like all sorts of, um, all sorts of New opport…

AI assessment note: “Probably in robotics or embodied AI.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I could talk to you all day. I do want to move into a quickfire answer. I say a short statement. You give me your immediate thoughts. What's one widely held belief about AI? That you think is wrong?

A I don't think we'll see rapid takeoff. I think we'll see incremental continuous improvement in AI. I actually think this is good for the world because if What we believe happens, which is all types of digital knowledge work gets automated. I think humanity needs time to prepare its workforce. I think we could use the extra time to upskill humans, to rethink education, to make sure there isn't massive job displacement. Um, I also think in the steady continuous improvement in AI models, There'll be value realized every step of the way. Unlike self-driving cars, I feel like people have this wrong model for AI that comes from self-driving cars where you get it, 99% of the way accurate, and you, if you can't solve the last one percent, they're not useful. AGI is not like that. I think when we automate the job of an underwriter or a claims processor or, or a CEO, there's incremental value that's unlocked for every percentage improvement in the models becoming more reliable. So I believe in slow and steady takeoff, and that's actually going to be great for the world.

AI assessment note: “I don't think we'll see rapid takeoff. I think we'll see incremental continuous improvement”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q The world seems to be moving to closed models. Is that good or bad?

A I think it depends on the Uh, application. Um, I feel like in the, firstly, in enterprises, it's often a mix of between proprietary and closed models. I wouldn't, sorry, between closed and open models. Um, we do see demand from enterprises that want either. Um, the closed models often are easier to Uh, get started with. Uh, but there are some cases where enterprises prefer open models to, um, for cost, customizability. Um, and I'm talking about the small language model regime between the half a billion parameters to ten billion parameters. I worry a little about frontier models. I feel like for frontier models, there is some value in keeping some of the technology closed. The, just because of how powerful they are, and, um, I feel like the US labs are extremely responsible and safety conscious in how they think about training these models, deploying them.

AI assessment note: “I think it depends on the Uh, application.”

Answered raw tape D 4 · C 5 · P 5 · Cm 4 4.55

Q What does that mean? First mile schlep and last mile schlep?

A When I say first mile schlep, I mean, um, for that underwriting co-pilot example that I gave you for that insurance company, I painted a pretty rosy picture of like how you take this model, you fine tune it on your proprietary underwriting data. In the real world, it doesn't work that way. First mile schlep is, let's say I'm talking to the CEO of this insurance company or the CTO of this insurance company. They'll say, our data is a mess. It's in silos. It's super fragmented. Some of the data is in spreadsheets. Some of the data is in a file that Bob has, and Bob doesn't work here anymore, right? The data kind of is all over the place. You first have to acquire the data, convert the unstructured data into structured data into a format to fine tune LLMs. You might want to set up good infrastructure for evals. You'd want to create good evals for the models or agents. Um, you might want to build a workflow designed for partial autonomy. So this human underwriter that's about to, um, use this model to evaluate, um, these medical histories, you might want to build a cursor like interface for them so that they can work alongside the AI to do their job. Also training the humans in these new workflows. Um, you, uh, want to make sure you're collecting data the right way. Um, you'd, ah, the way, for example, we do deployments is, ah, like a tandem system where you'd have a human, ah, and…

AI assessment note: “When I say first mile schlep, I mean”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q How is that different? That's so interesting. So in the transition from chatbots to agents, how does the data required change with that transition?

A When you are training a chatbot, uh, you'd usually do a lot of, uh, SFT and RLHF. Uh, SFT where you're giving the model input prompts and output completions. You kind of teach the model to imitate experts. Uh, with RLHF, like you're basically, uh, teaching the model to produce responses that a human would tend to prefer. Uh, RLHF is used to train what's called a reward model. And then the model is trying to produce completions that, uh, give it a high reward, right? With agents, and let's define an agent. I mean, different people define agents in different ways. I would define an agent as, um, something that's capable of taking action, um, in the real world or in the physical world. Uh, something that's executing a multi-step workflow, calling different functions, Um, uh, the agent could be, um, operating a computer, uh, or making backend API calls to actually do stuff, right? Um, you might have an agent to file your taxes. You might have, uh, an agent to prepare your monthly financials. Um, so to train an agent, um, you would also want to teach the model how to do tool use. So you teach the model how to call other functions, how to use other applications to be more leveraged. Today, the dominant paradigm is reinforcement learning. Um, and oftentimes these agents are trained through reinforcement learning, where you'd build what's called an RL environment, which is like a mini …

AI assessment note: “in that RL environment, you have input prompts, output verifiers”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Does every role need to go through that pathway of cursor for X before it goes to full autonomy, or are there some roles like customer service where it just goes to full autonomy?

A I think for some roles, uh, where the, where you can see that the models are quite good at matching humans, I think, uh, we don't need that intermediate step. They, there are certain roles where, um, by virtue of how the models are trained, where they are trained with, they're pre-trained with tokens on the internet, and then of course with, uh, talent from, um, research accelerators like Turing that's fine tuning the models, but there are certain types of roles where the tokens from the internet give it sufficient intelligence to do the job well. Customer support is an example, but if you pick Other roles, like for example, if you picked the role of, um, let's say an AI researcher, or you picked the role of a lawyer specializing in, um, venture capital financing, uh, it's possible there's not enough of those tokens on the internet. So the models will be relatively weak there out of the box. And also the way, I don't know, the way a Wilson Sincini does financing might look different from the way A Cooley does it. Maybe they have their own way. So you might want to fine tune them on your own proprietary data, distill the proprietary intelligence of humans working there. Um, so, so I would, so, so for those things, like you may need to do some fine tuning, the models may not work very well out of the box.

AI assessment note: “for some roles... we don't need that intermediate step. Customer support is an example”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q If you were to invest in companies in your space, where would you invest?

A Probably in robotics or embodied AI. The vertical stuff, like, I mean, we are scaling up pretty massively in generating data for different verticals. So I don't, I don't see that as like a big white space. But I think everybody is relatively early with robotics. Um, and robotics is such a vast realm that there could be interesting things to do there. One way I see the space, uh, Harry is, um, Again, think of like a, like these three dimensions. The first dimension being the type of intelligence that you're baking into the models. That could be in coding, in STEM, in functional expertise like sales, marketing, software engineering, or vertical expertise like healthcare, legal, finance, et cetera. So the first dimension is the type of intelligence. I do a cross product of that with, um, the modality. So audio, video, image, computer use. So that's multimodality. That's the second dimension. The third dimension is multilinguality, like different languages, right? And, um, the fourth dimension is different learning paradigms, like imitation learning, reinforcement learning, pre-training as well, which is unsupervised learning. And all of those may require different platforms to be built. Like we've had to adapt our platform For imitation learning, for reinforcement learning, for multimodality. Um, so I feel like in this matrix, there's like all sorts of, um, all sorts of New opport…

AI assessment note: “Probably in robotics or embodied AI.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q The world seems to be moving to closed models. Is that good or bad?

A I think it depends on the Uh, application. Um, I feel like in the, firstly, in enterprises, it's often a mix of between proprietary and closed models. I wouldn't, sorry, between closed and open models. Um, we do see demand from enterprises that want either. Um, the closed models often are easier to Uh, get started with. Uh, but there are some cases where enterprises prefer open models to, um, for cost, customizability. Um, and I'm talking about the small language model regime between the half a billion parameters to ten billion parameters. I worry a little about frontier models. I feel like for frontier models, there is some value in keeping some of the technology closed. The, just because of how powerful they are, and, um, I feel like the US labs are extremely responsible and safety conscious in how they think about training these models, deploying them.

AI assessment note: “I think it depends on the Uh, application.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q trillion dollars that you said, it's ours. And if we don't, We operate in maybe a slightly larger software technology budget world, but by no means a world that we can have the valuations and the money that we have going in. When you look at that, are there any areas truly today where you're like, we have seen the full transition from human labor budgets to AI technology budgets?

A I think the transfer is pretty high in areas like, uh, customer support, um, copywriting, um, SEO, like some of these marketing related, uh, areas. As you would expect, the transfer is faster in these low risk to fail areas, like when it's, when it's relatively easy. Uh, but I encourage your listeners to Look up GDP Val, which is this paper by OpenAI, where they measured the impact of today's AI models in automating all types of economically valuable work. It's a lovely piece of research. I encourage everybody to read it. Where they did the study, where they looked at, I think, like, Uh, nine verticals and 44 occupations. They took like a very diverse sampling of different types of knowledge work. Everything from financial services to real estate to, um, uh, healthcare to, to law. And they took very specific occupations. And in those specific occupations, they took real tasks where like a real deliverable has to be produced. Imagine an engineer, like a civil engineer Creating a blueprint for a building they're about to build, or somebody who's at the set of a movie studio coming up with a schedule for how you organize your crews, like real work, right? And for coding, you can imagine like a real world software engineering project. Um, and they saw that today's models were quite good at achieving parity with, uh, the best human experts in that field. Uh, we were roughly at like,…

AI assessment note: “transfer is pretty high in areas like, uh, customer support, um, copywriting, um, SEO”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Of the eight largest providers, do they not spend with all of you?

A They spend with a handful of companies. Uh, they do that for, uh, to have some level of, um, uh, resilience. Um, and I imagine there's some price benefits to having more than one person that they could work with, but I think the resilience piece is important. I mean, we know what happened with the, you know, you know, when the scale investment happened again, like the, it, it was, um, Uh, the labs did benefit from having other partners that they could work with. I would say it's a small handful. It's a small handful that are trusted. And of course, there's a, there's probably a giant pool of smaller startups, but it's a small handful of big companies in the space.

AI assessment note: “They spend with a handful of companies.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Do we lose the phone as the interface to this world? We obviously see Sam and Johnny. I've, you know, there's rumors of pendants and some hardware devices. I'm not asking you to comment on that. I'm just saying, does the phone become, still remain the primary interface and design device?

A We'll have some type of a device, um, that is, that we'll carry that's, um, always on and processing multimodal tokens. For example, as I'm talking to you, If I were to envision my perfect device, it would be something that has, it should have cameras. So maybe it's a wearable as a glass, uh, or something that I'm, uh, having on me that's processing visual input because I want to be able to read your body language. I might have like, um, AirPod-like thing in my ear that's whispering to me that maybe says, Jonathan, as you were talking about multimodality, Harry seemed Less interested. His body cues suggest that he was losing interest. But when we were talking about AR, he perked up. So those types of feedback and cues I think would be good. So I envision a device that, um, I think of it in terms of sensors and effectors. In terms of sensors, like obviously it has to be listening to stuff. It has to be seeing stuff. But in terms of effectors, it'll probably also be speaking in my ear. Um, ideally it should be something that you can talk to and have it do things later. For example, I might say, remind me to follow up with Harry on that idea for, um, uh, using Turing to automate clip generation, right? Like the, so, so it has to like remember that and come back later. So I do think there'll be all sorts of new devices and Glasses, hearing, like these AirPod type devices seem obvio…

AI assessment note: “I do think there'll be all sorts of new devices and Glasses”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.