The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Shawn Wang argument clarity score 3.9/5 from 15 exchanges on raw tape · average scores: directness 4.2 · coherence 3.8 · precision 3.7 · compression 3.5 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
28exchanges match
28on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q was, like, quite a bit higher. So just, I think that kind of progress over time is what we're most interested in seeing is, you know, are models getting worse? Model's getting better? Are people still loving PG vector? Do people still love Mongo? You know, stuff like that. That I think is the most interesting thing, so. Do you two have any questions that you think we should ask?

A Um, off the bat, like, it's, it seems like you're very, uh, language model focused. Um, um, you know, I think that there's an increasing, um, interest in multi-modality, um, in, in AI. Um, and I don't really know how that is going to manifest. Um, obviously, GPT, Four Vision, as well as, um, Gemini both have multi-world capabilities. Uh, there's a smaller subset of open source models that have, Um, multimodal features as well. Like I, we just released an episode today, uh, talking about IdaFix from HuggingFace. And, uh, yeah, so I think I, I think I w I would like to understand how people are adopting or adapting to the different modalities that are now coming online for them. Um, what their demand is relative to, uh, for like, let's say generative images versus, um, you know, like just visual comprehension versus, um, audio versus, uh, text to speech. Like, uh, what do they want? What do they need? And what's the sort of, um, yeah. Forced, like, stack ranked, um, preference order. Yeah. Yeah.

AI assessment note: “I would like to understand how people are adopting or adapting to the different modalities”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do we have to try to fit large language models, the compute that it requires, into kind of the human box? No, we don't. Why does it have to be, well, if it's not exactly how humans learn, then it's not the right way.

A No, I absolutely agree with what you're saying there, which is, uh, I call this a sour lesson, right? Like, every time you try to make a human analogy to machines, you probably fail because machines develop very differently from humans. Uh, so I, no, I, I agree with that, and I also think that it can still be an alien form of more efficient learning. Doesn't matter. We just know it is super inefficient. Like, that, that is something that we know is a unmitigated negative. So let's make it more efficient and better, uh, and that means that, you know, if, in order, instead of 2000 examples to learn one thing, what about 20 examples? What about two examples? Um, and that scales a lot more, um, and that means, You know, we can, we can get, we can actually get to a point with continual learning that we can actually have, ah, agents that adapt and, and build up a real world model. Otherwise, we're always stuck to the pre-trained, post-trained paradigm that is probably hitting some kind of limit right now.

AI assessment note: “every time you try to make a human analogy to machines, you probably fail”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q generation Opus, uh, GPT-Five. The, the amount of value that can be extracted from those models still seems, at least to me, to be, uh, critical. And, and now we have a whole new A generation of model that we're even going to get more Model LiveWay hang from. What do you, do you agree with that? Do you think you need to keep building the tools around the model?

A AI engineer exists in the white surface area between the peak capability and deploying it everywhere else, right? So the more model research peaks and spikes capabilities in one domain, but it's not evenly distributed in all products yet, that's where engineers have a job. Forever, basically. Um, so I'm very pro that. Um, I, I, I think capability over overhang will exist for a long time. I think it does keep, have waves of consolidation where you're like, actually all this stuff I built out, I don't need it anymore because the next model has got it from just like a single prompt. So I'm going to throw it out. But like, we do spring cleaning every now and then, like that's normal. And we build it up on previous gen model assumptions that then go away. And I don't think we should feel any attachment to the code. At the end of the day, we're all just trying to like serve customers better, do work cheaper, faster, easier. Yeah.

AI assessment note: “so I'm very pro that. Um, I think capability overhang will exist for a long time.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do we have to try to fit large language models, the compute that it requires, into kind of the human box? No, we don't. Why does it have to be, well, if it's not exactly how humans learn, then it's not the right way.

A No, I absolutely agree with what you're saying there, which is, uh, I call this a sour lesson, right? Like, every time you try to make a human analogy to machines, you probably fail because machines develop very differently from humans. Uh, so I, no, I, I agree with that, and I also think that it can still be an alien form of more efficient learning. Doesn't matter. We just know it is super inefficient. Like, that, that is something that we know is a unmitigated negative. So let's make it more efficient and better, uh, and that means that, you know, if, in order, instead of 2000 examples to learn one thing, what about 20 examples? What about two examples? Um, and that scales a lot more, um, and that means, You know, we can, we can get, we can actually get to a point with continual learning that we can actually have, ah, agents that adapt and, and build up a real world model. Otherwise, we're always stuck to the pre-trained, post-trained paradigm that is probably hitting some kind of limit right now.

AI assessment note: “every time you try to make a human analogy to machines, you probably fail”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q generation Opus, uh, GPT-Five. The, the amount of value that can be extracted from those models still seems, at least to me, to be, uh, critical. And, and now we have a whole new A generation of model that we're even going to get more Model LiveWay hang from. What do you, do you agree with that? Do you think you need to keep building the tools around the model?

A AI engineer exists in the white surface area between the peak capability and deploying it everywhere else, right? So the more model research peaks and spikes capabilities in one domain, but it's not evenly distributed in all products yet, that's where engineers have a job. Forever, basically. Um, so I'm very pro that. Um, I, I, I think capability over overhang will exist for a long time. I think it does keep, have waves of consolidation where you're like, actually all this stuff I built out, I don't need it anymore because the next model has got it from just like a single prompt. So I'm going to throw it out. But like, we do spring cleaning every now and then, like that's normal. And we build it up on previous gen model assumptions that then go away. And I don't think we should feel any attachment to the code. At the end of the day, we're all just trying to like serve customers better, do work cheaper, faster, easier. Yeah.

AI assessment note: “I'm very pro that. Um, I think capability overhang will exist for a long time.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q know, one of the things they said was we haven't seen anything that's kind of eyeopening to see people going to 200 dollar tier on this sort of thing. Haven't seen anything else like that in the space. Cause I, I think this is very new because of the new model capabilities, right? Where people, you know, it makes sense. Like you're willing to pay more money for this stuff.

A So this is something I've talked about before in terms of matching the dollar amount of spend to the capabilities of the AIs. The chart that I published in the past was, you know, OpenAI has like five levels of AGI-ness and level, level one is sort of like a chat boss. Level two is reasoning. Level three is agents. Four is organizations. Five is something super, superhuman. I don't remember what the exact levels are, but each, each, you can sort of each match each of them with like tiers. Like, uh, 20 dollars is like the ChatGPT tier. 200 dollars is where you're at. 2000 dollars is higher to 20,200 thousand, right? Like, you can see levels where it makes sense. I think Brightwave is also there, by the way. Like, I don't know what Brightwave charges, but it's higher, right, than a ChatGPT. And like, you have to deliver more value for that. But you, you can do it now.

AI assessment note: “matching the dollar amount of spend to the capabilities of the AIs”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q was, like, quite a bit higher. So just, I think that kind of progress over time is what we're most interested in seeing is, you know, are models getting worse? Model's getting better? Are people still loving PG vector? Do people still love Mongo? You know, stuff like that. That I think is the most interesting thing, so. Do you two have any questions that you think we should ask?

A Um, off the bat, like, it's, it seems like you're very, uh, language model focused. Um, um, you know, I think that there's an increasing, um, interest in multi-modality, um, in, in AI. Um, and I don't really know how that is going to manifest. Um, obviously, GPT, Four Vision, as well as, um, Gemini both have multi-world capabilities. Uh, there's a smaller subset of open source models that have, Um, multimodal features as well. Like I, we just released an episode today, uh, talking about IdaFix from HuggingFace. And, uh, yeah, so I think I, I think I w I would like to understand how people are adopting or adapting to the different modalities that are now coming online for them. Um, what their demand is relative to, uh, for like, let's say generative images versus, um, you know, like just visual comprehension versus, um, audio versus, uh, text to speech. Like, uh, what do they want? What do they need? And what's the sort of, um, yeah. Forced, like, stack ranked, um, preference order. Yeah. Yeah.

AI assessment note: “I would like to understand how people are adopting or adapting to the different modalities”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q now, right? Into latent space maybe, right? Um, then, you know, then having access to that institutional knowledge of how things actually happen, the actual decisions in big organizations, Feels really valuable. And so those are the four quarters. And then I can talk about kind of more context graph specific stuff. But first of all, do you agree with the four? Do you see a fifth or a sixth?

A I feel like the, the, the three out of the four is obvious. Sorry. I strongly agree with. And then the, the fourth one is agentic memory actually, where I'm like, that's a little weaker. That's a little less established. That's a little smaller. It's like the, the one, two, and four are very strong, very large. Categories where I know exactly how to architect it and everything. The third one is like the, uh, I don't know, something, something memory. And I feel like there could be more. And when I would think about two by two, right? OLAP OLTP is great. That's like a dimension of how wide you're querying, what volume of transactions you're doing. The other dimension should be ideally, the other axis should be orthogonal. And I don't necessarily know what that axis is. Uh, and so it doesn't exactly fit a normal two by two, which usually means that maybe there's a third dimension that's kind of being sort of mixed into the, into the, into the equation here. Because agentic memory, let's call it, is probably mostly personal, maybe some organizational, whereas context graphs is fully organizational. Um, uh, and yeah, so, so that's my, that's my initial reactions, but I'm happy to just only talk about context graphs, unless, unless you have more sort of.

AI assessment note: “three out of the four is obvious. Sorry. I strongly agree with.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q now, right? Into latent space maybe, right? Um, then, you know, then having access to that institutional knowledge of how things actually happen, the actual decisions in big organizations, Feels really valuable. And so those are the four quarters. And then I can talk about kind of more context graph specific stuff. But first of all, do you agree with the four? Do you see a fifth or a sixth?

A I feel like the, the, the three out of the four is obvious. Sorry. I strongly agree with. And then the, the fourth one is agentic memory actually, where I'm like, that's a little weaker. That's a little less established. That's a little smaller. It's like the, the one, two, and four are very strong, very large. Categories where I know exactly how to architect it and everything. The third one is like the, uh, I don't know, something, something memory. And I feel like there could be more. And when I would think about two by two, right? OLAP OLTP is great. That's like a dimension of how wide you're querying, what volume of transactions you're doing. The other dimension should be ideally, the other axis should be orthogonal. And I don't necessarily know what that axis is. Uh, and so it doesn't exactly fit a normal two by two, which usually means that maybe there's a third dimension that's kind of being sort of mixed into the, into the, into the equation here. Because agentic memory, let's call it, is probably mostly personal, maybe some organizational, whereas context graphs is fully organizational. Um, uh, and yeah, so, so that's my, that's my initial reactions, but I'm happy to just only talk about context graphs, unless, unless you have more sort of.

AI assessment note: “I feel like the, the, the three out of the four is obvious.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q How did it actually do it? I'm actually curious.

A Oh, usually it just like takes a screenshot and then it reads the screenshot by vision. So this is what I do for my, my zoom upload thing, right? Because I, I have paper club sessions that I need to upload to zoom and I wanted to automatically, uh, title them and do show notes and everything. So it just takes screenshots and try to try its best. It wouldn't probably benefit from transcribing, which it's doing by, it's operating my pure vision now, but it's good enough. Yeah. And then I, I do have to call, uh, out to Nano Banana to do images. So unless you guys do images for me. Uh, I have to call other people with images.

AI assessment note: “usually it just like takes a screenshot and then it reads the screenshot by vision”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q How did it actually do it? I'm actually curious.

A Oh, usually it just like takes a screenshot and then it reads the screenshot by vision. So this is what I do for my, my zoom upload thing, right? Because I, I have paper club sessions that I need to upload to zoom and I wanted to automatically, uh, title them and do show notes and everything. So it just takes screenshots and try to try its best. It wouldn't probably benefit from transcribing, which it's doing by, it's operating my pure vision now, but it's good enough. Yeah. And then I, I do have to call, uh, out to Nano Banana to do images. So unless you guys do images for me. Uh, I have to call other people with images.

AI assessment note: “usually it just like takes a screenshot and then it reads the screenshot by vision”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yeah. Interesting observation. What, what, uh, what are some of your ideas or thoughts? You have

A some specific... Riverside is the closest that has come to it. Descript is number two. Descript bought a Riverside competitor, and, and, uh, as far as I can tell, it's not very, it's not been very successful. Descript just, like, has, like, has a very, very good niche, very, very good editing angle, and then just has, hasn't done anything interesting since then. Although Underlord is good. It's not great. Like, your chapterization is better than the scripts. Again, like, they should be able to beat you. They're not. And Riverside is good, also. Very, very good. Very, very, very good. Like, so we, we actually recently started a second series of podcasts within Latent Space that is YouTube only, because you only find it on YouTube, and it's also shorter. So like, this is like a one and a half hour, two hour thing. It's a remote only, 30 minutes, chop chop, send it on to Riverside. Riverside Pretty good for that. Not great. It doesn't do good thumbnails. It doesn't do, the editing is still a little bit rough. It has, like, this auto editor, where, like, you know, whoever's actively speaking, it focuses on the editor, on the active speaker, and then sometimes it goes back to, like, the multi-speaker view, that kind of stuff. People like that. Ok, but, like, the shorts are still not great, you know, like, I still need auto, I need to manually download it, and then republish it to Yo…

AI assessment note: “Riverside is the closest that has come to it. Descript is number two.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q know, one of the things they said was we haven't seen anything that's kind of eyeopening to see people going to 200 dollar tier on this sort of thing. Haven't seen anything else like that in the space. Cause I, I think this is very new because of the new model capabilities, right? Where people, you know, it makes sense. Like you're willing to pay more money for this stuff.

A So this is something I've talked about before in terms of matching the dollar amount of spend to the capabilities of the AIs. The chart that I published in the past was, you know, OpenAI has like five levels of AGI-ness and level, level one is sort of like a chat boss. Level two is reasoning. Level three is agents. Four is organizations. Five is something super, superhuman. I don't remember what the exact levels are, but each, each, you can sort of each match each of them with like tiers. Like, uh, 20 dollars is like the ChatGPT tier. 200 dollars is where you're at. 2000 dollars is higher to 20,200 thousand, right? Like, you can see levels where it makes sense. I think Brightwave is also there, by the way. Like, I don't know what Brightwave charges, but it's higher, right, than a ChatGPT. And like, you have to deliver more value for that. But you, you can do it now.

AI assessment note: “matching the dollar amount of spend to the capabilities of the AIs”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q we underwrite as the future. And the future means that I'm more concerned about what's happening next. What are the new models? How do you gain market share? What has to be done? What are the new products that are going to be built? I'm less concerned about like where it's at right now in terms of market share, but that's just me. I don't want to speak for others.

A Yeah. I think the new models are, are really good. I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months. Uh, and, uh, it's, it's really interesting. I think OpenAI and Gemini are in this sort of price war a little bit with the, the Pareto frontier that I, I track, uh, in terms of like LMSYS versus the pricing. And Claude can still charge a premium, but still like, Have a lot of market share, obviously. And I think like that's just because they have a better model and like people just naturally gravitate to it, especially for coding, but also other things. And, um, I, I just think like articulating what makes a model good is just very, very difficult. Obviously this is benchmarks and evals and everyone has like, okay, today it's your turn to be best at SweetBench. And then like tomorrow is my turn. Uh, but like, it's, it's really stupid. Like we're, we're just like talking about like, you know, .12 differences in, in like SweetBench, but, I wonder, you know, if you're talking about like, okay, I am investing thirteen billion dollars in Anthropik for Series F to underwrite cloud five, right? What, what does it have to do? Like, what kind of, what kind of conversation does that look like? I have no idea. I'm not saying that, you know, but I'm just like.

AI assessment note: “I think the new models are, are really good. I mean, Opus 4.1”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q and They first, they break it down, and it's like, that is something they've trained into the models. Like, I don't think Deep Seek has, doesn't have it built in, but it probably could do it. And just thinking about that interface between, like, if the model needs it to be able to do the task end to end on its own, like, can it do that sort of thing?

A I think that my, my challenge with this whole reconciling this approach with the no harnesses thing is that I think a lot of the way that people, especially engineers, want to model it is that the plans and, and the memories are tools, and there are no special plan tokens. There are no special memory tokens. It's just context, or it's just, you know, whatever. Specifically for planning, because then you can do fan out to other agents. For tool calls and stuff. So it doesn't have to be sequential, but I'm just like, is this a fork in the road? Like, or, you know, do we have to make a real choice here as to do we outsource things to tools or do we keep it native within the model's tokens?

AI assessment note: “do we outsource things to tools or do we keep it native within the model's”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q and They first, they break it down, and it's like, that is something they've trained into the models. Like, I don't think Deep Seek has, doesn't have it built in, but it probably could do it. And just thinking about that interface between, like, if the model needs it to be able to do the task end to end on its own, like, can it do that sort of thing?

A I think that my, my challenge with this whole reconciling this approach with the no harnesses thing is that I think a lot of the way that people, especially engineers, want to model it is that the plans and, and the memories are tools, and there are no special plan tokens. There are no special memory tokens. It's just context, or it's just, you know, whatever. Specifically for planning, because then you can do fan out to other agents. For tool calls and stuff. So it doesn't have to be sequential, but I'm just like, is this a fork in the road? Like, or, you know, do we have to make a real choice here as to do we outsource things to tools or do we keep it native within the model's tokens?

AI assessment note: “do we outsource things to tools or do we keep it native within the model's”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Yeah. Because you're essentially offering free storage or compute or something, right?

A Yeah. Um, for those interested in this idea, there's a fourth one, which is basically the control plane or like the off layer, the IAM policies and all that. Uh, and, and that's like, maybe you can call that security as well. It's like the fourth layer that people kind of pay for as its own independent thing. For those also interested, uh, HashiCorp has More breakdowns from this, like David McJanet, um, which I think is very interesting if you're, if you're just in the business of running a cloud infrastructure company, you should know these things. Many hundreds of businesses have run into the exact same problems. You should just not repeat them and just learn whatever the best practice is.

AI assessment note: “Yeah. Um, for those interested in this idea, there's a fourth one”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Yeah. Because you're essentially offering free storage or compute or something, right?

A Yeah. Um, for those interested in this idea, there's a fourth one, which is basically the control plane or like the off layer, the IAM policies and all that. Uh, and, and that's like, maybe you can call that security as well. It's like the fourth layer that people kind of pay for as its own independent thing. For those also interested, uh, HashiCorp has More breakdowns from this, like David McJanet, um, which I think is very interesting if you're, if you're just in the business of running a cloud infrastructure company, you should know these things. Many hundreds of businesses have run into the exact same problems. You should just not repeat them and just learn whatever the best practice is.

AI assessment note: “Yeah. Um, for those interested in this idea, there's a fourth one”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q models are bad at and the labs don't focus on enough. Like if you want something solved, one of the levers that you have is send us a couple of prompts on it. We might be able to get a category going on it. And this thing that we were talking about earlier, right? That once things get measured, they can get targeted. You can make that work for you.

A For me as a content creator, infographics, very needed. I took the latest deep seek paper and, uh, I, you know, they had some descriptions of their search agents and their coding agents. And I, Put it in and create an infographic. And, um, I, I just think, like, there's an industrial use case that doesn't require a lot of, I guess, design taste, but just requires some, like, you need to conform to some preset references, which is something that, uh, that is increasingly important, especially in, like, the Nano Banana series. But, um, yeah, and I think, like, OpenAI is releasing Image Two soon, which, um, is going to have that. So I, I think, like, it's, it's all, like, of a kind where, um, People need to incentivize like workhorse use cases and not just art. I don't know.

AI assessment note: “For me as a content creator, infographics, very needed.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q we underwrite as the future. And the future means that I'm more concerned about what's happening next. What are the new models? How do you gain market share? What has to be done? What are the new products that are going to be built? I'm less concerned about like where it's at right now in terms of market share, but that's just me. I don't want to speak for others.

A Yeah. I think the new models are, are really good. I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months. Uh, and, uh, it's, it's really interesting. I think OpenAI and Gemini are in this sort of price war a little bit with the, the Pareto frontier that I, I track, uh, in terms of like LMSYS versus the pricing. And Claude can still charge a premium, but still like, Have a lot of market share, obviously. And I think like that's just because they have a better model and like people just naturally gravitate to it, especially for coding, but also other things. And, um, I, I just think like articulating what makes a model good is just very, very difficult. Obviously this is benchmarks and evals and everyone has like, okay, today it's your turn to be best at SweetBench. And then like tomorrow is my turn. Uh, but like, it's, it's really stupid. Like we're, we're just like talking about like, you know, .12 differences in, in like SweetBench, but, I wonder, you know, if you're talking about like, okay, I am investing thirteen billion dollars in Anthropik for Series F to underwrite cloud five, right? What, what does it have to do? Like, what kind of, what kind of conversation does that look like? I have no idea. I'm not saying that, you know, but I'm just like.

AI assessment note: “I think the new models are, are really good. I mean, uh, Opus 4.1”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q the military or the government and be able to, like, give them advice and, like, shape what they're doing. And so, yeah, there's a lot of, like, weird, like, moral conflicts, I think, with these companies and, like, Who they take money from or, like, who they allow to use their models. That's very tough. But yeah, well, what did you think of the Dario memo? Were you, like, surprised?

A You know, you see these things that, you know, everyone knows like when you put something like this out, it's going to be all over the press. And like, why would you put anything potentially embarrassing? Like you just, you just didn't have to feed your, your critics. You know, you can just say like, Hey, we're, we're, we're looking at the Middle East. You don't have to say like, we're compromising our principles or whatever to, to go into the Middle East. Anyway. Yeah. I think more broadly, uh, you know, I, I don't have, I don't have a take on that. I think, um, everyone has to sort of find their way. Um, I think Google, you know, has had to deal with the don't be evil, A slogan that they had in the early days. And, you know, at some point, like, what does evil mean to you? Doesn't mean, doesn't, doesn't mean the same thing to me as if I live in a completely different country and context than you. So like, yeah, there's, there's one, there's one form of discussion. I think the other thing that happens a lot with big founders, with big CEOs, which, uh, you know, I've had sort of offline conversations. A lot of them like want to build AGI, but like that, that I approve of in order to fight the AGI that I don't approve of. So like, it's, you know, the only thing that can be the bad AGI is a good AGI. And, and so they, they rather at least be part of that debate rather than, I gue…

AI assessment note: “why would you put anything potentially embarrassing? Like you just, you just didn't have”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q You also don't want to pay Fable prices for everything. Why would you?

A Yes, like, even, like, I'm trying to live in an infinite budget world, but the, the budget I don't have is time, right? Like, this, this, this limiting budget that everyone Has. I don't care if you're working at a frontier lab, you still, you still have time as a, as a, as a limiter. So, um, yeah, if that's like, if a 20 trillion model is it, there's no 200 trillion, there's no quadrillion scale model, then that's kind of it as far as usability is concerned. Living in an infinite budget environment. So therefore you just need something different. Whatever is it, maybe it's like thinking machines and stuff. Maybe it's Together AI's SSM stuff. Whatever it is, we don't have it yet.

AI assessment note: “I'm trying to live in an infinite budget world, but the budget I don't have is time”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q You also don't want to pay Fable prices for everything. Why would you?

A Yes, like, even, like, I'm trying to live in an infinite budget world, but the, the budget I don't have is time, right? Like, this, this, this limiting budget that everyone Has. I don't care if you're working at a frontier lab, you still, you still have time as a, as a, as a limiter. So, um, yeah, if that's like, if a 20 trillion model is it, there's no 200 trillion, there's no quadrillion scale model, then that's kind of it as far as usability is concerned. Living in an infinite budget environment. So therefore you just need something different. Whatever is it, maybe it's like thinking machines and stuff. Maybe it's Together AI's SSM stuff. Whatever it is, we don't have it yet.

AI assessment note: “the budget I don't have is time, right?”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q I agree, but to be fair, and so it actually doesn't really help the growth significantly today. We've had this kind of conversation with like investors and other people. It's like, how do you convert people from open source?

A The open source business conversation is so All over the place. Right. Okay. I'll just like for listeners who maybe they haven't thought this through, a lot of people say, oh, it's a free tier, right? Like, oh, if you run it yourself, but if, when you get serious, call us. Right. And then other, uh, and then, uh, me personally, because of my temporal experience, it actually is the way that it's the, it's GTM into some of the largest companies where we wouldn't pass their review process, maybe because we're too young of a company or like, there's like parts of the stack that we, Haven't like that just doesn't work with them, but because it's open source, then they, then they adopt it. And then later on, we figure it out. That's the low end and the high end. I don't know if it.

AI assessment note: “because it's open source, then they adopt it. And then later on, we figure it out.”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q Yeah. But like, yeah, no, I mean, we had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something. So yeah, so, like, the question is, like, your revenue is great, but, like, how about those margins? So I think that's very fair. Yeah.

A Yeah, and, like, it's gonna be the major factors of, like, compute. Personnel is now an increasingly key one. Um, data. You know, we, we have this concept of the four wars of, of AI, and it's, it's all around, like, the, the key battlegrounds that, that people have. The lesson is multimodality. And so like, uh, I, I think like there, there's, there's a lot of like really interesting, uh, work there that's, that's being done that like, uh, I think it's like a balancing act. Like you, you just kind of have, you need a visionary founder. Successes like, you know, two to three times a year. Um, and you just kind of keep the plate spinning while, uh, while you're growing this thing. But like, I really wonder what the sort of quote unquote terminal value, you know, just use the finance term of what all these things are.

AI assessment note: “it's gonna be the major factors of, like, compute. Personnel is now an increasingly key”

Redirected raw tape D 3 · C 3 · P 3 · Cm 3 3.00

Q the military or the government and be able to, like, give them advice and, like, shape what they're doing. And so, yeah, there's a lot of, like, weird, like, moral conflicts, I think, with these companies and, like, Who they take money from or, like, who they allow to use their models. That's very tough. But yeah, well, what did you think of the Dario memo? Were you, like, surprised?

A You know, you see these things that, you know, everyone knows like when you put something like this out, it's going to be all over the press. And like, why would you put anything potentially embarrassing? Like you just, you just didn't have to feed your, your critics. You know, you can just say like, Hey, we're, we're, we're looking at the Middle East. You don't have to say like, we're compromising our principles or whatever to, to go into the Middle East. Anyway. Yeah. I think more broadly, uh, you know, I, I don't have, I don't have a take on that. I think, um, everyone has to sort of find their way. Um, I think Google, you know, has had to deal with the don't be evil, A slogan that they had in the early days. And, you know, at some point, like, what does evil mean to you? Doesn't mean, doesn't, doesn't mean the same thing to me as if I live in a completely different country and context than you. So like, yeah, there's, there's one, there's one form of discussion. I think the other thing that happens a lot with big founders, with big CEOs, which, uh, you know, I've had sort of offline conversations. A lot of them like want to build AGI, but like that, that I approve of in order to fight the AGI that I don't approve of. So like, it's, you know, the only thing that can be the bad AGI is a good AGI. And, and so they, they rather at least be part of that debate rather than, I gue…

AI assessment note: “I don't have a take on that. I think more broadly, uh, you know”

Partly raw tape D 3 · C 2 · P 3 · Cm 2 2.55

Q Yeah. I think much improved, uh, post training all over to, uh, Make for a better coding model?

A Yeah, I think there's, like, different kinds of coding, right? Like, um, it's, it's interesting for me to observe that there, for example, uh, so I'm just gonna pull it up on the chart here, because I always like to show people visuals. Um, you're at 55 on SweetBench, and O-one gets, like, a 41, but then on Um, oh, I, I don't think I, I don't think I have the others, like, but, but either, um, it is, it is less, uh, uh, it is not at one level. And so I, I think, I, I think I struggled to get some kind of intuition of when, like, like what are the different elements of coding? I guess there is like, you know, single file edits, whether it's like a diff or, um, um, or whole file, and then there is entire project edits. Is that, um, a reasonable split? Are there more to this?

AI assessment note: “I think there's, like, different kinds of coding, right?”

Partly raw tape D 3 · C 2 · P 2 · Cm 2 2.30

Q Yeah. But like, yeah, no, I mean, we had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something. So yeah, so, like, the question is, like, your revenue is great, but, like, how about those margins? So I think that's very fair. Yeah.

A Yeah, and, like, it's gonna be the major factors of, like, compute. Personnel is now an increasingly key one. Um, data. You know, we, we have this concept of the four wars of, of AI, and it's, it's all around, like, the, the key battlegrounds that, that people have. The lesson is multimodality. And so like, uh, I, I think like there, there's, there's a lot of like really interesting, uh, work there that's, that's being done that like, uh, I think it's like a balancing act. Like you, you just kind of have, you need a visionary founder. Successes like, you know, two to three times a year. Um, and you just kind of keep the plate spinning while, uh, while you're growing this thing. But like, I really wonder what the sort of quote unquote terminal value, you know, just use the finance term of what all these things are.

AI assessment note: “it's gonna be the major factors of, like, compute. Personnel is now an increasingly key”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.