The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Elad Gil argument clarity score 4.0/5 from 10 exchanges on raw tape · average scores: directness 3.3 · coherence 4.3 · precision 3.9 · compression 3.6 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
10exchanges match
10on raw tape
3redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q ways. So I guess it depends on what you mean by external validation. Like in my mind, again, like I often think about things from a perspective of, are you trying to build a startup that's actually going to change the world? Like, do you have this big thing that you're dreaming of? And if you have this big thing that you're dreaming of, you, Like, why do you care?

A Maybe the way to think about it is in Sarah's context, like if you haven't, say you're a YC founder, you haven't been at Google, you haven't been at Meta, you haven't been at Twitter, you don't have this network of engineers, you're a complete unknown, you haven't worked with very many people, you're straight out of school. How do you then attract that talent? And to your point, you can tell a story of how you're going to build things or what you're going to do, but it is a harder, um, obstacle to basically convince others to join you or for others to come on board or to have money to pay them if you haven't. If you don't have long work at history. So I think maybe that's the point Sarah's making.

AI assessment note: “it is a harder obstacle to basically convince others to join you”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q um, contract. And I think that's, I think that's right. I can't think of, I feel like you guys see more businesses, um, than I do. Um, and I hope, um, and then, but I hope, um, but, but, but like you guys will be able to like chime in on, on that more. Like, is that, is that unique to this industry or where else have you seen that?

A It's been a while since I've seen, uh, so many companies ramping so quickly, like, Uh, and sometimes they were fake ramps. So like, you know, in the internet wave of the nineties, it was kind of startups selling to each other and kind of bootstrapping off of venture capital. And then there was giant telecom build outs on like a five year cycle that Um, cause huge revenue uplift and then suddenly there was a glut and things dropped dramatically. Here it feels like things are ramping really fast off of products that are a couple months old, which sometimes suggests that there's not defensibility. Um, and so then the question starts to become, okay, how do you build defensibility and what does that mean? And how do things get commoditized and do they, and you know, um, so there's a couple different markets where suddenly you see three companies all go from zero to five or zero to ten million of revenue in a year. Um, yeah. And then you're like, okay, there's three of these companies and they all ramped at the same time, the same amount. And so there's enormous demand, but what does that mean in terms of, do they cannibalize each other? Can three more entrants come in and do the same thing? Like what, what is the basis for competition in that market? And so I think there's a lot of that happening too, which is at least for me, pretty unexpected. And I think it's just cause we have …

AI assessment note: “in the internet wave of the nineties, it was kind of startups selling to each other”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q So like for coding, for example, you can have principles like, did it actually serve the final answer? Or did it like do a bunch of stuff that the person didn't ask for? Or does this code look maintainable? Are the comments like useful and interesting?

A But, but with coding, you actually have like a direct output that, uh, you can measure, right? You can run the code. You can test the code. You can do things with that. How do you do that for medical information or how you, how do you do that for a legal opinion? Or, you know, so I totally agree for code. There's sort of a baked in utility function you can optimize against, or an environment that you can optimize against. In the context of a lot of other aspects of human endeavor, that, that seems more challenging. And you folks have thought about this so deeply and so nicely. I'm just sort of curious, you know, how do you extrapolate into these other areas where the, the ability to actually measure correctness in some sense is more challenging?

AI assessment note: “with coding, you actually have like a direct output that, uh, you can measure”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q now. Uh, if you ask ChatGPT how to solve a problem with Next.js, it tells you the solution for 2020. And I would love for that to be the solution for 2023. So please go and like help yourself to our docs. I can give you whatever format you want. But for other companies, it's going to be a challenge, right? Because they're expecting a different type of content negotiation.

A It seems like that's another place where tooling can become really valuable in terms of, you know, the ability to understand whether content that's provided in a corpus is, you know, falls under certain copyright laws or has other issues around it, or there may be other sort of tools that we increasingly will see from content owners or for content owners in terms of how you actually deal with this on the web. And, you know, we've already seen some early days versions of that around ImageGen. And some of the image generation models and people not wanting certain content included in that. Like Getty Images, I think famously pulled a bunch of data specifically to avoid this sort of issue or ask people to pull that data.

AI assessment note: “that's another place where tooling can become really valuable in terms of”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q right? It truly doesn't quite exist. So that's what I'm interested in, what we're interested in. So in that case means to build a good product, a good customer user experience, we need to think things more horizontally, more holistically. That means the decision making tend to be centralized in our design team, or like, tend to be centralized, right? So, um, less like more, so more Apple-like, less Amazon-like.

A It's funny because when I first met you, it was just you starting Notion, and as before you brought on Simon, and you talked about things that way even then, and I felt that one of the reasons I was lucky enough to invest or, you know, I came on board was because you had such a cohesive view of how you wanted to build software, and you had such a cohesive design aesthetic, and it was your mocks, but it was also how you were dressed and how that reflected into the product. I felt like it was extremely striking, you know, Like you're one of a very small number of people I've ever seen where that design aesthetic is just kind of permeated everything in a very cohesive way. And so that's one of the things that got me excited at the time. I was like, wow, this is capturing a, uh, aesthetic that could be an incredible product platform. But you also talked about things. Even then I remember in terms of like, okay, what's the, what's the cohesive Apple like thing that you can do for software?

AI assessment note: “you also talked about things... in terms of like, okay, what's the cohesive Apple like thing”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q that strategic tussle is the question of what do people actually want from search? Do they want hard facts? Do they want opinions? How much Lossiness. Are they willing to tolerate in order to get the compression of The LLM took 10 search results and summarized them for me, but maybe they hallucinated the summarization, right? It's, it's a really interesting place. Like what do people actually want, you know?

A Yeah, I think it's a really interesting open question and related to that too, you start, I've started to see, um, the degradation of some performance of some of the really early players in the market where there's a lot of qualifications or safety, or you start to see the LLM go out to I try and gather information, and it just really, in some cases, either slows down or makes the information worse, because to some extent, the reason you're interrogating an LLM is you're, you want to get to an answer, you don't want a web search, or you don't want that. You know, poor synthesis of bad data from the web. You want the good synthesis of bad data from the web. And so, yeah, it's been fascinating to watch. And to your point, I think a lot of people, um, have started to adopt perplexity for that specific use case because it has that middle ground of some form of IR plus LLM in a traditional sense. So yeah, I agree with you. It's, it's, it's going to be fascinating to watch how all the directions this goes. And I think the reality is people almost forgot that people from search, they actually want an answer. They don't. They don't want to do a search process. They want to get to a result in many cases or a list of results. It's almost like people forgot the user need in some sense.

AI assessment note: “the reality is people almost forgot that people from search, they actually want an answer.”

Redirected raw tape D 2 · C 4 · P 4 · Cm 4 3.40

Q stack starting from GPUs with NVIDIA and AMD and a bunch of others. We have a model layer with obviously OpenAI and Grok and Gemini and many, many others. We have an application layer. You had many, many, many of them on your podcast for agents to all kinds of software. How do we make sure this American stack is dominating that market share of tokens inference? Very good method.

A Yeah, it's really interesting because one other thing that you didn't mention, I feel, is a cultural exportation through the models. And so if you look at prior waves of sort of culture spread, it was the movie industry, it was social media, and then now it's these models, because a lot of people go to these models as a trust, a source of truth for history, for information, for other things. And there have been some famous examples in some of the Chinese models where there's a mission of Tiananmen Square or a mission of other facts. And relatedly, there's some things in some of the U.S. models that seem very politically slanted or otherwise not quite great. But it's interesting to also think about it from the perspective of a broader cultural export. So I just wanted to add that to your points on defense and scientific progress and other areas. I think that's another key thing.

AI assessment note: “one other thing that you didn't mention, I feel, is a cultural exportation”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q the name that they use is maybe generative reward model. I think that area goes beyond this kind of need for specific task annotations. We might need specific task annotations. Then the question is, how many labels will we need? And the hope is that in the limit, maybe you need as many labels as the user will provide the system when the user wants to teach it something new.

A Not to abuse it, but I, you know, a friend of mine used to use Nyquist Shannon sampling theorem as sort of a proxy for how much an intelligent person or machine can actually extrapolate the intelligence of something smarter than itself. And, you know, it feels like you're almost falling into some version of that where, you know, I think the theorem basically states that, um, You know, you have some frequency on a wave and you're sampling it and you need to sample, uh, above a certain rate to be able to actually reconstruct the wave. Right. And you could argue that that's some form of learning or intelligence or something else. And so you need to be smart enough to actually tell how smart you can be in some sense.

AI assessment note: “a friend of mine used to use Nyquist Shannon sampling theorem as sort of a proxy”

Not addressed raw tape D 1 · C 4 · P 4 · Cm 4 3.10

Q a lot of us, you know, see that and we're like, Hey, we, we remember being founders. We remember what it's like to have all bucks stop with you. And if there's a way that we can apply that level of decisiveness here, And accelerate teams through like the otherwise, like, you know, design by committee bullshit that kills a lot of companies, then like that's a net positive, right?

A The other piece of buying in companies is sort of integrating different systems and platforms. And often you see people do one of two things. They either let things run independently. That was WhatsApp for a long time at Meta, for example, or you integrate in all the infrastructure. And usually that makes things more performant. You can have features that cross over, et cetera. But the flip side of it is sometimes you see these giant projects that just stop a company from functioning, right? Like at Google, there was a famous sort of identity layer that was built, and for a year and a half, people just didn't ship things, right, in other areas. And so how do you all think about integrating and acquired companies and acquired technologies? But more generally, as you build and build and build and build, there's a big munch of stuff. How do you sort of consolidate that down in a performant way that's actually speedy to execute? And then I guess later on, maybe we could talk about whether AI has played a role in that or not.

AI assessment note: “The other piece of buying in companies is sort of integrating different systems and platforms.”

Not addressed raw tape D 1 · C 4 · P 4 · Cm 3 2.95

Q Claude's character of like, when Claude refuses, does it just say, I can't talk to you about that and shut down? Or does it actually try to explain like, this is why I can't talk to you about this. Or we have this other project led by Kyle Fish, our model welfare lead. Where Claude can actually opt out of conversations if it's going too far in the wrong direction.

A What aspects of that should a company actually adjudicate? Because the dumb version of this is I'm using Microsoft Word and I'm typing something up and Word doesn't stop me from saying things, which I think is correct. Like, I actually don't think in many cases these products should censor us or prevent us from having certain types of speech. And I've had some experiences with some of these models where I actually feel like it's Prevented me from actually asking the question I want to ask, right? In my, in my opinion, wrongfully, right? It's kind of interfering with, and I'm not like doing hate speech on a model. And so you can tell that there's some human who has a different bar for what is acceptable to discuss societally. And that bar may be very different from what I think may be mainstream too. So I'm a little bit curious, like, why even go there? Like, why, why is that? I have model companies business.

AI assessment note: “What aspects of that should a company actually adjudicate?”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.