Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What tools and what data connections come to mind when you say, what's interesting? What, what, what's notable work that people have done?
A Oh, okay. So, my favorite example on this is that until very recently, I would argue that it was Basically impossible to get an LLM to draft an email for me in any useful way, because most times you're sending an email, you're not just writing something for the sake of writing it. Chances are context required is a whole bunch of historical emails. Maybe it's notes that you've made. Maybe it's meeting notes. Maybe it's, um, pulling something from your, um, any of like wherever you at work store stuff So for me, like Google drive, one drive, um, and our super best databases, if we need to do some analysis or some data or something, Preferably. Model can be plugged into all of those things and can go do some useful work. Based on it, the things that, like, I find most impressive currently that I am somewhat surprised work really well in late 25 are that I can have models use Superbase MCP to read only, of course, run a whole bunch of SQL queries to do pretty significant data analysis and make charts and stuff, and can read my Gmail and my Notion.
AI assessment note: “models use Superbase MCP to read only... and can read my Gmail and my Notion.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q That's true. I never thought about that. I've been in a database data industry prior and there's a lot of shenanigans around benchmarking, right? So I'm just kind of going through the mental laundry list. Did I miss anything else in that, in this category of shenanigans?
A I mean, okay, the, the, the biggest one, like, that I'll bring up, like, is more of a conceptual one, actually, than, like, direct shenanigans. It's that the things that get measured become things that get targeted by labs that they're trying to build, right? Exactly. So that doesn't mean anything that we should really call shenanigans. Like, I'm not talking about training on test set, but if you know that you're going to be great at another particular thing, If you're a researcher, there are a whole bunch of things that you can do to try to get better at that thing that preferably are going to be helpful for a wide range of how actual users want to use the thing that you're building, but will not necessarily do that. So, for instance, the models are exceptional now at answering competition maths problems. There is some, uh, relevance of that type of reasoning, that type of work, um, to, like, how we might use modern coding agents and stuff, um, but it's clearly not one for one. So the thing that we have to be aware of is that once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the next couple of years. There's no, Silver bullet to defeat that other than building new stuff to s…
AI assessment note: “the biggest one, like, that I'll bring up, like, is more of a conceptual one”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Is the dataset public? Or what's, is it, is there a held out set?
A There's a held out set for this one. Um, so we, we have published a public test set, but we, we've only published 10% of it. The reason is that for this one here specifically, it would be very, very easy to, um, like, have data contamination, because it is just factual knowledge questions. Um, we will update it over time to also, um, prevent that, but with, yeah, kept most of it held out so that we can Keep it reliable for a long time. It leads us to a bunch of really cool things, including breakdown quite granularly by topic. And so we've got some of that disclosed on the website publicly right now, and there's lots more coming in terms of our ability to break out very specific topics.
AI assessment note: “There's a held out set for this one... we've only published 10% of it.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q And I think the general assertion or the message is that the efficiency from next gen NVIDIA chips is actually not four X. You have what, three X or four X? You have three X in here and it's, it's like two X maybe, or it's more of like a power story rather than like a share sort of compute tokens efficiency story. But yeah, what's going on in hardware?
A Okay, so the, the, the answer, unfortunately, uh, is it depends, and it just depends massively on, like, so many things across a bunch of different types of workloads and ways to think about it. So one of the simplest ways to think about this is to take single relevant model, to think about serving it at speeds that are realistic for what you actually might want to hit and can afford to hit, and then think about the throughput per GPU that you can achieve serving the model at those speeds. One of the reasons that's important is that there's a trade-off between the throughput per GPU that you can achieve and the per user speed that you can achieve, and as in, it costs more to serve stuff fast to users. When you run all of that, for especially big sparse models, you can get a lot better than two or three eggs gain going from hopper to blackwell generation to video. I am, this shouldn't be too controversial to say, but like, I'm pretty confident that Blackwell has delivered pretty enormous gains, and that the next couple of years of NVIDIA's roadmap are going to continue to deliver quite enormous gains, and that those will actually come through as lower total cost per token to the companies that are running models on them, and will allow bigger models, will allow way more tokens to be made for lower cost, and that that's going to continue. These things also stack on all of the sof…
AI assessment note: “you can get a lot better than two or three eggs gain going from hopper”
Answered raw tape
D 4 · C 4 · P 2 · Cm 3 3.35
Q What are we going to be talking about next year? Like what's, what's like, what's emerging that you're seeing and like maybe not in the discussion?
A The first answer that I'll give to that is the, the boring answer is that on most of our charts, the lines go in a particular direction and our overall prediction is the lines are going to keep going in that direction. We're going to do a lot and do a lot to Be as useful as possible to developers and companies to measure what's important on every one of those and along those lines. But I think we're going to talk about similar stuff. It's just that we're going to have continued on this trajectory for another year and things are going to feel pretty different because of that happening. I know this is the boring answer to that question.
AI assessment note: “I think we're going to talk about similar stuff.”