The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Sarah Sachs no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Do you have, um, failing evals that you're just hoping that will have success eventually when a good model comes out?

A I mean, yeah, so I think, I mean, I could talk about this for 60 minutes, so I will limit myself. I think it's a real issue when people say evals, and it's just like, that's quality, that's like unit, I mean, it's like saying testing, it's not just unit tests, right? So we have the equivalent of unit tests, regression tests, those live in CI, those have to pass a certain percent, you know, within some stochastic error rate. Then we have, as you're building a product evals of these aren't passing right now, and this is launch quality. So we have a report card, and we need to, on these categories, you know, be at 80 or 90% of all of these user journeys to launch. And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a point where our evals were saturated, and we weren't able to really give insightful feedback other than it wasn't worse. And not only is that not helpful for our partners, it's not helpful for us to understand where the stream is going, you know, going back to that analogy. And so we spent a lot of time thinking about what Notion's last exam looks like, right? Not just Humanity's last exam, but Notion's last exam. And, um, there's a lot of, you know, dreams about w…

AI assessment note: “we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Like, is there ever a discussion of, like, we're not going to ship it because we're not able to tie it down? Or are you happy to just, like,

A No, I mean, there are a lot of things where we choose not to use MCP because we want to add more high touch to quality. I think search and agentic find is like the largest instance of that, where we have, um, slack and linear and JIRA search and notion that is not using necessarily the search MCP functionality that is provided by those companies. And that's because it's quite critical. We think to how our agent trajectories work is for us to have, A little bit more control on the functionality of the search journey, and so it usually comes from quality, and there's a long tail of things, and that's why we built an MCP client, or an MCP server, excuse me, so that people can connect whatever they want. There is that long tail, right? But we, for search particularly, I would say that's like the primary entry point, but there are other connections as well that it's a little bit of secret sauce about when we are okay with like MCP functionality and user-driven auth, and when we actually want We want to carry a lot more ourselves.

AI assessment note: “there are a lot of things where we choose not to use MCP”

Answered raw tape D 4 · C 5 · P 5 · Cm 4 4.55

Q Now it's free. It's amazing. I'm worried about when I have to start paying. How do you think about, so you have Notion credits as a payment for this, which is like separate from the usual tokens that the model generates. How do you design pricing, value-based pricing based on the task and things like that?

A So they are, um, the credits and payment structures associated with the token usage. The reason that we had to make it not just throughput of tokens is that it's not always priced that way. Like our, um, fine-tuned open source models are served on GPUs. Web search is priced differently. You know, if we were to host sandboxes, those are priced differently. So we had to think of an abstraction above tokens, and it's also not just tokens. It's the token model, um, and serving tier trade-off, right? Because we can have priority tier processing. We can have asynchronous processing. The cash rate could be different, um, depending on who uses it when, right? And so we wanted to, um, from the get-go commit to making sure that customers were getting the fair deal. Not necessarily that we were making a ton of money off of it, but that customers were paying for what was reasonable. That's the fundamental of where we started. And also, you know, we're selling enterprise SaaS. So if we sell credit packs and you get discounts, if you're an enterprise and you buy a certain amount of credit packs and things like that. So it also just helped the sales motion, um, work a little bit easier. So that's the answer on the abstraction of credits to dollars. Now, Was the question how we decide how to price it, or?

AI assessment note: “So we had to think of an abstraction above tokens”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q No, just like anything sort of interesting technically, right? Like I think you had, you had some, uh, book points. I always call these like check marks along the way when, when the guest says something that they want to return to later, I just like check mark it. I'm like, okay, we'll get back to it.

A Meeting notes was one of those things where at first we were nervous that we'd have to teach people a different way to work. And we were nervous that that was a lot of user friction. I think one of the reasons why, I mean, they're one of our biggest growth levers. I think they're one of the most like, In terms of virality of adoption and retention, quite strong. Um, and so we've invested more and more as we did that. I think what's really powerful about it is, again, Notion is the system of record of where and how you work. The way that I use meeting notes is every one-on-one meeting I have is meeting notes. When I do my performance review for myself, my self review, I say, primarily look at all my conversations with my manager and like write up what I did this year, right? Because if I didn't talk about it in my one-on-one with my manager, it probably wasn't relevant for my performance review. So It also just adds a ton of signal on prioritization. That's really helpful for a good system of record. That's really helpful for like our agent. It's also like caused a lot of scaling for search and for the agent. Um, and you know, it's, it's just an explosion of content when you have transcripts like that. Um, how we do compaction, a lot of that was triggered by meeting notes, pass into context, things like that. Um, so it's been a good impetus for us to think about longer form, um,…

AI assessment note: “Meeting notes was one of those things where at first we were nervous”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Yeah. Do you believe in internal hackathons for this stuff?

A I think there's like two different versions. So one is like, we just have a, a solid bench of senior engineers that come and go on what we call the Simon vortex and productionizing what we built, right? Because when you're in the Simon vortex, the velocity is super high. The direction changes daily and it's meant to be like the equivalent of a Skunk Works Lab. We don't need to do hackathons for that. We need to have senior engineers that we trust to come in and out of those projects. For instance, like management boundaries are really loose. Like you report to him, but you work for her right now. Like, That is something that when we hire managers, it's important they don't care about because we tend to form orc structures.

AI assessment note: “We don't need to do hackathons for that.”

Partly raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q What's the size of the team today, both engineering and overall?

A I manage, ah, the team that's what we'll call like core AI capabilities and infrastructure. That's about 50 people. But then we have, I partner teams that do packaging, so how it shows up in the corner chat versus custom agents versus meeting notes, that's another 30, 40 people. And then every team that has a product service at Notion that a user can interface with owns the tool that the agent interfaces with, the editor team. The team that did CRDT for offline mode is the same team that handles how two agents, um, edit competing blocks, right? It's the same problem. The team that built the underlying SQL engine is the same team that owns how the agent asks it to run a SQL query and it does it performantly. And so from that regard, anyone working on product engineering is tasked with making them work for customers that are humans and agents. Because over time, the majority of our traffic will be coming from agents using our interface, not humans. And so Our objective is to make it so that the whole product org is building for agents.

AI assessment note: “core AI capabilities and infrastructure. That's about 50 people.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.