The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Samuel Colvin no published score: only 5 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 5 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
5exchanges match
5on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yeah, that is, seems like a nightmare if you allow that. So you started Monty over Christmas. What's the inspiration? What's that origin story?

A So I actually had, had a like very early version of this that I had done a couple of years ago and had completely abandoned. And then I spoke to maybe four different people at Anthropic and each of them independently. I like did my standard thing about how important type safety is because I, you know, I think type safety is important for humans, but it's critical for AIs. And I say this to people all the time and some of them nod and some of them don't nod, but like these four people I spoke to at Anthropic each independently said, Yeah. Type safety is super important if you're chaining tool calling, or if you're like using, using code for tool calling. And when the first fourth of these people spoke to me, I was like, They're obviously thinking about something. I find Anthropic hilarious because they're like the most secretive company, and yet everyone gets excited about whatever's going on inside Anthropic, and suddenly everyone, and they all hint at you about whatever it is that's going on at the moment.

AI assessment note: “these four people I spoke to at Anthropic each independently said, Yeah. Type safety”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q This is a side tangent on Cloudflare. Is Python supported a full? I actually wasn't fully aware that Of what the status of that thing is.

A Yeah, so, so Pyodide, which is Python running inside the browser in Scripton, is supported now by Cloudflare. They just, they basically, they're having some struggles working out how to manage, ironically, dependencies that have binaries, in particular Pydantic, because these workers, where you can have thousands of them on a given metal machine, you don't want to have a different, you basically want to be able to have a shared, shared memory for all the different Pydantic installations, effectively. That's the thing they, they work out. They're, they're working out. Who's my friend, who is the primary maintainer of Pyodide, works for Cloudflare, and that's basically what he's doing, is working on how to get Python running on, running on Cloudflare's, like, network.

AI assessment note: “Pyodide, which is Python running inside the browser in Scripton, is supported now”

Redirected raw tape D 3 · C 5 · P 5 · Cm 4 4.25

Q a founder too, like, how did you think about the AI interest rising? And then how do you kind of prioritize, okay, this is worth going into more of the, and we'll talk about Pydantic AI and all of that. What was maybe your early experience with LLAMPS and when did you figure out, okay, this is like something we should like take seriously and focus more resources on it?

A I'll answer that, but I'll answer, which I think is like a kind of parallel question, which is Pylantic's weird because Pylantic existed obviously before I was starting a company. I was working on it in my spare time. And then beginning of 22, I started working on the rewrite in Rust, basically. And I worked on that. I worked on it full time for a year and a half. And then once we started the company, people came and joined. And it's a weird, it was a weird project because that would never get signed off inside a startup. Like we're going to go off and three engineers are going to work full on for a year in Python and Rust. Writing like 30,000 lines of rust just to release open source free Python library. The, the result of that has been excellent for us as a company, right? As in, it's made us remain entirely relevant. It's like Pydantic is not just used in the SDKs of all of the AI libraries, but I can't say which one, but one of the big foundational model companies, when they upgraded from Pydantic v one to v two, their number one internal metric of performance is time to first token That went down by 20%. So you think about all of the actual AI going on inside, and yet at least 20% of the CPU or at least the latency inside requests was actually Pydantic, which shows like how widely it's used. So we've benefited from doing that work, although it didn't, it would have never h…

AI assessment note: “I'll answer that, but I'll answer, which I think is like a kind of parallel question”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Yeah. Any thoughts on all the other people trying to build on top of open telemetry in different languages too? There's like the open LLM tree project, which doesn't really roll off the tongue, but how do you see the, yeah, the, the future of this kind of tools? Is everybody going to have to build, why does everybody want to build their own open source observability thing to themselves?

A I mean, we, we are not going off and trying to instrument the likes of the OpenAI SDK with the new semantic attributes, because at some point that's going to happen and it's going to live inside the hotel and we might help with it, but we're a tiny team. We don't have time to go and do all of that work. So open telemetry, like interesting project, but I suspect eventually most of those semantic, like that instrumentation of the big, of the SDKs will live, like I say, inside the main open telemetry repos. What happens to the agent frameworks? What data you basically need at the framework level to get the context is kind of unclear. Uh, I don't think we know the answer yet, but I mean, I was on the, I guess this is kind of semi-public because it was on, I was on the call with the OpenTelemetry call last week talking about Gen.ai and there was someone from Arise talking about the challenges they have trying to get OpenTelemetry data out of Langchain where it's not, uh, like natively implemented. And obviously they're having quite a tough time. And I was realizing, hadn't really realized this before, but how lucky we are to primarily be talking about our own agent framework, where we have the control rather than trying to go and instrument other people's.

AI assessment note: “I suspect eventually most of those semantic... instrumentation of the big... SDKs will live”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Okay. Well, we will, we will kill IPMB at some point. And yeah, any other takes? Uh, I was going to ask you just like broadened out just about the London scene, you know, what's it like building out there, you know, over, over the pond?

A I'm, I'm an evening person. And the good thing is that I can get up late and then work late because I'm speaking to people in the U S a lot of the time. So I got invited just earlier today to, uh, some drinks reception about AI at, uh, Tim Downing street with the prime minister. So I'm, I'm feeling positive about the UK right now on AI, but I think, look, like, everywhere that isn't the US and China knows that we're, like, way behind on AI. I think it's good that the UK is, like, beginning to say this is an opportunity, um, not just a risk. I keep being told you should be at more events, you should be, like, you know, hanging out with AI people more. My instinct is, like, I'd rather sit at my computer and write code. I think that, like, is probably a more effective way of, uh, getting people's attention. I'm, like, a bit of me thinks I should be sitting on Twitter, not on, not In San Francisco, chatting to people, I think it's probably a bit, a bit of a mixture and I, I could probably do with being in the, in the States a bit more. I think I'm going to be over there a bit more this year, but like, there's definitely the risk if you're in somewhere where everyone wants to chat to you about, about code, where you don't write any code and that's, that's a failure mode.

AI assessment note: “I'm feeling positive about the UK right now on AI”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.