The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Lukas Biewald no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And you, you guys launched this prompt suite in April. Like, can you talk us through the sort of, you know, thought process of, Hey, like, you know, I, I really admire this as a leader and as a technical person, you're like trying to stay really plastic about what is actually changing in machine learning. How'd you think through this change?

A Well, it's really hard, right? I mean, so what happened was we have a great business that, you know, makes like an ML, uh, set of ML tools for training models. And we actually helped most of the LLMs out there were built using weights and biases. And then we started to see like, wait a second, some of these ML tasks, you could just ask the LLM, right? So instead of doing like a sentiment analysis model, you could just be like, Hey, like, is this document positive or negative sentiment? Like, For structuring documents, you can just be like, hey, find all the names like in this document, and it actually works super well. And a little piece of me is a little bit sad about that because we have this like great, simple, relaxing business that grows revenue every, every month that I always dreamed of, right? So, you know, part of me is like, shit, this is actually our kind of first real existential threat, I think, you know, and, and, um, and, you know, I went to my like leadership team and I went to my board and I was like, I think there's like a real existential threat here. And I think they were like, Hey, you know, we don't like see it in the data. Like, are you sure? Like, maybe you're being paranoid. And I guess I do feel sure. And I don't want to say I'm like the only one or like pay myself as the hero. Like, you know, my co-founder is also seeing this and you have people talki…

AI assessment note: “we started to see like, wait a second, some of these ML tasks, you could just ask the LLM”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q One of the things that's common to people or to developers is that they love to write their own tools and they tend to really enjoy using open source over closed source solutions. How did you think about the open versus closed source approach and how did you think about, you know, making something that's valuable enough and good enough to overcome that natural inclination to just do it yourself?

A What's funny, like, I think the tools thing, I've always felt like, I've always felt like, kind of proud of making tools for developers. Like that's always felt like really good because I think developers sort of know what quality is. Like, I mean, it's like, I, I kind of like making a tool for someone that could make the tool themselves because it kind of raises the bar. And this stuff, like my grandfather was like a pattern maker, which is like a sort of, you know, like the person who makes a pattern for other machinists. And he had the same attitude of like, look, I'm making this stuff for like other Engineers and like, there's like an honor in that. So I definitely feel that pressure and love it. The open source first closed source thing was really just like, we didn't know how to make an open source business. So, so like we kind of started off closed source cause we just, we actually wanted to have like a working business and it's had like a pro There's been a major pro, which is that all our competitors are open source. And what that means is that they don't get to see how users actually use their software. And so I think our software is a lot more ergonomic because we have like metrics on what people actually click on. If people aren't clicking on a button, we remove it. If people like, you know, pick an option all the time, then we know to like make that the standard op…

AI assessment note: “The open source first closed source thing was really just like, we didn't know how”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Do you still pay attention, I'm sure you do, actually, to, like, annotation? Like, what do you, What do you think happens to the data, data annotation space and like, you know, the land of LMS and LHF and such?

A You know, I'll be like honest, actually, I'll just be like totally honest. I find it like incredibly stressful because I still feel bad that we lost the scale. Like I still like, it's just like lingered with me and I, I admire scale. Actually, I know hard that, that businesses. So I have just like deep admiration for their like execution, but as a competitive guy, I kind of can't get over it. So I've like always inundated with questions from VCs, like whatever any annotation companies raising, I know about it because everyone like calls me, but I, I honestly try. I know I should be closer to it, but I try to stay away from it just because it caused me so much anxiety to look at what's going on that I, uh, I just can't deal with it.

AI assessment note: “I try to stay away from it just because it caused me so much anxiety”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q did differently with the second company? I feel like, you know, I've started two companies and with the second one, there was all sorts of lessons I applied immediately. Were there two or three key takeaways that when you started Weights and Biases made the second time around easier or was it harder? How did you think about, you know, key, key, key learnings or how to apply new things?

A Yeah. I mean, I think like one thing was like extreme clarity about who we were serving. So I'm surprised I don't hear this more because like the, the ways that biases started with a, with a customer profile. And I think it's actually a nice way to start a company because, you know, especially as like a founder, you have to spend so much time with your customers. You have to seek them out. Like picking a customer that you love, I think is a really good thing for your like mental health, you know? And so That was like a big thing, and then I think like, I think I've just been a more confident person in myself. Like anytime I start thinking like, okay, like long-term or short-term, it's just like you always want to think long-term. Like everybody wants you to think short-term. Like everyone's going to push you to think short-term. They wouldn't say it like that, but it's like, you know, it's like people can see like ARR growth. They can see like user growth. It's harder to see like product quality, right? And so I think like, I think I'm a competitive guy who likes, you know, metrics and likes accountability. But I actually think that can get counterproductive for me, where, you know, you start, like, sacrificing short-term things to grow these external facing metrics, and I just really try to fight that myself. I think everybody, like, chases, every entrepreneur chases, like, sh…

AI assessment note: “one thing was like extreme clarity about who we were serving”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q some cases they're, they're training their own instance of LAMA-II or whatever they're using. Do you think that's where the world is heading or do you really think things kind of collapse onto some of these proprietary models like over time? Like it's six months from now, it's a year from now, it's years from now. I'm just sort of curious about how you think about adoption of open source.

A You know, it's funny. I, I feel like lately what I've been telling people is like, I'm just trying to see the world clearly as it is today. I can't predict the future and I can barely keep track of, you know, what people are doing today when I consider it like my, my full-time job. So I, I'm like scared to prognosticate like what, you know, might be coming, but I think you're right that that's what's happening now. I think like there are like a bunch of things that could change, right? Like, I think like, you know, GPT is way far out ahead. And it's hard to fine tune it. Not even possible with, with GPT four. And I think that that is like a little, that's not like a technical limitation. I guess sort of like a business model, um, you know, limitation. So that might change. I think that there's a lot of hidden costs to running your own model. I think people are really enamored with the idea of running their own model. And I, I've kind of seen this before where I think at the end people do rational things, but that kind of takes them a while. So I'd rather sort of, Support what looks like the rational workflow. I mean, I think the insane thing must be crazier to be an investor in this world is like very, very few people have LLMs in production. Like there's probably more companies that have raised money as like LLM tools than companies that have LLMs in production, which is like …

AI assessment note: “I think that there's a lot of hidden costs to running your own model.”

Partly raw tape D 3 · C 4 · P 4 · Cm 4 3.70

Q I don't think we got a taste of that, but we did talk about whether or not probabilistic graphs are, are coming back a little bit. How did you, how'd you go from, you know, Stanford to founding figure eight?

A Yeah, you know, it's funny. I actually really struggled doing research with, with Daphne. Basically the things that I tried just barely, barely worked. Like, you know, I published a couple of papers that I feel kind of ashamed of where it was sort of like, go from like, 68% accuracy to 70% accuracy in a task nobody cares about by throwing like a thousand x to compute. And by the way, like, kind of guessing the most likely answer is probably like 64% accuracy. So, um, you know, it just, It, it felt honestly kind of pointless and sad. Like I love the idea of like computers learning to do things, but it's hard to sort of sustain the enthusiasm for that when everything you try just completely, you know, doesn't work. And even the things that do work, you kind of wonder if you're like p-value hacking, like, okay, I tried a thousand things, you know, so I guess something's going to be like a little bit more accurate than, than a baseline.

AI assessment note: “I actually really struggled doing research with, with Daphne. Basically the things that I tried”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.