The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Brandon Duderstadt argument clarity score 4.3/5 from 11 exchanges on raw tape · average scores: directness 4.5 · coherence 4.8 · precision 4.2 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
11exchanges match
11on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q was great. Uh, the, the open source, uh, AI world right now seems to be just completely exploding. It's super exciting. Is that, is that, uh, like, for, for you guys who are, like, deeply into the space, is that, is that the same impression as, like, the, The rest of us, like what's your overall sense of the, of the health and vibrancy of the open source AI ecosystem?

A Yeah, it's, uh, all time high as I would say right now, like we've got a ton of VC money flowing in the, you know, companies that are shipping things open source. So like, I love that that's happening. I love that, you know, I think on a podcast right after GPT for all came out, I think it was the weights and biases podcast with Lucas. I said something like the biggest challenge for open source is going to be getting like the monetary resources to do like a seven billion model or these bigger models. And so it's amazing to see companies like Mistral actually going and doing what they say they're going to do and releasing these amazing models. And Meta as well with the release of Llama I think has played a big part here. And I think also one of the reasons it's kind of popping off right now is we really are at this kind of new frontier in terms of the discovery of what these things can do. And so it's totally possible that some Random person, you know, in the middle of nowhere that's just like somewhat interested in this stuff, you know, spends enough time poking at it. There's so much new stuff to find that they can discover things that are really amazing. And so you see some of these really incredible techniques coming out of, you know, not maybe what you might call like the Royal science or like the, the universities, but some random person will be like, oh, hey, like I did t…

AI assessment note: “Yeah, it's, uh, all time high as I would say right now”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Any, any, any lessons learned there from a go-to-market perspective? Is that, is that just, um, you know, generally comparable as selling developer tools, meaning you need to be authentic and you need to be responsive, or what have you learned?

A Yeah, so I think one of the most interesting things I've learned was be open to your, your thesis about core audience shifting. So when I originally built, you know, and envisioned Atlas, Um, and sort of brought it to Andre and the team to, to help me realize it. I very much thought that the, you know, ICP ideal customer profile would be a machine learning engineer. And, you know, a lot of our early users were machine learning engineers. We still have a lot of them using the system, but something that struck me as very interesting is Atlas started to get picked up by a lot more, um, kind of like, I don't want to say less technical teams, but an archetype that's more of a business analyst, um, or someone doing business intelligence where, Um, we see, you know, consulting companies that get these data sets from their, from their customers, and maybe they have domain expertise, but they can't code. And so never before have they been able to interact with this data, uh, with this level of velocity and this level of granularity. And so we've seen, you know, a much wider adoption in terms of how technical the, uh, the actual users of the systems are than I originally anticipated, which has been really, really interesting to see.

AI assessment note: “be open to your, your thesis about core audience shifting”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q bearing in mind that, uh, you're a young startup and, uh, in a super fast moving industry, uh, how does that all work? Like those two pieces, how do they work together, including, uh, from a go to market perspective? Because GPD for all presumably is more of a developer audience. Is, is Atlas more of a kind of like enterprise audiences hugging face, like an enterprise customer for you?

A Yeah, so Atlas is definitely a B to B enterprise tool. Um, we make it, you know, very, we have very generous limits for individuals, like power users, academics as well, especially, um, to make it so that they can leverage it, but it is kind of the core growth engine behind the monetization of Gnomic. Uh, and I think that puts us in a really, really interesting position relative to a lot of other open source companies, because it means that we can keep GPT for all as this, like, Very pure, like, love letter to the community open source, uh, sort of project. Um, and we won't sort of be pressured to eventually find a way to kind of, like, squeeze it, squeeze it for dollars. Um, but I think the two interact also very well from a, from a funnel standpoint. So the kind of person that is interested in tooling to help them understand and curate massive unstructured data sets is the kind of person that's either one, using that to train models or two, generating a lot of that kind of data. And That's exactly the kind of person that is interested in, in tools like GPT for all. And so, you know, a lot of the, you know, first enterprise sales that we had with Atlas were inbound in the GPT for all discord. And that's, you know, maybe this is just what building a company in 20, 24 is, but our sales funnel is literally like Twitter to discord to enterprise sale. It's insane.

AI assessment note: “our sales funnel is literally like Twitter to discord to enterprise sale”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I remember you, you guys burst onto the scene with, uh, GPT for all, which, um, I remember like felt from the perspective of, of, of an outside, uh, you know, person looking at, looking at the project, you felt like a, like an overnight success and just like crazy open source traction. And, uh, that was an, a weakened project, right? Wasn't it like, what was the story there?

A I mean, that's the, the model that Zach was talking about training. Um, it's also funny how, you know, Everything can look like an overnight success, kind of, from the outside, but I think a key part of what made GPT for All so good is we had spent the last year building Atlas, and it was the first time we really applied Atlas to the task of, like, cleaning training data for a model, and so, you know, within a weekend, we could go in and, like, if you look at the original GPT for All paper, we show, like, oh, here's how you identify, like, the sections of the training data when a model's refusing to respond, and, you know, oh, like, Here's where you identify the parts of the training data where the model's giving one word answers. And like, obviously you want to remove those so the model that you're training doesn't learn those behaviors. And so it's one of those things where, you know, which I think is so true in startups where like the effort payoff curve is super non-linear. It's like very much like a series of step functions where you put in all this effort and then only when you pass some threshold, your, your payoff shoots up.

AI assessment note: “we had spent the last year building Atlas, and it was the first time we really applied Atlas”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Excellent. So that's, uh, GPT for all. Uh, but as you mentioned, there was, uh, actually a product before GPT for all called Atlas. What does that do?

A Yeah. So Atlas is a tool for exploring and interacting with massive unstructured data sets using only a web browser. Uh, it was born out of a lot of the work that my co-founder Andre and I were doing at this, you know, radiology, uh, generative AI startup, Rad AI, Where a lot of the work on the models involved exploring these massive data sets, finding places where maybe the doctors had made mistakes or, you know, the model had made mistakes, uh, looking for patterns in those areas, attributing that back to the training data, looking for distribution shifts, um, you know, in, in the actual production data. Uh, when we were at RAD AI, there was this super interesting, you know, global health event called COVID. You may have heard of it that happened. That was this massive distribution shift in our query stream and having, you know, The, at the time we didn't have the, the tooling like Atlas to, to deal with that in sort of like a very flexible way and it, you know, caused a lot of headaches, certainly. Um, and so effectively the interface that we sort of converged on is, you know, what would be ideal for this is this sort of massive scatter plot that you can run in your web browser where every piece of data in your data set is a point and two, uh, points are close together if the pieces of data are similar. And what we found is this, uh, sort of layout is particularly useful for…

AI assessment note: “Atlas is a tool for exploring and interacting with massive unstructured data sets using only”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That's, that's awesome. Why do you think people, why don't people do that? Is that, um, is that like a reproducibility kind of like, uh, is that, is that, is that, what is it? Is that marketing? Is that coming across as open source, but not being very open source?

A There's a couple of answers here. The, um, I think the, the truest answer is that the data is often the secret sauce. Like when people talk about like, oh, what is the secret sauce of your AI model? It's almost always the data. Uh, sometimes it's like clever scaling and, you know, you'll have, you'll have architectural tricks that move the needle, but like almost always it's the data. And so. You know, if a company wants to appear open source without maybe really being open source, they'll release the weights, not the data. Um, for some companies, you know, that we're now seeing, you know, you're seeing like Voyage pop up for instance, and, and Cohere is another one that are Really like investing heavily into like embeddings as a service. That's their IP. That's their core IP. And so, you know, for their business to work, they can't release it. Um, but then you start to get into these like very weird conflict of interest scenarios where it's like, okay, but if they don't release it one, like, you know, there's an audibility question on that model. And then there's another question, which is if, you know, they don't release the data and they don't have necessarily the tooling to actually comb through all of it. It's really hard to make sense of if the benchmark scores that they get are like super legit or not. Like maybe they're legit, but like, you know, you, you, you can't act…

AI assessment note: “if a company wants to appear open source without maybe really being open source”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Great, and maybe as a level set for, to make this interesting for everyone, um, so what was ChatGPT for all? What, what was the concept of the product? What did it do for the model?

A Yeah, so it started as basically just a reaction to GPT for becoming, you know, increasingly closed source. Um, you know, we as a company full of like people that basically met through open source and like have been open source contributors through our careers, We're like pretty frustrated with the fact that like this, you know, quote unquote open lab was like, yeah, we're not going to tell you anything about the data. We're not going to tell you anything about the training. Uh, and we'll get into this later when we're talking about benchmarking, you know, closed versus open models, but, um, it was born out of that frustration, I think. And so, um, the original inspiration was like, let's just make a sick open source model that everyone can use and, you know, show people how to do science. Right. Um, but really what it grew into now is sort of this like incredible open source ecosystem where you have these, you know, all these companies around the world now shipping models, uh, in the open source, like Mistral's doing really awesome work there. Hugginface is doing really awesome work there. Repl is doing really awesome work there. And we're now part of this kind of weird kind of, uh, almost, um, emergent, like assemblage of like hackers and like hobbyists that are like taking these models and like running around and quantizing them. Or like actually training or like slurping ne…

AI assessment note: “help to ensure that those models can run on a wide variety of different hardware”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Oh, there you go. VC money actually helping. Who knew? So you, you all have a, uh, relationship with LAMA CPP. Do you want to talk to that?

A Yeah. So this has been, you know, as Zach was talking about earlier, um, one of the biggest parts I think that was integral to the success of the original GPT for all is that it was quantized. We use LAMA CPP for that. And. I think one of the things that we've been really proud of here at Gnomec is the fact that we've been able to actually, like, bring staff on full time, um, to work on things like improving lawless CPP. And so, you know, right now, uh, we have Gnomec staff working back and forth with Gergi on, you know, getting new compute backends merged in to improve the, you know, support for these models across a wide variety of hardware. Um, and I think the lesson here for the broader open source ecosystem is, like, You know, as economic pressures start to apply and like, you know, the, the, the VC money starts to run a little bit drier, like staying collaborative and like making sure that, you know, as a whole, the, the open source ecosystem itself continues to be able to move forward. And like, you know, if someone else's repo, like is one that's useful, you know, in your projects, like being open to like collaborating with them and like being flexible about where those boundaries lie is like really, really important. And so. Uh, that I think is a lesson that I really hope all of the people building an open source right now take to heart.

AI assessment note: “bring staff on full time, um, to work on things like improving lawless CPP”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q you, you're now a three product company and product slash models. You got Atlas, you got GPT for all. And you got, um, uh, no meek, uh, embed. Uh, so how does that, uh, fit in, uh, especially the latter part? We talked about the other two, um, from a go to market perspective earlier, but like, how does that latter part fit into the role go to market model?

A Yeah. So one of the key portions of the Atlas pipeline is actually embedding all of your data. And so one of the things that we're excited to start doing is like, as we build up this, you know, uh, core competency of building, you know, these incredible embedding models, Uh, we can release them to the public and also incorporate them to improve the Atlas system. And so there's sort of direct product value, uh, and sort of direct audibility on the Atlas side of things that I think is really important. And beyond that, you know, going back to, you know, if you want to talk about go to market and how GPT for all fits into that, this idea of if we can do these things where, where we're really just like giving as many people as possible access to this technology, be it, you know, what you need for retrieval augmented generation with nomic embed, or What you need to actually get up and running with the generative models themselves with GPT for all, regardless of your hardware, that's just going to bring more people into the AI ecosystem faster. Um, and I think that demand sort of directly translates to, you know, uh, need and, and sort of like usage of Atlas, because again, as people are producing massive unstructured data sets, uh, and consuming massive unstructured data sets and tweaking these models and, you know, Really starting to learn the lessons that, you know, we constantly …

AI assessment note: “brings more people into the AI ecosystem faster.”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q company, like fast forward a couple of years or three years, um, is the idea that you guys are going to To be very prolific in releasing new products based on where the market's going, and then, and then what's the end result? Is there like a suite of, of different products that you're going to integrate into some kind of a platform? Like, how do you think about it?

A Yeah, it goes back to the core mission of, of Nomec. You know, our objective function is to improve the explainability and accessibility of AI models. On the accessibility side, we want to, as much as possible, continue releasing, you know, new open source tools, you know, for the community for free to bring more people Uh, you know, regardless of where they're at, regardless of, you know, what resources they have into, you know, the post AI world. I think this addresses, you know, what I would say the number one risk of this technology is, which is, um, the sort of unequal access element to it. Like the, the version of the world that I fear most is like, there's two or three mega companies that like lock everything down in like the early days. And so the sort of like, I don't want to say inequality necessarily, but the inequality induced by that only, you know, gets worse and worse and worse. That I think is like the number one future we have to be like watching out for. Um, and then the explainability side of things really is, I think, Atlas. Um, I cannot emphasize deeply enough how important looking at and understanding your model inputs and outputs are, um, you know, for actually building models that are safe and doing the sorts of things that you intend them to do.

AI assessment note: “On the accessibility side, we want to, as much as possible, continue releasing”

Redirected raw tape D 2 · C 4 · P 4 · Cm 3 3.25

Q problem and something that was like, yeah, no, you need this thing at, at, at a moment to, you know, to input data into the vector database, but it's no big deal kind of thing. And it seems like over the last few months, it's become like a, like a, Uh, huge area of, uh, activity and, uh, with a lot of people competing. So what happened to that space?

A Yeah, I think, you know, the fundamental thing here is, um, people realize that there's like a new data primitive, uh, you know, and that that data primitive is here to stay, and that's the vector of the embedding. And, you know, just to bring us back to Atlas for a second, in the same way that we see vector databases taking that new data primitive and trying to play out the, the implications of that at the database layer of the stack. I think we're going to see, you know, a radical shift like that at all layers of, of the business software stack. And the, the current conception of, you know, what I believe Atlas to be in the limit is the logical extension of what happens when you make, you know, embedding vectors a fundamental primitive at the sort of visualization or, or business analysis, business intelligence level of the stack. Um, so internally we love to throw around calling it sort of like a tableau for unstructured data.

AI assessment note: “just to bring us back to Atlas for a second”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.