Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q and like looking at all these new, uh, Hit new headlines is, is really helpful. But then, um, it only gets you, like, a very surface level understanding. Then you still need a process to decide which one to invest in. Um, so I'm, I'm trying to dig for, like, what is your formula for, like, deciding, you know, what to go deep on and what to kind of skip?
A From a practical standpoint, as a company, like, I already know the, there are, like, three to five Things that will be valuable and useful to us, and then there's other stuff that's like out of scope from, from, for different reasons. Some stuff is like out of scope from, um, hey, this is not going to impact or help us, and then other things are out of scope because we can't do it. You know, like the, the stuff, like different tech, so a really good instance for that is, um, specific algorithms for, um, you know, Improving extremely large-scale distributed training. Like, that's that, we're not gonna have the opportunity to get 2000 each 100. If we do, it'd be really cool. But, like, I'm just saying, like, as for now, like, you gotta, you gotta reach for the things that would be useful. Things that would be useful for us, for instance, um, are, for everybody, actually, to be honest, is like, um, evaluations, uh, different post-training techniques. And then synthetic data, uh, uh, construction. Like, we're always on the, I'm always on the look for that, and then how do I figure out whether these things, um, you know, which new piece of news is actually novel? Um, well, that's sort of my, like, mental cache to a certain extent. Like, I've built up, like, this state of, like, I already know, like, all the things that have already been written, uh, for the state of the art, uh, fo…
AI assessment note: “I already know the, there are, like, three to five Things that will be valuable”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q topic, um, did you have any, um, exploration of synthetic data at all? Um, you know, use, use Mistral to rephrase some existing part of your data set to generate more tokens, anything like that, or, or any other form of synthetic data that you, that you choose to mention? Um, I think you, you also mentioned the large world model paper, right? So, um, yeah, yeah, anything like that?
A Yeah, yeah. So, um, yeah, we, we used, like, GPT-IV to, uh, rephrase, uh, certain aspects of, um, the chat data, um, reformatting it or, uh, kind of generating new types of, uh, tokens and language and, you know, types of data that the model could see. Um, and, uh, Also, like, trying to take the lower, uh, correlated instances of out-of-domain data and, and that we wanted to inject it to the model, too, as well. So, um, I actually think a lot of the moat is in the, the data pipeline. Um, you, you'll notice, like, most papers just don't really go into deep detail about the data set creation because they probably know, I mean, there's, there's some aspects that are, like, uninteresting, right? Which is, like, we paid a bunch of people and, like, generated a lot of good data. But then the synthetic data generating pipeline itself, that, that, you know, sometimes that could be, like, 25% or, or 50% of the entire data set that you've been used to do pre-training.
AI assessment note: “we used, like, GPT-IV to, uh, rephrase, uh, certain aspects”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q cool. Well, you can leave that in the call to action at the end. Um, I just want to, you know, we have a couple more questions to round this out. Um, you mentioned a lot of papers in your work. Uh, you're also building a company. You're also looking at open source projects and community. Um, what is your daily or weekly routine to keep on top of AI?
A Ah, um, So, one, subscribe to AI News. He didn't have to pay me to say that. I actually really think, like, it's a good aggregator. I, I think it's a good aggregator. I'll tell you why. Most of the fastest moving, like, um, research that's being done out there is, like, it's showing, it's mostly on Twitter. Like, my Twitter is, like, I, I wasn't a power Twitter user at all before three years ago, but I had to use it, and I had to always check it in order to keep on top of, like, early work, right, that people wanted to talk about or present, because Nothing against, uh, submitting research papers to, like, ICLR or ICML. Like, no, knowing the state of the art, like, those are, like, six months, uh, uh, late, right? Like, people have already have it, dropped it on archive, or they're just openly talking about it.
AI assessment note: “So, one, subscribe to AI News... I had to always check it”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q non-deterministic and it's, it's hard to control. Um, yeah, I mean, so like, you know, I, I think what's the, uh, founding story of Gradient? Like, uh, how, you know, of, uh, of, of all the problems that you chose, like, uh, why choose this one? Um, you know, uh, how did you get together your co-founders? Anything like that, that, that, that bring us up to the present day?
A One of my co-founders is Chris, and he's, he's a really good friend of mine as well. I don't know if you intersected with him at Penn as well, but, um, yeah, yeah, Chris Chang, who, he was, he was at Penn as well, did banking for maybe one or two years, and then, um, you know, was a, uh, software engineer at Meta, um, also was at Google, and then most recently he was like a director at Netflix and product, and, uh, We always wanted to do something together, but we felt the, you know, what really came to fruition was wanting to develop something that is enterprise facing for once, um, mostly because of our experience with internal tooling and, and, and inability for something to like, uh, basically, um, exist through like a migration, right? Like all, all the time with every ML platform, That I've ever had to experience, or he had to experience. It's like a rebuild, and you rip it out, and you have a new workflow or automation come in, and it's this huge multi-quarter, maybe even multi-year, uh, project to do that. And, uh, we also teamed up with, um, a former co-worker of Chris's from Open Door Forest, uh, who's also, um, on Google Cloud Platform, and, um, you know, him seeing the scale Uh, and actually the, uh, the state of the art in terms of Google was using AI for systems before everybody else too, right? They invented a transformer, and their internal set of tooling was ju…
AI assessment note: “One of my co-founders is Chris, and he's, he's a really good friend”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Yeah, you mentioned, uh, quite a few. You mentioned Ruler, Lugo, Infinite Bench, Bamboo, Zero Scrolls. Um, like, do you wanna, do you wanna give us, like, uh, maybe two or three of, of those that, that you thought were particularly interesting or challenging, and, uh, you know, what made them stand out for you?
A There's just so many, and they're so nuanced that I would say, like, yeah, Zero Scrolls was the first one I'd ever heard of, uh, coming out last year, and it was just, like, extent, like, making, it's, it's got, it was more of, like, tracking Um, uh, variable over long context. Um, I'll, I'll go into Ruler, because that's the freshest in my mind, and like, we're just scrutinizing it so much and running the evaluation in the previous two weeks. But like, Ruler has, um, basic, uh, has four different types of evaluations. So the first one is exactly needle in the haystack, except you throw multiple needles. So you gotta retrieve, you know, multiple key value pairs. There's another one that basically You need to differentiate.
AI assessment note: “I'll go into Ruler, because that's the freshest in my mind”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q on Crusoe since you just brought it up. Um, yeah, like, can you explain what Crusoe is? Yeah, you know, I, I have this mental image of putting GPUs on top of oil rigs. Uh. What, what, what is it? What do they do? Uh, how do you work with them? You know, just any, anything, anything nice. I'm sure they appreciate nice things that you say about them too.
A Oh, for sure. For sure. Um, so, uh, they, they came to us in, in, through a collaborative effort where we basically were in search of, um, uh, cloud, you know, a GPU provider. I don't want to call cloud service provider quite yet because then, you know, you think about the hyperscalers, but for them, uh, you know, they're one of the biggest alternative GPU, uh, Uh, cloud providers, and, um, they were offering up, like, we want to do a collaboration to showcase, uh, their technology, and, uh, it, it just made it really easy for us to, like, scale up with their L-Forties, and those are the specific, uh, GPU instances we used, and, um, uh, coordinating that effort with them to get, you know, that dedicated cluster, uh, first to do the project, um, it, it, Became a really good relationship, and we still work with them today because, like, we're trying to evaluate more of these models and possibly train more of them, and anyone could go up to them and, and basically, um, get your compute from them, and they, they have a lot of, uh, GPUs available for the, for those type of projects.
AI assessment note: “they're one of the biggest alternative GPU, uh, Uh, cloud providers”