The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

1,847exchanges match
1,797on raw tape
133redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And isn't cooling technology, uh, that is being used as something that's, uh, well understood and it's just getting deployed, or is there like fundamental new things happening in cooling right now?

A I think liquid cooling has been around for some time, ah, but has never been deployed at this scale, and so the innovation is more around how to make it reliable, how to make it, ah, cheaper, ah, more scalable, right? So there's a lot of innovation around that. There's also a lot of new innovation, new kinds of liquids, new kinds of materials that can absorb heat better, ah, because anything that can improve the efficiency of heat transfer Uh, is very important for data centers. So we can then run the chips hotter, right? And there is a direct correlation between running a chip hotter and how powerful the compute is. So the hotter the chip, the more memory bandwidth you get, the more flops you get. And so there's more, there's a strong payoff. If you can cool well, that also means you can produce more intelligence.

AI assessment note: “innovation is more around how to make it reliable, how to make it cheaper”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Is what you just described, uh, called latent MOE, or is that a different concept?

A Latent MOE is a specific, uh, innovation that we have in NemoTron three family. And, uh, what it does is actually, um, reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token produces a vector and the idea is like, we're going to take that vector and learn a way to compress it. And then send that compressed thing through the network, and then we're gonna uncompress it at the other end. And as a result, we save on network bandwidth, and we also get four times the number of experts for the same inference cost. So you could think about it as like, you know, our library of books got four times bigger, um, and we get to, you know, read four times more books, uh, at the same inference cost because of, because of this particular innovation.

AI assessment note: “Latent MOE is a specific, uh, innovation that we have in NemoTron three family.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Yeah. That's the highly optimized, uh, storage. What else? The networking part and what other pieces?

A So, so I, I, I was talking about this one cloud cluster product that We've got, and the, the way of just for everybody to think about this is like, okay, well, look, you've got a bunch of GPs. Let's say you've got a cluster of 10,000 GPs. Well, I want to partition that cluster up, and so what it is, is it's a bunch of GPUs, some CPU servers as well, because you need to have an orchestration, uh, fleet as well, and then you've got some storage, and, um, all of the CPU servers and the storage servers and the GPU servers are interconnected with the, the storage, so they can quickly read and write from it, and, um, so there's And that, that communication happens over what's called, you know, the in-band network. And then there's the compute fabric, which is where I was talking about where all of the sort of weights and, uh, feature activations are being shared, uh, throughout that compute fabric. And then there's an out-of-band monitoring network where you've got access to whether it's BMC or, uh, some of your DPUs. Um, and When you are trying to create a sub-partition of a 10,000 GPU cluster, you need to simultaneously partition the in-band, the out-of-band, and the compute fabric. Okay, so, like, that complex coordination between we've got a bunch of bare metal systems to, hey, we've got a virtualized system that has, you know, what's called RDMA, you know, RDMA, remote direct me…

AI assessment note: “simultaneously partition the in-band, the out-of-band, and the compute fabric.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Do you want to talk about, uh, mega kernels and Together Atlas for a minute?

A Together, uh, mega kernels, Together Atlas, these are both, um, projects along these lines. So let me dive into the mega kernels first. To understand this, uh, the, the first thing when we say kernels is we usually mean we are going to write a specialized GPU program for a single operation in a model. Um, you can think of a model as one of these train models is like ABC different operations in a row, and there'll be hundreds of these. And the way that we've been writing kernels for the whole history of, let's say call it NVIDIA hardware, is that you really specialize a single kernel for a single operation. With these mega kernels, we're doing something quite interesting, which is We can take the entire model, however many billions of parameters and put it into a single GPU kernel. Um, and with that, you can start to do a lot more fine grained optimization than you were able to do before. Uh, it actually starts to make the NVIDIA GPU look a little bit more like a Cerebris chip or look a little bit more like a Samba Nova chip in terms of the, the optimization that you're able to do. And this is really critical at inference time. So we're able to see two X, sometimes three X speed ups, um, over even highly optimized inference engines. Um, so we're working on bringing that to, um, uh, to, to really work in production, bring, bring it to fruition and use it across our whole stack. T…

AI assessment note: “these are both, um, projects along these lines. So let me dive into the mega kernels”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Great. We'll go into, um, all of this in a, in a minute, but before doing so, uh, we, we alluded to, to some of your background. Let's go into it. Starting from the, from the beginning, what was your path to becoming a top researcher?

A I grew up in Russia, in Moscow. Starting from like middle school, high school, I was really interested in mathematics, and I was thinking I will be a mathematician or, or engineer of some kind. I was interested in machines, uh, and, and eventually, uh, computers, and I got into an undergrad in computer science, and I was still thinking that I'll be doing some kind of theoretical, you know, applied linear algebra, tensor methods, things like that. Uh, but at some point, I, Kind of discovered, uh, machine learning. There was this, uh, professor that we had, uh, Dmitry Vetrov, who had one of the, like, leading labs in machine learning in Russia at the time, and I was lucky enough to join that lab and start doing some research on machine learning in my undergrad. So that was around 2013, maybe 2014. I initially was working on non-neural network machine learning, uh, methods, so Gaussian processes. That's kind of By now, you know, nobody really talks about that anymore, but eventually I, I got into a PhD thinking I would still be doing a Gaussian process, but I ended up working on deep learning, and that was actually quite, I'm happy that I didn't work on Gaussian process. I worked on some things related to kind of core machine learning, methodology, optimization, probabilistic methods, questions related to generalization and how the models learn features. After I finished my PhD, I…

AI assessment note: “I was lucky enough to join that lab and start doing some research on machine learning”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And then, uh, within that world, your work specifically has focused on weak to strong. Uh, can you explain what that is?

A That is, uh, the project that we did, uh, back at OpenAI. That Work was focusing on the future scenario when we will be trying to align models that are above our own capability on certain tasks. Already now, if you take the frontier LLMs, they are extremely capable, and on a lot of domains, we need expert humans to be able to tell which responses are good, which are correct, which are not correct. But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model. So imagine you ask it to make a repo for you for some, like, you know, new startup idea and just implement it from scratch entirely, and then it gives you, you know, 10,000 lines of code. You have no way of checking if all of this code is correct, if all of this code is safe to use. And so that's the problem of supervision. We are moving to this future when Like it's very hard for a human to supervise, uh, the models, uh, directly. And so instead we studied a simplified setting where we used a small model to try to supervise a larger model.

AI assessment note: “we used a small model to try to supervise a larger model”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q So I'd love to do a little bit of a deep dive slash educational part on, um, the whole reasoning model aspect, uh, because, uh, as you just mentioned, since it's so new, uh, some people truly understand how those work. Many people Don't. At a very simplistic level, what is a reasoning model and how is that different from, um, your sort of base LLM?

A So a reasoning model is like your base LLM, but before giving you the answer, it, it thinks what people call in the chain of thought, meaning it generates some tokens, some texts that's meant not for you to read, but for the model to give you the better, the better answer. And while it does this these days, it is also allowed to use tools. So it can, for example, in its thinking, so-called thinking process, go and browse the web and to give you a better answer. So, so that's the superficial part of the thinking models. Now, the deep part is that you start treating this thinking process as part of the model, basically. So it's not something the model generates and it's an output for you. It's something you want to train, right? You want to tell the model you, you should think well, you should think so that the answer after this is good in, in whatever way. And this leads you to a very different way of training the model because models were in usually trained with just gradient descent, the way deep neural networks are trained, meaning you say, predict the next word and you do a gradient, you Differentiate your function from the model. They're not fully differentiable, but you approximate it, and you train your weights to, to do that. And that, it was quite amazing that doing just that, you could make a chat. But with the reasoning model, you can do that because there is this rea…

AI assessment note: “a reasoning model is like your base LLM, but before giving you the answer, it, it thinks”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q uh, we've done a bunch of great episodes recently with Sholto from Anthropic, Jerry from OpenAI, and then, uh, Julian from Anthropic, uh, if you're curious to learn more. Let, let's move on to the business, uh, of AI. Uh, you mentioned in the report that the business of AI Finally caught up with the hype. What caught your attention in terms of fact stats in the last 12 months?

A Yeah, a couple of them. Again, like, where we came from one or two years ago was just tons of money going into this segment, building models, a lot of usage, but not clear where the revenue would come from. I think it was maybe OpenAI was making fifty million dollars or something two years ago, and it was very Unclear how they would ever hit like billions of revenue. Um, and nowadays, I think if you sum sort of the top 20 or so, uh, major AI companies from the labs to the most popular kind of vertical applications, you know, across them, they're making tens of billions of dollars of revenue. Um, you can look at the smaller scale companies, which, you know, are growing from zero to twenty million or twenty million plus, uh, as a group, they generally grow about 60% faster on a quarterly basis than non AI companies. Um, we've all seen like the famous charts about ARR or non ARR. It's unclear, uh, but, uh, you know, very steep curves for various coding companies. Um, and perhaps most interestingly across a segment of 43,000 or so, uh, US customers, we work with ramp to show that retention of, uh, subscriptions on AI products across this customer set has really improved markedly since 2022. On 2022 is around the 50% after 12 months. Uh, and now in 25, it's hitting around 80%. Um, and the second stat in that analysis that was interesting was the total spend on AI products, uh, per c…

AI assessment note: “retention of, uh, subscriptions on AI products across this customer set has really improved”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q riff on that theme a little bit, uh, IP rights, safety, regulatory, uh, a little bit like the sustainability, uh, thing that we were discussing earlier. It sort of feels like that, that whole world, uh, as, um, Sort of slow down in terms of, like, progress, maybe starting with regulatory. Do you think that regulatory is anywhere near catching up or providing an adequate response to what's going on?

A Yeah, I'd say like a big one on that one. I mean, clearly the Trump administration unwound a lot of the Biden era policies, uh, whether that was on, uh, diffusion, you know, trying to push a lot of state level legislation against AI. The, uh, over in Europe, like the EU AI act has had, um, delays in implementations, only three member states that have actually implemented it. And now we're finally seeing how even its authors are saying, uh, maybe we went too far, uh, particularly as we look at progress, Uh, the speed of progress in the U S and China compared to Europe. Um, you know, famously this bill in California, um, you know, rate limiting AI progress was really watered down into what eventually became SB. Um, there were, you know, many, many proposed bills, I think over a 1010% of them actually made their way into laws. Um, so it's still kind of patchworky, but like at a meta level, it looks like we traded regulation for just going faster. It was perhaps like best encompassed by, uh, by the shift between the AI safety summit and the UK, which was at Bletchley, which basically pledged like a whole network of, uh, AI safety institutes and conferences that would happen over the coming years, um, to then the subsequent event in Paris, which was called the AI action summit, completely different than AI safety summit. And JD Vance saying something along the lines of basically lik…

AI assessment note: “at a meta level, it looks like we traded regulation for just going faster.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q your ability, uh, to truly understand, uh, how models work. So the general field of mechanistic interpretability, you alluded to the fact earlier that, um, RL, if I understood correctly, uh, sometimes make it a bit harder because it does, uh, things occasionally in a more inscrutable way. My words, maybe not, not yours. Um, so what is the latest and, uh, indeed does RL make things harder or easier?

A Oh, so what I meant before is that debugging RL in general, you know, completely unrelated to interpretability. It's harder because there are more moving parts. But it is also true that if you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts, and then you could also have a thought that, oh, maybe I should use that as a reward signal in RL and punish the model if it thinks the wrong thing, but then suddenly you completely destroyed your interpretability angle, so you sort of have to be careful that, yeah, you don't Do RL on the signals that you actually want to use to interpret what the model is thinking of doing. That said, I think, yeah, there are some extremely exciting interpretability things happening, including mechanistic interpretability. I think, like, actually, last year, I think before John Anthropic, maybe even, there was a super cool Golden Gate Claude model, where, you know, they found the neurons in Claude that were responsible for the Golden Gate concept, and then modified them to make a version of Claude that really loved the Golden Gate Bridge in San Francisco. And so that's like a really vivid example of, oh, you know, we really understand what's happening in…

AI assessment note: “if you're not careful with RL, you can make interpretability harder.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And because it's bird base, was that less of a massive compute data crunching effort, or was it still intense?

A I mean, uh, less of yes and still intense yes. Like definitely wasn't all smooth on the infrastructure side. Um, we had to build a custom tokenizer and optimize it for Stripe events. Um, we had to scale our data pipelines to grow to the very large data sizes. I mentioned earlier, like previous models just hadn't trained on such large amounts of unstructured data all at once. We also just like had to Build custom data loaders to make sure that GPU utilization was high. Um, earlier versions actually resulted in like pretty low GPU utilization because data loaders, you know, became the bottleneck. And so, yes, that made training more expensive, but also it made it slower. And so, yeah, I mean, this was, this was something bigger than we trained before. We had to add a bunch of checkpoints to make our runs more robust, intermittent failures, you know, the kind of stuff that you would be doing anyway if you were like an AI lab, but we are not first and foremost Um, and AI lab. And so, um, those were, those were all, um, sort of progressive builds for us.

AI assessment note: “less of yes and still intense yes. Like definitely wasn't all smooth on the infrastructure side”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Let's unpack some of it. So, let's start with features. You just mentioned Canva code. You guys had a very impressive last 12 months in terms of shipping velocity, and I wrote down Matic Studio, Dream Lab, Canva AI, Canva code, Canva sheets. Maybe give us a little bit of a Tour of what those different things do at a, at a high level? Uh, so Magic Studio to start.

A Yeah, so Magic Studio is a collection of AI-powered tools, uh, both text and visual. Uh, there are text-based tools that allow you to, for instance, um, uh, modify text to be in your, in your particular voice, or, uh, in a, in a brand voice, which is really powerful. There's text summarization, text expansion, Uh, we've got a whole suite of image editing that's really powerful, whether that's infill or outfill. Uh, we've got Magic Grab, which is an amazing feature, which basically allows you to treat a rusted image as, um, something that you can decompose. You can pick objects up and move them around. Uh, Magic Arrays uses similar technology to be able to remove and then infill, um, in a very smart way. Um, We did launch a Sheets product in, um, in April. Uh, that's our take on, uh, on, on a spreadsheet product. It is designed to be, uh, very, very easy to use, uh, fully integrated into that, into the visual communication experience of Canva. Uh, and it will be our kind of data backbone going forward. So when we want to do more data-driven products, we'll be using Canva Sheets as the kind of data source for that.

AI assessment note: “Magic Studio is a collection of AI-powered tools, uh, both text and visual.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q have like come out of nowhere in the last, you know, sort of like five, six months in, in particular, like for, from your perspective that, um, you know, is somewhat obvious idea that was sort of, um, you know, it was a matter of time until, uh, it was going to be implemented and, um, You know, it's good, but not completely groundbreaking. Or is that a major development?

A Well, it's been worked on for a very long time. So we've known it was coming for, um, for years, for years, for the past three years. Um, and it is kind of obvious. It is kind of obvious because if you think of like the pre reasoning world, um, The input space to a language model is everything. It's all of language. Uh, and so you can ask it very simple questions like one plus one or extremely complex ones like, um, go cure cancer. Like those are two strings or requests that you can ask it, and you really don't expect it to spend the same amount of energy and time on those two different problems. Um, One, it should respond immediately. The other one, it should probably, uh, you know, it might take years of thinking and trying to accomplish. Um, but we didn't have that reality before reasoning. Um, we had an input and then an immediate response. And so both of those got the same energy and effort put into them. So it had to come at some point, this notion of different amounts of energy or time being spent on problems, test time, compute, um, I think the effectiveness of it was surprising. Like it was really quite incredible to see how much gets unlocked. Like these models actually can with very little supervision, uh, very little data from humans saying, this is how you think through problems, sort of figure it out for themselves. Um, so that's been incredible to watch. I think …

AI assessment note: “It is kind of obvious because if you think of like the pre reasoning world”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q anyone that may be curious about the world of data infrastructure, and then we can go into all sorts of, uh, technical details, but like to start with, what world do you operate in if, if we think of all of this as, as, uh, you know, databases, so databases, data warehouses, how would you sort of compare and contrast the, the, the, the various, uh, databases of the world?

A So I think of the world, uh, the database world As being really divided into two halves, uh, the analytical side and the transactional side. The analytical side are systems that are built to be very read optimized, so reading data as fast as possible. The transactional world being more oriented towards writing data, uh, fast and consistently. So when you think of a transactional system that's maybe powering an application, you know, a, a, a canonical example would be like a, an ATM that needs to record a debit and a credit Uh, very quickly and has to be consistent every single time. An analytical application would be something like, uh, how many customers bought product X last, last year? And slicing and dicing the demographic profile of your customers and understanding their journey, you know, through the website all the way to a transaction at the end. Um, and so those are, those are broadly speaking the, the, the two worlds.

AI assessment note: “database world As being really divided into two halves, uh, the analytical side and the transactional side.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q friend, could be foe, depending on how you describe it. And, and, and where you are in the stack, but certainly from a frontier model perspective, the fact that, uh, you know, Meta is putting all its weight behind, um, something like Llama four, uh, has to be frightening. So that's how we think about that general purpose model. Now, maybe let's talk about specialized models. What do we think?

A Yeah. I mean, I think that's probably a more interesting area, at least from our, um, perspective and how we think about the world. Um, you know, Having models that are specialized on either a specific modality, so, you know, image, video, audio, et cetera, um, or on a specific vertical or industry, you know, think pharma, life sciences, material sciences, um, or even generally automating specific use cases across, um, horizontal industries. So, you know, thinking, you know, for a minute about some companies that, that we've been fortunate to partner with, um, Synthesia and kind of the first category around, um, you know, AI for generative video. Um, and then on the enterprise automation side, we also were fortunate to partner with a company called H, um, based out of Paris.

AI assessment note: “Having models that are specialized on either a specific modality”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And so you're fully distributed, but, uh, you have a large office in Paris these days. Is that, is that fair?

A Yes. So we're a hybrid company, uh, so people can work from anywhere in the world. Usually when there are a couple of people, three, four, five people, they take an office in the city that they are at. So I think we have maybe, uh, 1213, uh, small offices all over the world. Uh, the biggest one being Paris, where we have, I think around 20, 25% of the team based, based here. The three founders of Hugging Face are, are French. So we have quite, quite a lot of, uh, French roots, uh, but we consider ourselves like an international company. So not an American company, not a French company, but an international company with people from all over the world.

AI assessment note: “Yes. So we're a hybrid company, uh, so people can work from anywhere”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q All right. So going from the infralayer of the cake, Or hardware layer of the cake to, uh, the software layer? Maybe let's start with foundation models. So sort of the same questions, like those are huge, um, kind of dollar at play. Is that, uh, something you're interested in?

A We're, so we study it a lot because it matters, but I, okay, so here's the thing, right? You look at like Lama three, eight billion, trained on 15 trillion tokens. The, I run it on my Mac book. It's awesome. Like latency super low. It works really well. It's basically on par with the Mistral eight by seven billion mixture of experts model. And I've stopped going to open AI because I can run on my machine. It's faster and it's integrated in my email client. And they're like, there are a bunch of advantages to it. And so, I, I think the small language models, particularly for B to B applications, if they continue to improve their performance like this, make a lot more sense. Uh, the, the, we ran this analysis, I mean, if you look at the pricing page for, like, OpenAI, and you look at the 4.5 turbo, and the cost per inference, and compare it to the next most recent model, there's a 160 X difference in pricing. And so what does that tell you? Well, a model that's six months out of date loses its pricing power rapidly, right? Commoditizes really fast. And then you have this open source dynamic with Lama where Meta and OpenAI are, are, are competing. So I think we'll see a lot of innovation there. I think we'll probably see a bifurcation of the model sizes where like the Lama three, eight billion parameter model on an MLU basis, which is the high school equivalency is doing pretty we…

AI assessment note: “we study it a lot because it matters”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q now as, uh, with your theory ventures hat on, uh, which is really interesting to me because BI of all parts of the modern data stack is really To me, the one that sort of felt like the kind of like unloved child where, where you've seen less innovation. Uh, so what's the story about Omni and where do you think, I guess, uh, BI, modern BI should be going?

A Yeah. So, yeah, a sixteen billion dollar category. You're right. It's a super competitive, uh, hard to differentiate category. I think it's broadly misunderstood. Um, but the, uh, the way that we think about it is it has swung between a pendulum of control to Um, empowerment. So, uh, during the 2000 era, there were four centralized BI companies, MicroStrategy, Cognos, Business Objects, and Hyperion. And then Tableau came and unbundled the visualization layer. So it swung from really, really tightly controlled, centralized control to anybody can do whatever they want with their data. Then the cloud data warehouses came and Looker said, well, let's go back to centralization with the data modeling. And what Omni is trying to do is Narrow the, those swings and say, you can have the control and the modern data model And an individual marketing analyst who wants to define a particular kind of cost of customer acquisition can use an Excel-like UI to do that. If the metric is awesome, they can say, I want to promote this to the team, to the group, or to the entire company, and then that folds into the underlying data model. So the data team can define metrics, but also an individual person in the company can define metrics, and their approval flows to marry the two.

AI assessment note: “what Omni is trying to do is Narrow the, those swings and say”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q learned, you know, for all the founders who may be listening, uh, right now that, you know, maybe part of the group of companies that, um, You know, you said should be acquired by Data Break of the Snowflake. Like, how does, how do, how do acquisitions happen? Is that something you can engineer? Is that something that you need to position yourself for? What have you seen and learned?

A Yeah, so the, the answer is, um, someone's putting their career on the line. Uh, so when someone buys a business, not an acquihire, but let's say like 50 to fifty million plus, someone in a company, in the acquiring company is saying, I am betting my career, effectively, in this company, that we should buy this company at this price. And that doesn't happen overnight, right? It doesn't happen in the course of like two or three months. It probably happens over the course of 12 to 18 months. It's a big, long enterprise sale. You can think about it that way. And There, it is a multi-party sale, where typically you have a GM or a head of product who decides, I really need this business as part of the product portfolio, and rather than building it, which would be much easier politically to do, to navigate that budget, I think we should buy. And I'm willing to make a case and go first to my manager, then to my VP, then to the board, To justify it, involve legal, corp dev, and many other teams, and manage what must be a very difficult cross-functional effort to get this through. And so somebody really has to care, there has to be a very strong personal relationship between the founder of a business and that buyer. Because that buyer is putting together a three to a five year plan to demonstrate some positive ROI on that acquisition. So they'll take the startup's plan and then discount…

AI assessment note: “someone in the acquiring company is saying, I am betting my career”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Okay, so maybe to, uh, double click for the more technical people, um, in the, in the audience. So, like, how does that, how does that work? How were you able to, um, sort of do this no compromise kind of approach?

A Sure. So every scale out system that I am aware of before VAST is based loosely speaking on a concept called shared nothing sharding. You have a lot of nodes in a cluster and each one of them has direct access to a bit of the namespace and each one of them has responsibility for a piece of the pie for a piece of that namespace. And we realized that that architecture is reaching the end of its rope. Um, it starts to see diminishing returns in performance as you scale beyond a certain limit. Uh, resilience is very problematic. If a node fails, you need to recover that node's responsibility, and it can take a week, and during that time, you can't have another failure, um, and so that limits the scale of it. What we needed to do is the opposite. Instead of direct attached having drives in the nodes, we disaggregate. We put the drives on one side of the network, we put the logic on the other side, and we leverage a new protocol called NVMe over Fabrics to make it look like All of those drives are directly attached. That allows us to scale capacity independently from performance. It allows us to scale with dislike parts over time, so you never have to migrate your data between systems. More importantly, it allows us to move from shared nothing, sharding, to shared everything. Every node can now see the entirety of the data set. Data on low cost flash, metadata on storage class memory…

AI assessment note: “Instead of direct attached having drives in the nodes, we disaggregate.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q How did you Uh, bootstrap it initially, like to get the, the first few tens of thousands or however many you needed to start?

A So when we started, we, the, to train the prostate AI, even before we launched Ezra, uh, we used the public data set available, um, called the prostate X data set that the NCI put together to build the prototype, to show to investors that we actually know how to build an AI and so on. When we launched our first full body scan, it was about 75 minutes. It was not yet using any AI, and we were just, um, selling it as a, as a high quality full body protocol without AI in order to build the data set to be able to train the eyes. And then we got our first thousand full body scans in the first year or so, and that enabled us to start, uh, training, and then we've since obviously grown that many, many, uh, uh, faults.

AI assessment note: “we used the public data set available, um, called the prostate X data set”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q of UX? Um, so you want to give the, the, the users, or your customers, users, um, you know, a choice in terms of like how they navigate and how they get to the answer. Um, is the idea that, uh, you can almost like pick AI versus a human. You can fast forward to human directly. Uh, how does that, how does that all work and any early lessons?

A In our manifesto, we say, like, we believe the future of support will be humans plus AI. And what we mean there is the AI will, uh, so basically some conversations will be answered entirely comprehensively by AI. How do I reset my password? Here's how. Uh, how do I get a refund to click this link? That type of thing. Just a complete answer. Customer gets exactly what they want and they leave. The next set will be things that the AI will attempt to answer, but ultimately might fail over to a human. And then some, the AI will just be like, I'm not touching that. It sounds like a sales query. I'm going to hand that straight over to a human, right? So that's the first piece. The second is there's a symbiotic relationship between the, the, like the bots and the humans, if you like, right? The, or the AI and the human, which is when are, you know, humans are in the inbox, the AI can help them. It can do things like summarization. It can like, you like, look stuff up for them and all that sort of stuff where we're investing a lot there. But then also the humans help the AI because when, when you, you know, one of the features of Finn is you can go through all the answers it's given and be like, oh, you got that one wrong. Let me teach you. And you can give Finn new facts to learn from so that it doesn't make its mistakes again. So I think you'll see this kind of relationship where lik…

AI assessment note: “we believe the future of support will be humans plus AI”

Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q I saw an answer, something called tools. Did we cover that already? Is that, is that, is that, so we talked about product engineering, we talked about fine tuning.

A So we haven't actually covered tools yet. So, you know, LLMs by default are text in text out. And so they're limited in what they can do in terms of taking action in the world. And the idea of tools. Or what OpenAI originally called function calling and we sort of called it tools. And now I think we, we won the war on the naming because everyone, they renamed it to tools now as well. But essentially what this is, is that if an LLM wants to take an action in the world, you essentially allow it to do that by exposing a set of APIs. So a set of different programs that the LLM can use. And you say to the model, Hey, these are the tools that are available to you. So for example, web browsing might be a tool that's available to the model. And the model then can output a request to that API, which is just a JSON string. So it's just another piece of text. And that JSON string will specify which tool it wants to use, what question it wants to ask of that tool or how it wants to use that tool. That output is then taken, used to run the tool itself. So maybe you'll go and do a web search. The result of the web search is then passed back to the model and the model then uses that output to make another decision. And so suddenly you go from a system that's only text in text out to something that can be action taking and that can actually take advantage of external APIs in the world or do in…

AI assessment note: “So we haven't actually covered tools yet.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And in terms of channels, so you're particularly prominent on LinkedIn. Uh, again, if I'm a startup, uh, doing technical data science stuff or AI stuff, uh, would you recommend I start with, with one channel and just focus all my energy on that, or do all the things that you mentioned, so like YouTube and podcast, or like, how do I, how do I start?

A I would repurpose content as much as possible to make it as easy on yourself as you possibly can. So if you're already writing blog posts, Why not convert that into some sort of newsletter that you can put on LinkedIn and on Substack, right? You're literally just copying and pasting stuff, and you're kind of trying to reach a couple different new people, but you think about it as just like top of funnel inbound. You'll deal with it later and see who comes through those channels once you get a little bit more established, right? You just kind of trying to start someplace. Um, You already, you have a newsletter, right? So cut it up into some smaller pieces for some LinkedIn posts as well, right? Like take some paragraphs, take some high level learnings, ask people some questions. LinkedIn is a great place to get to know your audience as well. People are really involved in the comments and leave thoughtful comments from my experience. So that's where you can start to ask people about what they're excited about, what part of this resonated with them. You'll notice who responds to what kind of content, where on your LinkedIn. Now you've got that kind of going, let's say you found some LinkedIn posts that worked really well. Well, why not hop on a video and read it and read through some of the comments and discuss what you learned. Now you can edit that into short form clips for Inst…

AI assessment note: “I would repurpose content as much as possible to make it as easy on yourself”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Yeah, that's great. Let, let, let's get into this. Like what, what are the five pillars?

A Uh, yeah. So, you know, first is freshness. Um, so freshness is relating to, um, Uh, uh, the freshness of the data. So for example, you know, it talks about, uh, media companies, you know, you can probably think about e-commerce companies or even FinTech company that relies on thousands of data sources, um, you know, arriving, let's say two to three times a day. Um, how do you keep track and make sure that thousands of those data sources are actually arriving on time? There has to be some automatic way to do that, but that's sort of a common reason for why data would break. Um, so freshness is one. The second is volume. Um, So pretty straightforward. You know, you'd expect some sort of volume of data to arrive from that data source. Has it arrived or not? Um, the third is, uh, distribution and distribution sort of refers to at the field level. So let's say, um, there's a credit card field that is getting updated or social security, um, number field that gets updated and suddenly it has, um, uh, letters instead of numbers that would obviously be, um, something is, is incorrect. So you actually need tests for that at the, at the field level. Um, The fourth is schema. So actually schema changes are a big culprit for data downtime. Um, oftentimes there's engineers, um, or other team members actually making changes to the schema. Uh, maybe they're adding a table, changing a field, c…

AI assessment note: “first is freshness... second is volume... third is distribution... fourth is schema... fifth”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And maybe double clicking on the, on the latter part. So governance, like why does governance matter?

A Yeah, totally. I mean, governance matters, uh, in, in those two ways, right? The first one is around productivity. So your analysts and your product managers and anybody who's going to make use of data during their job is effective at using that data. They may have skills like writing, being able to write a SQL query or interpret a dashboard that are generic, but they don't have the organizational context. And on governance, on the other side, which is compliance with regulations, Those are actually, um, the domain dependent. So there's the regular, uh, regulations here, like GDPR and CCPA that apply to almost all organizations. And then depending on the domain, if you're a financial company, like there are many in New York, you may have certain other regulations that you have to fulfill. So these could be around, I'm reporting this data to a certain auditor, and I want to be able to prove that there's no Manipulation that's happening outside of what I already know during the process, or this particular system is regulated and should not have any sensitive data outside of these bounds, right? So understanding what your data is in that system and, uh, what are the bounds there so you can alert, um, our other style of compliance and governance requirements that are pretty top of mind for users.

AI assessment note: “governance matters, uh, in, in those two ways, right? The first one is around productivity”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Yeah. And you have a, you have a multi-cloud offering now, right? Uh, what is that?

A So we're the only player where a person, a company that allows you to run, well, one, we ran on all three clouds, but before you would run, say, app A on one cloud, and app B on another, and app C on another. Today, you, with our new announcement, you can run app A across multiple cloud providers. Now, why is that important? Well, if you're a customer in Australia, or you're a customer in Canada, Amazon, or say, Azure may only have one region. And if you want geo, if you want diversity and you only want, you only care about that market because you're a local Australian or Canadian company, the only way you get geo diversity is by going to another cloud provider because, and we've seen a lot of high profile outages, you know, and so they worry about like, am I going to get the diversity? The second thing is, you know, this, this market, the database market has a lot of baggage with Oracle's past behavior. So no one wants to get locked into any one vendor. So especially, Say someone like AWS who's got very broad ambitions as a company. And so one day you might find AWS competing with you and your core business. And so given that, um, companies do like the fact that they can leverage MongoDB to run on any cloud provider and, and leverage the best of breed services because the cloud providers are competing against each other. So Google may have some capabilities that are better tha…

AI assessment note: “you can run app A across multiple cloud providers”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q controversy around Elastic. It seems to have, like, gotten itself tripped up in this whole conversation. So with this whole complexity of, like, open source and cloud and the relationship, like, what, what do you think of, um, I guess open source in 2021 for a young company. If you're a young startup, uh, is that like the way to go? Is that, uh, or is there an evolution there?

A Well, I think you first have to ask yourself, why are you trying to be open source? What is your, what's the strategy behind it? You just, you shouldn't just be open source for the sake of being open source. You should just say there's a clear strategy. For example, Snowflake is not an open source company. Datadog is not an open source company, right? So you can thrive as an infrastructure company without being open source. You have to ask yourself, what is it? Why is open source an important element of your strategy? And be very, very clear on that. For us, it's all about, you know, a freemium strategy to get people to use, um, MongoDB and to try it because the market's too big for us to try and market to the whole mark, you know, the whole market ourselves. We want to leverage the virality of open source to get broad adoption. So I think that's one. Second thing is you have to recognize that if you, if, if you're open source, the best way to monetize open source is open source as a service. If that's the case, then how are you going to make sure you have a moat around your, your product? And that the cloud providers don't come in and, and, you know, take your free version and compete with you. So you have to think through your licensing around like how you plan to license your product and make it simple for customers so that it's not confusing to them or to the market about, …

AI assessment note: “You shouldn't just be open source for the sake of being open source.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And, um, maybe to, to, to put that in context further, just some of the, some of the names of, uh, you know, famous data warehouses, and a lot of folks are going to know this on, on, on, on, on the Zoom, but, uh, what are some examples of the main data warehouses?

A Yeah, the most popular data warehouses today would be, uh, Snowflake, who everyone's talking about right now because they're about to IPO, um, Redshift, Uh, which is a data warehouse that you can buy from Amazon Web Services. Redshift was incredibly important because, um, it was, it came out in 2013. It was in the AWS console. It wasn't the first really good fast data warehouse, but it was the first one that was cheap. Uh, and so a lot of people bought Redshift who previously would not have been able to buy one of the enterprise data warehouses that existed before that. And then Google BigQuery. Is another important data warehouse, uh, today that a lot of companies use.

AI assessment note: “The most popular data warehouses today would be, uh, Snowflake”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Um, so, so you mentioned one key project is essentially search, making everything searchable, so everything that's outside of Slack, searchable from Slack?

A Yeah, so I would say it starts off with, uh, just looking at the messages that are within Slack today. Uh, I don't know how many of you actually search regularly. Uh, it's, it's a start, I would say. As Stuart might say publicly, it's, Kind of a piece of shit, but it works well enough. Uh, he's pretty self, uh, deprecating about the product. Uh, right now it really works if you know exactly the message that you're looking for. So you remember maybe that someone said, all right, clarify, uh, you know, speaking event, and there's one message that you're looking for, and maybe you use some Boolean operators to find it. Uh, obviously that is not how the next set of people who are going to use Slack are used to searching. Uh, I think their expectations are kind of set in the consumer sphere, right? It's like they're used to being able to use Google, and they can type in a bunch of words, and optionalization will work, and synonymization will work, and it'll do query expansion, and all these things that just help you actually find what you're looking for. Uh, we're thinking about how to do that with all the information that's within Slack. Uh, starting with the messages that people actually type, uh, and figuring out how to actually rank those, and rank those so that you and I probably see very different results within an organization. Uh, and then moving on to people. So, thinking a…

AI assessment note: “and then eventually, yeah, to all the third party integrations that we have.”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.