Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You mentioned that initially GitHub Copilot ran on Codex, which was the early, um, coding model from, from OpenAI, uh, in, I think at the end of two of last year, 2024, you, uh, introduced GitHub models, uh, which, uh, feels like a library of different underlying models you can use. Walk us through that. How does that work? And what can you do as a developer?
A In 24, we did two things and GitHub models was, um, early August and that gives, um, developers access to catalog of models. Integrated into the GitHub platform. And in fact, last month, uh, in May, 25, uh, we brought, uh, these models into the repository. And so you can, you know, integrate these models into your repository to build AI scenarios into your own applications. And, um, that's really cool because you don't have to go to another place, you know, and sign up for a new account. And then you have your model stuff here and your code on the other side. Um, GitHub ultimately was always About, uh, millions of developers collaborating on, on a project together. And so for lots of years, we bring what we call the primitives, um, that developers need, um, into the repository, you know, issues, wikis, pull requests, actions, and now models. And then late, um, I think it was late October, uh, we announced multi-model choice for Copilot, moving from just having one model provider, OpenAI, to having multiple model providers, Uh, for Copilot Chat and for now Copilot Agent Mode. And so we added back then Anthropic Claude, 3.5 Sonnet, and now it's Anthropic Claude Sonnet four and Opus four. And we added Google Gemini back then 1.5. Now we are 2.5 pro. And what really, you know, it is about choice. Um, we fundamentally at GitHub believe that we need to offer developers choice, right?…
AI assessment note: “integrate these models into your repository to build AI scenarios into your own applications”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And the, the agent works with prompts, right? It's vibe coding where you describe what it is that you want it to do and it does it. Is that, is that right?
A It's prompts when you use it within the IDE. Um, although even there you can start with, um, brainstorming Uh, a cycle first. Uh, and so that, for example, that works really great, uh, with, with Claude, uh, Sonnet or Claude Opus for, as you can first ask it, you know, how would I build this and what are the, what's the system design for, for this feature, for example, and then have it first write a markdown file with bullets doing the engineering together with the model. And then you take the first task of that and feed that into the agent mode to write the code. For the coding agent, because it sits, you know, on GitHub platform, the starting point is an issue, and the issue can be, you know, the description that comes from a product manager or from a user of your open source project, but it's also all the comments in the issue, you know, attached images, um, uh, file references or, uh, web pages, and of course the coding agent and agent mode both can use MCP servers and tools, and so you can then You know, further connect into additional context. Uh, but yeah, fundamentally it's a prompt. It's just that the prompt when it's, it's no longer just one input field, you know, with the three lines of code, it can be, you know, a long description, uh, a specification, uh, from a product manager, just as they would fight for, for the human developer. And again, the product manager t…
AI assessment note: “but yeah, fundamentally it's a prompt.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And for the enterprise, do you allow people to fine tune their model based on their data? How does that part work?
A We do not today. Um, we, um, have pursued that path in the past. Um, so I think 20, 23, 24, a number of companies in our space saw fine tuning as the next, the next opportunity. The challenge from my perspective on fine tuning is that, uh, A, as there's another model every other day, uh, you're, you're ultimately always going to be behind with your fine tune model. And we are putting you then in a position where you're having the choice between the model that you fine tuned Uh, that is based on an older version of, you know, of the base model, or you can pick the new model, but that isn't fine-tuned yet. Um, B, I think most, uh, customers' code bases, um, especially if you look into, you know, the individual project, you know, most companies that are at a certain scale have not just a single programming language and a unified stack. That's the dream every CIO or CTO is talking about. Like, I want a standardized stack for all my developers, and Then you look into a 30 person startup, and of course, even they don't have that because as soon as they go from web development to iPhone development to Android development, they already have three stacks. And so then if you look at the individual repository or set of repositories, that code base isn't actually big enough to have a meaningfully fine tuned model. And lastly, I think this is where fine tuning kind of like got left behind b…
AI assessment note: “We do not today.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Let's double click on that for a few minutes. The launch of that agent mode, uh, was, uh, a major announcement, major new step in the history of, of GitHub that just happened at that build 2025, a few weeks ago. What does it do? Do and what tasks in particular would you suggest people should direct it to?
A So the coding agent, the way it works is that you can just give it a task, coding task, um, documentation, generating test cases, or simple things like find all the bugs in my code base. The agent then, you know, goes off in the cloud and spins up a virtual machine and checks out the repository, installs other tools, and Solves that task for you. And the magic here is that in the meantime, you can keep working on your part of the code base on a different task, a different issue on your local machine. And so effectively the coding agent is like a new member of your team that can take on certain tasks. And you can obviously, you know, not only assign one task to one coding agent, you can assign 10 tasks to 10 versions of, of that coding agent, and they can all run in parallel. And when they're done with their work, they submit a pull request exactly like one of your human team members would do. And then they alert you and say, Hey, this pull request is ready for review. And then you go in and you review the code and you can comment on it and the co-pilot will pick up those comments and, and keep iterating. And so if you don't like the code or, you know, I tested it yesterday with one of my hobby projects and I realized the README is, Uh, completely outdated because, you know, I didn't spend time on writing a README for myself and just told the coding agent, look at the code base,…
AI assessment note: “give it a task, coding task, um, documentation, generating test cases”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Alright, so that's a model layer. Um, on top of that, uh, do you have an API layer? I mean, it seems that part of the business is to provide the models to companies like Oracle or Notion. Um, is that the model themselves? Is it a separate product?
A No, we provide the whole Platform. So serving of the models, optimization on different hardware, like AMD, um, we're up and running on Cerebris, Grok, of course, Nvidia. Um, so all of that we provide and we can deploy anywhere. And that's been quite unique. I think one thing that's different about agents and AI compared to other SaaS is that usually you're trying to do something that a human is doing in the organization. And To do that, you need the con, you need the same context that human has. And so that means you need to give very broad visibility, like our north platform to be fully useful needs to see all of your internal communications, all of your emails, all of your documents, your customer records, your et cetera, et cetera, et cetera. And that is a huge security risk, very unique one compared to like CRM software or HR software, that type of thing. Um, and so our security posturing, the fact that we don't say, Hey, send your data over to us, like hit our API, trust us. We're socked to the fact that we don't say that. And instead we say, we're going to ship our models directly on your hardware, whether it's in your VPC on a cloud or, you know, for regulated industries in your data center, that's been a huge unlock. People get comfortable plugging in much more. And so they can actually do more with the product.
AI assessment note: “No, we provide the whole Platform. So serving of the models, optimization on different hardware”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So speaking of that, so You know, what's, uh, what's, what's happened over the last couple of years? Like any, you know, metric, uh, including vanity metrics that you can share, fundraising history, number of customer, number of documents, whatever it is that you want to share to give a sense to people for the reality of the company as of today.
A I can't forget that you're a VC, so I'm not allowed to share any, any metrics over here. Uh, but there, there's some public ones. Uh, and, and I think things that I'm, I'm really proud of are really lasting or first and foremost, our team. So we built a team of a hundred amazing, uh, incredibly smart, uh, folks all five days a week, sometimes six in office in New York city. And we are now just starting to become a multinational corporation. Uh, and so we were opening SF and we've already opened a London office, uh, which is incredibly exciting with, with, uh, goals to end the year at 300 to 400 employees. Uh, lots of exciting growth. A lot of it here, uh, in Silicon alley. And unfortunately, Silicon Valley there, we have to, we have to move out there a little bit. But I think on, on like AI metrics, or a little bit more about the product and how it's used, one of our favorite things to track is the amount of unstructured data, the amount of pages that are processed by the platform. And a really interesting thing is, hey, last year, Hebbia and probably all of the other major consumer model providers processed around a hundred million pages. Probably around, whatever, hundreds of years, maybe thousands of years of, of, of reading. This year, we're already on track to process around four to five billion pages. Uh, so, you know, somewhere around 50,000 years of reading, uh, for, fo…
AI assessment note: “This year, we're already on track to process around four to five billion pages.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So we alluded to, uh, Influx three dot O, which, uh, effectively you all just announced, uh, congratulations. So it's from the outside. It looks like a major effort at rewriting a bunch of things. Why did you do that in the first place, and, uh, what is it that you did?
A We started writing, depending on how you look at it, we started writing in FluxDB, I don't remember when the first repository opened, let's call it three, three and a half years ago, so it's been a really big project. And there were certain assumptions after being in the market for, you know, at that point for seven years and, you know, probably at that point having 12, 1300 customers, we just learned a lot about what customers were doing and what issues they were having. And so we set out to solve, let's call it four or five problems. So the first problem we set out to solve is, is, um, is the issue of cardinality. So when you get a database that's describing a rich set of data around a time series This thing, you can run into cardinality problems that can choke off, um, can choke off the performance of the data.
AI assessment note: “we just learned a lot about what customers were doing and what issues they were having.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And maybe to drive it home, you know, again, very basic way. What's a couple of examples of, uh, I'm an analyst in a company. And I'm in charge of BI. What kind of queries do I get and need to answer through my BI tools?
A It's been around for so long. I think the general thing that a BI person gets asked is one of two things. Can you please update this dashboard? By which they mean, I want another data source included, or I want a different type of visualization or something like that. Uh, or the second request they get is, can you extract this data set to me and just email it to me? That's what BI people, I think, typically do. I think what they want to do is figure out for the business what is actually going on in the data. You know, what's happening to our marketing performance? Is it getting more or less efficient? What's happening to our inventory? But they end up servicing a bunch of very rudimentary, relentless requests from users who are one step removed from the data that they want. So I think both sides are very unhappy with this process. But unfortunately, I think that's probably a reasonably good description of what we see people doing in classic BI jobs today.
AI assessment note: “Can you please update this dashboard?... or... can you extract this data set”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Fantastic. So to get into the meat of the product and the more slightly technical discussion, a big part of the idea is that you live on top of Snowflake, all modern data warehouses. Do you want to talk from a product and technical perspective about how that works, where the data lives, where the computer lives? How do you do it?
A Yep. So one of the things that changed, obviously, with the Databricks and the Snowflakes and the Big Queries and so forth is that the data volumes got huge. One of the things that I knew before, I was working in infrastructure, and one of the things I knew before joining Sigma was that the price per terabyte of storing data in AWS for five years in a row was -65%. So you don't have to, you know, sort of graduate with an economics tree to understand that if you have declining prices year over year, you're going to have increasing volumes. That's exactly what was happening. Snowflake and the like came around and just figured out how to make, or better said, I like to think of it as organize that data so you could use it. So the first thing that Sigma had to do architecturally, uh, is abandon everything BI products had done in the past around managing for performance. All products a priori were built with caching layers. We do, we do not do any caching. So everything you do in Sigma is live query on the warehouse. Our customers have up to trillions of records. Uh, and our overhead on that transaction is less than one second. So you log in through Sigma, you build your assets in a Sigma controller layer, but we push down every bit of the query. So the data never leaves the warehouse. This has a number of very salient, uh, beneficial impacts. Number one is super high performance, o…
AI assessment note: “we push down every bit of the query. So the data never leaves the warehouse.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And clearly you're all in on cloud. There is a little bit of a theme around cloud repatriation, maybe with AI, you know, people wanting the models to be very close to the data. Is that something A, you see, B, you worry about?
A I think that the models are going to live next to the data in these warehouses, and we're already seeing that. Um, Anthropic, uh, doing, you know, if you know Dario, doing that deal with Snowflake is a clear attempt to make sure that they're as close as they can be to the enterprise data, that they're getting great performance, that they're adhering to security models. Uh, I think that's a trend that we are going to see. I also think we're going to see people, uh, building, uh, smaller models in these warehouses as well. Either tailored to whatever business they're in or trying to manage cost. So I think that's inevitable. Uh, there's some trade-offs, of course, to not working on the, on the, on the premises side, but I think the, I always think about things as a train track in this particular case, like the trains left the station, it's gaining momentum. There's no point in building backwards.
AI assessment note: “I think that the models are going to live next to the data in these warehouses”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q you think about, um, building your own models, which you've done, uh, to some extent with Arctic or to a very large extent with Arctic, uh, versus partnering? So you mentioned, uh, OpenAI and Cloud, and you just announced, uh, in the last few weeks, uh, major either partnership or expansion of partnerships with both. Microsoft to deploy OpenAI models, and Enthropic to deploy Claude. How does that all work?
A We got out of the foundation model business early last year. It just looked like an impossible challenge. At the end of the day, we are a smallish public company compared to the likes of Google and AWS and Microsoft, or even OpenAI in terms of how much money we are able to put for things like model training. We shifted our folks to focus more on things like post-training, where we felt we had, we still had leverage. We also have a very good inference research team that specializes in making inference super cheap and super fast for, for our needs. I think that's the right place to be. That also drove our partnerships. By the way, these are deep, meaningful partnerships, meaning the anthropic models run within our deployment, so I can With confidence, look at our customer and tell them that their data is not leaving the Snowflake deployment, and it's similarly with Microsoft and Azure and OpenAI. And so these are meaningful integrations that we have done working with these companies and the cloud providers. I feel like that's a much better use of kind of our resources, which we have to husband a little bit than continuing to Try to invest in foundation models. We got priced out. It's fine.
AI assessment note: “We got out of the foundation model business early last year.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And just to complete this kind of, um, architecture, uh, product tour, uh, Cortex, uh, the analyst, like what, what are, what are those in, in just a few words?
A We said we wanted AI To be a core part of Snowflake. Roughly the way we did that in practice was we said we will host a model garden inside every Snowflake deployment. Snowflake essentially runs in what we call deployments, which you can think as a point of presence in every major data center that AWS and Azure and GCP have. These are instances of Snowflake that are running everywhere, but it's a single cloud. It's completely connected. Data can move seamlessly from one place to the other. And, um, we run a model garden inside each of these deployments, and we offer a set of products on top of it. That's the name Cortex. Cortex AI is an umbrella of products. In practice, what this means is anyone that can write SQL or Python in Snowflake can use language models just as part of their data processing. If you want to do sentiment detection on customer feedback you have in Snowflake, it's as easy as calling a single function. Similarly, if you want to create a chatbot on unstructured data, you create a Cortex search index, um, and then you use Streamlet, which we talked about, to create a user interface and an application that you can deploy. And Cortex Analyst is the, is a structured data solution. It's the idea of being able to ask a question about a structured data set that sits in Snowflake, and Snowflake Intelligence is the uber structure on top of all of this That helps with …
AI assessment note: “Cortex Analyst is the, is a structured data solution.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q ties with your own work? So, you know, things like, uh, deep learning guided program synthesis, what does that mean? This, this, uh, what does program synthesis mean? And this, um, you know, how, how, how deep learning helps? Um, and then, you know, other, other, some of the other type of approaches were test time training, and then combining program synthesis with transductive models. What does that all mean?
A Yeah, no, I think ArtPrize last year did a tremendous job at highlighting, you know, what are the current best approaches to create models that actually have fluid intelligence, and they're all test time adaptation methods. Um, so in particular, one category of method that, uh, that, you know, really got big via ArcPrize is this time fine-tuning, or this time training, um, where you're using an LLM that's doing, you know, transduction, meaning that it's looking at the task and trying to directly predict the answer. Uh, that's, you know, you can think of it as opposed to doing induction, which is that you look at the task, And you try to predict a program that will, that will turn the task into the answer. So in one case, you just directly predict the answer. In the, in the other case, you try to predict the process or program that gives you the answer. Um, and so you're, you're looking at these, uh, these transductive models. And at this time, you're gonna, um, generate, uh, uh, Input-output pairs from, from the current archetype that you're trying to solve, uh, specifically for fine-tuning, I'm going to fine-tuning your model, uh, to map the one, one input to the outputs, and then, and then you're going to run that model, uh, on the test input, and I'm going to see what it gives you. Um, so that's one thing. Um, another category of approaches, uh, that, uh, that got big with A…
AI assessment note: “program synthesis is this idea that you have some, The language that you're working”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And, and, and to this point, so the, the one of the versions of the, Is it the private version? Correct me if I'm wrong. You have to publish exactly what you do. You have to say limit on the amount of compute. Is that fair?
A Yeah. I mean, the, the, the base idea is that there's one track for self-contained approaches with no internet access that are very efficient. So they have very limited compute budget and authors must open source them at the end of the competition. So this is really designed to incentivize, uh, open sharing, uh, and, and, you know, get, get as many ideas as possible. With a big focus on efficiency. We believe, like, efficiency is not just a good feature to have in your system. It's actually at the heart of, of, of intelligence. Um, and the, the other track is to provide continuous benchmarking of frontier models to be able to track, okay, like, if you look at the, at the best currently available commercial frontier models, things like, you know, right now, for instance, O-one Pro, soon the future is going to be O-three and so on, Gemini three, whatever, uh, How much fluid intelligence do these models actually have? Um, mostly independently from, from the efficiency consideration. We still do have, uh, um, this notion that we want to monitor efficiency, so we're going to be reporting, uh, results on a two D plot. Uh, so we're not just looking at the score as a scalar. We're looking at the score associated with, uh, uh, the cost per task, and it was required to achieve this score. And of course, uh, a model that gets you the same score, But at much lower cost per task is a smarte…
AI assessment note: “they have very limited compute budget and authors must open source them at the end”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And the GPUs themselves, they sit on some cloud or like, how does that, um, how does that work? Like, do you run on like all sorts of NVIDIAs and MDs and like, are you agnostic that way? Or like, how does, how does that work?
A Yeah, similar to PyTorch IDEA, we want to provide, we want to build on top of the best hardware possible, and the different hardware skill has different strengths, so we build on top of NVIDIA GPUs, um, from the, um, many, many generations to the most recent one, B-TU-Hundred. Uh, we also build on top of AMD's hardware, uh, and we optimize on top of ROCCM. Is there, um, CUDA equivalent layer? Um, so yeah, so we have been deployed to many, more than 15 regions globally, uh, currently on top of five different clouds. We, we plan to expand to possibly 10 different clouds and, uh, a lot more regions, uh, globally. Um, and in terms of deployment, we have our hosted, managed, fully managed single tenant API. Um, we can also deploy into your VPC, uh, and into your cloud, uh, for privacy, um, concerns. Or we can connect VPC to VPC through private link or VPC peering. So all these possible way of deployment we have enabled for various different enterprises and startups.
AI assessment note: “currently on top of five different clouds”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And then any thoughts on, uh, the future of, uh, the industry at large? Uh, like the next 12 to 18 months, is that the year of agent that everybody's talking about? Is that the year of adoption? What, what's your, uh, what's your prediction? What are your predictions?
A Uh, so there's no doubt. 2025 is the year of agents. Literally, no matter where you go, people are building all kind of agents. So, uh, I think there will be thousands of agents being built, focus on solving specific problems. Um, at the same time, I also think 2025 is the year of open models. And it's clear that Deep Seek has, has created a big dent there, but the meaning is not about Deep Seek itself by itself. It's the precedent it has set in the world about setting a huge, like a very solid baseline for open models, for open model providers, um, that it's just Need to be better before anyone can open source new models. Um, it also set a presence for all model providers to be better. Um, I'm very bullish on open models because I have seen the power of open source that PyTorch has been able to leverage. Um, DeepSeq, um, for example, just within one month of releasing their new models, There are, despite DeepSeq model, extremely hard to tune and optimize, extremely hard. There are 500, more than 500 variants published on Hugging Face, optimizing for, um, for local device, optimizing for cloud infrastructure. People tune and DeepSeq model, people distill all in DeepSeq model into various different kind of models. Companies like Perplexity also tune DeepSeq model Uh, for.
AI assessment note: “2025 is the year of agents. Literally, no matter where you go”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q into, uh, RAG, uh, and the architecture in some detail, and maybe into, I guess, the one dot O version, and then we'll spend time talking about two dot O and Gentic RAG and what contextual, uh, does, uh, but, uh, maybe as, uh, as an introduction slash deep dive, uh, how does that work? So this, uh, retriever, there's a generator, all those good things. What's the core architecture?
A Yeah, the core architecture is very simple. You have a language model, so that's the G, And then you want to give that context, and the way you do that is by augmenting it, the A, uh, using retrieval, the R. So that's R-A-G-RAG. Um, and, and so how you do the retrieval, that has been changing constantly over time. Um, so in the initial paper, we used a vector database or a face. So the, the, the words, uh, vector database didn't exist at the time. Um, but so face was the first vector database. Um, and Uh, I think people over time have started figuring out that that has all kinds of limitations, right? So how a vector database works is you, you just have embeddings. So you encode pieces of information or chunks of documents, you encode them as a vector, and then you do basic dot product similarity search. Um, but, uh, that, that has issues where you're just looking for chunks that are similar to the question, but you don't necessarily want to find chunks that are similar to the question. You want to find chunks that are relevant to answering the question, right? So you, you need to do different things with the representations, with the embeddings. So a lot of modern RAG deployments are very different from the original, uh, ideas in the paper, right? So you still have a vector database, but the way you encode things is, is very different. Um, then, uh, you usually also have a spa…
AI assessment note: “Yeah, the core architecture is very simple. You have a language model, so that's the G”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Replit or, or REPL. What, what is REPL in, in REPLit?
A You know, uh, for example, a DOS or like a UNIX command line is a form of a REPL. So, uh, it comes from, um, Lisp actually in the 19 fifties at MIT, um, where, uh, you can create, uh, the simplest programming environment with like one command, which is read, eval, print, loop. So read reads the command from the command line. Eval, evaluates it, like a lisp string, uh, print prints out the results, and then goes back to the start. So this is like the simplest IDE in the world, right? Um, and so when I was building, so it, you know, the, the way sort of Repl.it was born, I, I was going to school in, uh, in Jordan, uh, studying computer science and, um, I, I didn't have a laptop. So every time I would want to do some homework, I would have to reinstall the development environment. And that was like the, you know, it was so annoying, right? You have to download gigabytes of software and packages and, and there's always something that goes wrong. Um, and so I was like, you know, I'm, I'm using everything in the browser. At the time, Google was putting docs and Gmail in the browser.
AI assessment note: “one command, which is read, eval, print, loop.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And like, how do you, how do you, yeah, how do you position?
A Yeah, I think it necessitates being, uh, even crisper on the differentiation between those things, and so for us, there, there are a few things that we point to with customers. Number one, we're one of the only hybrid players, so Databricks and Snowflake are cloud only, so if you happen to have data on-prem, uh, we're pretty much your only bet, and you know, it just so happens that that turns out to be most of the Fortune 500, uh, almost the entirety of the financial Social services sector in particular, and so we do a lot of business in those industries as a result. Um, the second thing that helps differentiate us is the openness of the platform. So we're an open engine querying open formats. Uh, and while there's been widespread embrace, I would say, especially last year in 2024, around open formats and Iceberg in particular really winning that format war, um, that's new. That's new for this industry. We've been doing this Forever, though, and that's, you know, the first queries run on Iceberg were, were Trino or Presto queries. So, um, that pairing in the open source community of Trino and Iceberg, uh, has been kind of a reference architecture for years at this point, and that gives us an advantage because we can help manage your Iceberg deployment holistically. Everything from streaming, you know, ingest, uh, uh, loading that data into Iceberg tables, Maintaining it, doing …
AI assessment note: “for us, there, there are a few things that we point to with customers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q it's been fascinating actually to see the, the Dell like stock price. Yes. Everybody's talked about Nvidia. Yeah. You know, for all the right reasons, but, um, it's like amazing how Dell has like the quintessential on-prem player, uh, has, uh, seen their fortunes accelerate as well last year. Um, okay, great. So that's, um, Starburst Enterprise and, um, Starburst Galaxy, the managed version had a How does that work?
A Yeah, so that's a classic, you know, SAS, um, uh, product that's hosted and managed by us. It is connecting to your storage, uh, so it's your own S three buckets, your own, you know, RDS, your own MySQL database, uh, but the compute and the control plane is managed by us, and so we're able to offer a very seamless, easy to use, turnkey, push button type of approach, while still giving you all the performance and functionality Uh, that you need and that you're looking for, and that product has actually evolved very, very quickly for us. Um, we've been able to get a lot of new interesting features and functionality there, and one of the cases, one of the use cases where we've seen a lot of, uh, adoption and interest is actually customers who are building their own data applications and using, uh, Galaxy as the embedded engine, uh, where, you know, there's, uh, Data analytics portion of the SaaS app that they provide to their customers and the analytics are essentially powered by our engine. And, you know, we were digging into like, you know, why are they choosing us, uh, for this? And I think it goes down to, you know, if you're gonna be part of, uh, someone else's margin, essentially their cogs, right? Uh, you need good TCO. I think what we're seeing is, um, data application developers are choosing iceberg for storage because that gives them a lot of flexibility. They want Uh, s…
AI assessment note: “that's a classic, you know, SAS, um, uh, product that's hosted and managed by us.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q maybe we'll, we'll, we'll put in the video, uh, where, you know, you have the classic, uh, sources on the left, and then, uh, you know, Starburst in the middle and magic on the, on the, on the right side. Uh, so walk us through that. So in terms of, of sources, how do the data go into, uh, or how does Starburst access the data? What kind of data?
A Sure. Yeah. So at the heart of, uh, Starburst is this notion of connectors. And, uh, we think of everything as a connector. So even if you're just accessing S three and you're going to be querying iceberg tables, That's, that's technically our S three connector to our data lake connector, uh, to access that. Uh, and so every connector is basically just connecting to the underlying catalog of the system that you're connecting to. And then as soon as you've connected, which is like a, a one-time, you know, setup thing, uh, now you can run queries and you can do that at the command line, like just start to write SQL queries, joining tables across different systems, or, uh, you can use a BI tool like Tableau or ThoughtSpot or others. Um, or, and then this goes to the sort of data application side, we're seeing customers, you know, build more programmatic ways to interact with, uh, the data, um, that we have access to. Um, one of the features that we've built that's really nice, especially for, uh, internal purposes is, uh, something called data products, which is basically allowing you to stitch together a view of your data across these different data sources, and that's where you really start to get some, Interesting optionality because you can decide to materialize that view or not materialize that view, and, and there are trade-offs to both. You know, if you're querying the data…
AI assessment note: “at the heart of, uh, Starburst is this notion of connectors”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But like, to which extent, in your opinion, does that help you take control over that?
A Yeah, I think this is a really interesting question. Um, I don't think it does allow them to take over the project, to be perfectly honest. Uh, I think that the market is actually resolutely determined to ensure that it continues to be independent, which is important actually for Iceberg, and that's what made Iceberg popular in the first place over Delta, you know, which was Databricks' you know, own format. The market wants an independent Standard. That's what they want, independent of any vendor. And so, fortunately, there's enough groundswell of people like ourselves, like Snowflake, uh, like some of the others you mentioned, where it is, and it's also, by the way, an Apache Software Foundation governed project. So, you have the Apache Software Foundation also ensuring independent governance, which is important and makes it truly, I think, independent. I mean, Ryan works for Databricks now, and he will for some period of time,
AI assessment note: “I don't think it does allow them to take over the project”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q you worked at Facebook, um, and you also wrote in one of your posts about the reality, uh, of data engineering at companies outside of Silicon Valley. Can you, can you maybe compare and contrast, uh, if I'm a data engineer, do I do different things if I'm at Facebook or Meta or Google, uh, or Uber versus, uh, what I would do in a company outside of Silicon Valley?
A Yeah, I don't think necessarily the goal changes. I think there's, there are aspects of the fact that there are different tools. One, um, when you work at, Companies like Facebook, uh, and other big tech companies, their infrastructure is very mature, right? Like they've, they've generally done at least most of the ones that I've either seen or worked at and case with Facebook, they've done a good job of like creating a centralized set of tooling that everyone's using and everyone's on. Um, which makes your job very easy. Uh, when I worked at, and when I've worked at companies outside of Facebook, you've spent a lot more time kind of just setting up and configuring whatever tools you need to manage. Uh, at Facebook, you essentially have your internal airflow that a different team runs, and you can just push files to it, uh, and it picks up the data pipelines, you know, automatically. Whereas in most other companies, you might be the person that has to even manage, you know, your airflow instance, which is managing all your data pipelines. Uh, you go into a different company, they're gonna have Six different, uh, like data storage or data platforms. One team's gonna be using Databricks. One team's gonna be using, uh, Snowflake. One team's still somehow using a SQL server instance. Um, you've got all these mergers and acquisitions happening, which also cause kind of further layer…
AI assessment note: “there's less support if you're a data engineer outside of Silicon Valley”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q folks, so, um, hopefully we can cover some key concepts and definitions and all those things, as well as going into more technical stuff so that there's a little bit for, for everyone. Um, But maybe before we get into this, a little bit of your background, your story. So you're Ben, but you're also known as Seattle Data Guy. So do you want to go into all of this?
A Yeah, sure. No, great. Thank you for, for the, the light intro. So yeah, like you said, uh, Seattle Data Guy. Now, now I live in Denver, so it's currently just snowed, I think the other day. So it's white outside. Uh, that's new for me from Seattle. Uh, I've been kind of in the data engineering space for nearly a decade at this point. Started at a hospital doing a combination of like data science, data analytics, data warehousing work. Um, I think like most people around 20 12, 20 15, I got bit by the data science bug and I was like, I want to do data science. It's, you know, the sexiest job, uh, of the 21st century. So I got into that, but eventually found data engineering and really enjoyed it. Uh, from there I worked for a healthcare analytics startup, um, that did everything from like fraud detection, some stuff with like opioid monitoring and a few other Uh, things around like population health, uh, again, around data engineering and building kind of a massive, uh, data warehouse essentially of, of pretty much like, I think it was like over half the US, uh, in terms of like who, who existed in that data warehouse. And then eventually did that at Facebook as well, data engineering. Um, uh, all throughout that time, I've been consulting as well and, and putting out a ton of content. So I think I started putting out content somewhere around in Um, so I've been writing for a l…
AI assessment note: “I've been kind of in the data engineering space for nearly a decade”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q where there's like every year like a more, you know, even more gajillion, you know, new tools kind of thing. Why is it so Complicated. What, why, from your perspective, why are they, why do you need to put together those, like, chains of, like, tools, and sometimes it works, sometimes it doesn't. You mentioned, like, broken pipelines, that seems to happen all the time. Like, why are we here?
A That's something I've definitely been thinking about more recently, because when I first started in the data world, the common approach for, for a lot of companies I worked with would be, like, okay, we'll set up, um, Some cron jobs to run some SQL scripts essentially, or some Python scripts that call SQL scripts, or PowerShell scripts that call SQL scripts. And we'll just time it all correctly so that they, you know, run in a specific order or set up dependencies somehow. And then what a lot of people would end up doing is just essentially building Airflow, um, but internally. And so we knew that the problem, like we've run into this problem where it's like these things don't talk to each other and that causes issues. So then when Airflow I think came out, I think that's why it was popular for, for a lot of people. It was like, okay, cool. Now we can like write scripts, uh, and not have them bump into each other or have them, you know, make it very easy to backfill and all these things that, you know, we find painful, uh, as data engineers. And then somehow with the MDS, the modern data stack kind of push in We kind of ended up here again where it's like, okay, we have one tool that does extract. Okay, we have one tool that from extract does You know, transform. Okay, from there, you know, you kind of can keep parsing it down. We have another tool that does data quality. Um, p…
AI assessment note: “part of this works because I think it grew the pie”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What are some of the use cases? If I'm a, you know, an enterprise potential customer today, what do I use Rider for? And then we'll get into the, the, how it works behind the scenes.
A So if you are a major investment bank, let's say you are the top investment bank in America, um, you are using us for everything from earnings calls, summaries, to deep dives on specific companies and sectors. Um, you're using us for search over merger proxies and M&A docs. If you are a Salesforce, they're a major customer. Um, you're using us for dozens of different kind of custom applications plugged right into Slack to review content for compliance, to automatically, uh, rewrite for SEO, to produce marketing and email. Um, if you are a, uh, insurance company, CSAA, uh, uses us to build Uh, knowledge assistance for agents who are answering the phone. So our, our focus area is, and the verticals that we focus on are financial services, healthcare, and retail CPG. And, you know, the, the, the types of use cases are, um, everything from your mission critical, like reviewing, uh, insurance policies and doing claim adjudication on them, um, to helping Salespeople sell more. Now, when we are in call centers and support, it's not, you know, chatbots for your website or helping exchange mediums into smalls. Like, there's lots of companies doing that. It is really in supporting, ah, a high-end kind of knowledge task. So, if you are, um, a CPG company and somebody calls into the call center, To ask whether, you know, there's phthalates in the, uh, Avino product, right? You are, as some…
AI assessment note: “you are using us for everything from earnings calls, summaries, to deep dives”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q As any, um, you know, great entrepreneur, you've been very focused, and you seem to have said no to, like, self-hosting kind of scenarios. Where, where do you, first of all, is it correct? And second, um, uh, just walk us through the thinking.
A Yeah. And, and it's just something we're like still debating to a large extent, right? And by the way, I think of it like self-hosting and not just, you know, kind of a binary thing, it's kind of a spectrum, right? There's like, there's like on-prem on one side, there's like sort of, you know, fully multi-tenant on the other side, and there's a bunch of stuff in the middle. There's, you know, VPC peering, private link, cloud prem, you know, BYOC, ring your own cloud. So, so, so, so it's actually kind of more of a spectrum. Uh, but, but we've been sort of cloud maximalists and we've been all in on like, let's just build like a multi-tenant service where we just run everyone's code in our cloud environment. And, and to be clear, like a lot of customers is not, you know, a lot of potential customers is not, you know, comfortable with that model. I tend to think about that as like the cloud itself, you know, I'm old enough to like, I, I, I, I, you know, started my career pre-cloud and I remember the first time I heard about AWS, you know, in 2007, 2008 or something when they launched, like my first reaction was like, How could anyone ever run their code in someone else's computer? That's nuts. Right. And then like just a few years later, I was like doing it myself and I was like, this is awesome. Uh, and, and so, and then like, I think there's like sort of a similar thinking, you k…
AI assessment note: “we've been sort of cloud maximalists and we've been all in on”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What are you building next in terms of a roadmap you can share?
A I mean, like, I, I hope to build this for the next 20 years. So like, I think there's so much stuff we want to build. Like right now, like we're, we're focusing a lot on just like taking the existing compute platform and, and improving it in various ways. Like one thing, for instance, like I'm, I'm very focused on like, How can we get into more like real time, like low latency use cases, which is like kind of an annoying, you know, annoyingly complex technical problem. Like we, we've sort of, you know, relied on like a centralized control plane running in one single region, US East one. Uh, but we sort of recognized in order to get to sort of latencies below, you know, a 152 hundred milliseconds, we probably need to split that up and run like a decentralized control plane. So that's going to take a long time.
AI assessment note: “run like a decentralized control plane”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Which was in, like, 2006, right? Yeah, 2004 to 20 years.
A 2006. Everybody's brain just sort of, like, broke, and they're like, wow, in order to build systems that can handle the data sizes that we're seeing, you kind of, you have to just Dramatically change how you're building them. You have to, you have to run on lots of machines, lots of cheap, inexpensive machines versus, you know, these giant hyper expensive machines. And to be fair, that was a problem at the time. Um, but you know, nowadays, like, you know, I've got a Mac M two laptop. It's two years old. Um, it's like probably an order of magnitude to two orders of magnitude faster than the server machines were back when, you know, MapReduce came out and people started building Building these systems, you know, let alone like nowadays the server systems, you know, have hundreds of cores and, you know, can have terabytes of RAM. Um, and, uh, and so really kind of like if you were going to build something now, like why would you bother with all the complexity of scale out? Because the thing about the way we designed systems, like I was one of the people that helped start Google BigQuery, uh, and I worked on, you know, single store for a couple of years. So I, um, you know, I, Have been in, you know, with my elbows deep in, in, you know, building these, these, these complicated systems is that, like, there's just this huge tax that you pay to, um, to have to build a distributed sys…
AI assessment note: “2006. Everybody's brain just sort of, like, broke”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q When was that moment? Because having a, you know, a highly successful library is one thing, but then turning yourself into a platform to host a bunch of different libraries and models and all the things, when, when did that happen?
A It started quite, uh, quite early. Even in, kind of, like, the first days of, of Thomas releasing, kind of, like, the, uh, first port of, uh, BERT, I think, uh, contributors, open source contributors started to solve bugs, right, and kind of, like, improve some small things. Um, and then progressively as we added more models, more contributors started to, to add, uh, more, more models. Uh, so progressively the community, Contributed more and more, and we felt this, uh, this movement where the more we contributed to the community, uh, the more open source we, we did, uh, the more the community was giving us back. Um, so it, it validated us into, into this approach, uh, to today, as I mentioned, where we have, uh, five million AI builders using, using our platform, who collaboratively shared, uh, Uh, one million public models. Half of them have been downloaded in the past 30 days. Uh, so most of them are actually very useful and active for the community. They also contributed, I think it's, uh, more than 200,000 data sets to the platform. Um, so it's open data sets that anyone can go and use to fine tune, to customize their models for specific language, specific domains, Specific, ah, use cases. And collectively they built over 300,000 spaces which are the apps, ah, on, on the Hugging Face platform.
AI assessment note: “It started quite, uh, quite early. Even in, kind of, like, the first days”