Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Um, can you talk a little bit about just, um, like the intuition for how fine tuning or domain specific embeddings improves performance?
A Yeah. Fine tuning and domain specific embedding models are what we are very good at at Voyage. So just to have some context here. So, uh, what we do is that we start with a, uh, um, general purpose based embedding model, which is also what we trained from scratch. Uh, and from there we, um, uh, first, Fine tune or continue pre-tune whatever you call it, uh, on, um, some domain specific data. So for example, we, uh, fine tune on two trillions of code snippets, tokens, and then we get a code embedding model and we do the, uh, uh, fine tuning on one trillion legal tokens. And that's, uh, how we got the legal embedding model and this domain specific embedding models. I didn't use any of the preparatory data so that everyone can use them, But they really excel in one particular domain and the performance in other domains are not, uh, changed much. And the reason why we do this is because the number of parameters in the embedding model is a limited. So, um, because, um, you only have like a, you have a latency budget, uh, something like maybe 1:02, sometimes like a 200 milliseconds, you know, some people even want 50 milliseconds. Um, and then, um, basically it's, uh, it's impossible to use more than ten billion parameters. For embedding models. And we have limit parameters. Any customization is very important because the customization means that you use the limit number of parameter…
AI assessment note: “customization means that you use the limit number of parameters on the right tasks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So you guys actually started with this open source, um, model Bark. Can you talk about like what the idea was at the very beginning and how you ended up in music generation?
A We did our, we were doing all text at Kensho and we did our first audio project, um, after we were acquired by S&P Global, which was learning to transcribe earnings calls. So I'm sure both of you have read an earnings call transcript, uh, exceedingly likely it was done by S&P Global. Um, it used to be done completely manually. It was very painful and we could lend a lot of speed and scale by bringing automation to that. And we fell in love with doing audio AI. Like we happened to be musicians, but it kind of took this very honestly non-sexy project of earnings call transcription to show us how much we loved it. We also realized that certainly compared to images and text, audio was really, really far behind, and this was in 2020. And I think that's maybe even more true now, if you just look at everything that's happened in images and text in the last couple of years. Like I said, we never had a master plan. We, we made Bark and, um, as an open source project. And, um, even before we released Bark, we knew we wouldn't be focusing on speech. I think if I'm honest, a lot of people told us, go build a speech company. It is more straightforward. You'll build a Great B to B product and people will love it. And we couldn't help ourselves. We just love music too much. And so we decided to build a music company.
AI assessment note: “We just love music too much. And so we decided to build a music company.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Is it, um, like, a known space of, like, domains and algorithms one can learn, or is it really, like, learning to solve problems algorithmically really well, and it could be completely new problems or domains in every competition?
A Yeah, there's definitely, um, there definitely are standard algorithms that exist, so, like, you know, shortest path or something, or, you know, binary search trees or things like that, and so it helps a lot to learn the fundamentals, but the whole idea is that every problem in a contest is Um, it's totally unique. You know, it's, it's a new problem that's never appeared before. And, um, um, you know, the, the, the beauty of the contest itself is in the creative problem solving that you're doing to figure out the right algorithm. Right. And so, you know, while the fundamentals are very helpful, a lot of it is figuring out how to use each of these pieces and, you know, reduce the problem to a shortest path problem or how you, you know, modify, you know, certain algorithms to make them work for, for different use cases.
AI assessment note: “there definitely are standard algorithms that exist... but every problem in a contest is totally unique”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can I actually, um, just describe the two public demos so far and like why they're important?
A Yeah, we've been doing, um, kind of like two divergent set of demonstrations to the world. The first is we, we do plan as a business to start launching into more kind of industrial solutions, like more like the corporate labor market, uh, you know, manufacturing, supply chain, logistics, those type of areas. We think that'll allow us to, to build the AI data engine quicker because we're shipping robots faster and we'll, it'll help us build manufacturing volumes quicker, which will help cost. Those are like the reasons why we're doing it. There's another market that we're extremely excited about, which is in the home. And the home is a really messy place. It's very unstructured. Everything's different. There's like a higher variance of failures where we're like, you know, if we drop like the number one dad cup at home or the number one mom cup, like no, not great. We drop a bin in like a warehouse, like who cares? You know what I mean? It's a little bit different scenario. Also safety is impacted. There's pricing compression as we move into the consumer world. There's just a bunch of stuff that happens. Um, so, so we've done a few demos so far. First is we've done like bin moving, very traditional industrial solutions roles where we're taking bins from palace into conveyor systems. We're doing that fully autonomous end to end on our robots now, all bipedal. Um, and the second is…
AI assessment note: “First is we've done like bin moving... and the second is we're doing”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q the last time we saw each other in person was just how quickly, like, the AI, um, ecosystem and research field is evolving and what it means to manage an open source project through that. Can you talk a little bit about what you decide to keep stable and change when you both have, like, big ecosystem of users now and, like, very rapidly changing environment of applications and technology?
A That's been a fun exercise. So I mean, if we go back to the original version of link chain, what it was when it came out was essentially three kind of like high level implementations. Two were based on research papers, and then one was based on Nat Friedman's like Nat bot type of agent web crawler thing. And so there was some high level kind of like abstractions. And then there was a few like integrations. So we had integrations with, I think like open AI, coherent hugging face to start or something like that. And those two layers have, kind of, like, remained. So we have, you know, 700 different integrations. We have a bunch of, kind of, like, higher level chains and agents for, for doing particular things. I think the thing that we've put a lot of emphasis in, um, to your point around, kind of, like, what's remained constant and what's, uh, and what's changed is, like, a lower level, kind of, like, abstraction and runtime for, for joining these things together. One of the things that we pretty quickly saw was that as people wanted to Improve the performance, go from prototype to production. They wanted to customize a lot of these bits, and so we've invested a lot in a lower level kind of like chaining protocol, so lane chain expression language, and then in, in a different protocol lane graph, which is one, something we're really excited about, and that's more aimed at, ah, b…
AI assessment note: “what's remained constant and what's, uh, and what's changed is, like, a lower level”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q a combination of that and tree search and just, like, trying to be efficient with, like, your sampling at every step has shown, like, a lot of really interesting, uh, effective applications recently, and I think the, like, cognition is one example of, like, a surprisingly amazing agent has, has come out, like, Where else do you think agent, agentic applications will begin to work or that you've already seen?
A I think on the customer support side, that's a pretty obvious use case. I think Sierra, um, you know, has emerged there and is doing, is doing quite well there. Um, I think, yeah, the cognition demo was very impressive. I think they did a lot of things right. I think they really nailed a really interesting UX. Um, and that was maybe one of the things that, that I was most excited about. Um, and then obviously it seems to work very well. And so I don't know exactly what they're doing under the hood. Um, Uh, but, but those type like coding, coding problems in general, we see a lot of people working on. I think there's a really nice feedback loop that you can get by just like executing the code and seeing if it works. Um, and you know, as well as the fact that the people building it are developers and so they can, they can, uh, test it. Um, coding, customer support. There, there's some interesting stuff around like recommend, like recommendation, um, chatbots almost. Um, so I draw a distinction between that and customer support or with customer support, you're maybe trying to explicitly kind of like resolve a ticket or something like that. And the, um, and the recommendation bit is a bit more focused on like a user's preferences and, and what they like. Um, and I think we've seen a few, uh, I think we've seen a few things emerge there. Um, but I'd say customer support and coding a…
AI assessment note: “I think on the customer support side, that's a pretty obvious use case.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q any given application, take your prompts and go from, you know, Um, anthropic to mistrawl to OpenAI to something else. Um, in, in reality, it feels like, you know, the way an application, uh, responds is probably going to be sensitive to the fact that these LLMs are actually going to predict differently. Like, what do you think about this? Can you, can you switch? Is that a real pattern?
A It's not as easy as it seems like it should be, and I think the main thing is that the prompts still need to be different, um, for each model. I do think Um, the prompts will probably start to converge in the sense that if you think the models are getting more and more intelligent than like, hopefully these small idiosyncratic sees don't matter as much. Um, and as more and more model providers start supporting the same things, um, then that will make it easier. And what I mean by that is, you know, so many prompts for open AI, which is, you know, the leading and most used one use function calling. Um, and, you know, up until some period ago, like no other models did. And so you just like, couldn't use this prompts at all. Um, but now like Mistral has function calling and, and, and Google has function calling. And so I think they're a little bit more transferable there.
AI assessment note: “It's not as easy as it seems like it should be”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, I want to get into that and sort of what we can expect in terms of computing innovation if we're not just jamming more transistors on chips or we're unable to do that. Um, every one of our listeners I think has heard of AMD, but can you give like a very brief overview of the major markets you serve there?
A Sure. So AMD is a, a story company. It's been around well over 50 years. And it, uh, it started out really being, you know, a second source company, really bringing, uh, you know, second source on key components and x-eighty-six microprocessors. But you fast forward to where we are, uh, today, uh, and it's a very, very broad portfolio. Uh, when, uh, Lisa and Sue, our CEO and I were brought, uh, into the company just over 10 years ago, uh, it was with, uh, a mandate to, uh, get Uh, AMD back into very, very strong competitiveness, and so, uh, we started with the CPU line, brought the CPU, uh, very, very competitive, and then really across the portfolio, and just in February of twenty-twenty-two, acquired Xilinx, so that expanded the portfolio further. So AMD creates the world's largest supercomputers. It's got a massive install base now in the cloud, so many of your cloud operations That you're running, are running on, uh, AMD EPYC, uh, x-a-t-six CPUs. Gaming, we're, we're, we're huge. We're underneath all the, uh, Xbox, all the PlayStation, as well as, uh, many, uh, gaming devices that, uh, that, that you buy when you buy your, your, uh, add-in boards. And then across, uh, embedded devices with all of that rich Xilinx portfolio, as well as embedded x-a-t-six. And we, we acquired Pensando, so it extends that, uh, portfolio Ah, right into a networking interconnect that we need as …
AI assessment note: “supercomputers... cloud... Gaming, we're, we're, we're huge. We're underneath all the, uh, Xbox”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q traditional, uh, CNN, RNN, and other types of, um, neural network architectures, but also in terms of this shift to transformers and diffusion models and everything else. Um, can you tell us a little bit more about what initially caught your attention In the AI landscape and then how AMD started to focus more and more on that over time and what, what sort of solutions you've come up with?
A You bet. Well, uh, we all know the AI journey, you know, has been going since, uh, really the, uh, the race began when, uh, the application space for AI opened up, uh, and GPUs were obviously, uh, pivotal there. When you look at the, uh, the, the key work that, uh, you know, Uh, Henson had done in terms of showing how GPUs could drastically improve the, uh, accuracy of image recognition, natural language processing. Uh, and so that, that, that's been known, uh, for some time. And so what we did at AMD as, uh, we, uh, right away, uh, saw the opportunity. Uh, the question was plotting our course, uh, to be that strong player in AI. So it was a very, Uh, thoughtful and deliberate strategy because AMD, we had to turn around the company. So if you look at where AMD was, uh, in, uh, you know, 2012, uh, you know, through, uh, you know, really 2017, uh, it was largely all, all of the revenue was based on PCs and then gaming. And so it was about making sure that the portfolio, the building blocks We're competitive. Those building blocks had to be leadership. They had to attract people to get on that AMD platform for high performance applications. And so first we actually had to rebuild the CPU roadmap. And that was a Zen microprocessors that, uh, that we released in, uh, in both, uh, PCs with a Ryzen line, as well as Epic, our X-AVI server line. So that started the revenue ramp for the …
AI assessment note: “what we did at AMD as, uh, we, uh, right away, uh, saw the opportunity.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You mentioned Canopy is, um, to help enable more people to build RAG products. Like, where do you, where do you see developers or your customers struggle to get embedding space to AI products generally successful? Or what were you, what were you trying to achieve with, with Canopy?
A Yeah, so, the vector databases and Pinecone specifically are very foundational model, are very foundational pieces of technology. We're, we're very deep in the stack, and to build a, You know, a proper full end-to-end solution, say like Notion Q&A, there's quite a lot that you have to build on top of it. You have to ingest documents and, and what's called chunk them. You have to figure out how to break them into like factoids and pieces of information. You have to embed everything with models. You have to ingest them into the vector database. You know, when you get a query, you have to figure out how to manipulate it and how to embed that. You have You have to search over it. You have to re-rank. You know, there's, there's a lot. There's a whole system you have to build around it, and a lot of people told us that this is actually quite complex, and they're right, right? We put out Canopy as really an example. It is an end-to-end kind of cookbook. If you just take this, it should work. You should probably, once it works, you should figure out how to make it better for your own application, right? Because, you know, medical Data is not JIRA tickets, you know, and JIRA tickets are not Slack messages, and you might be building a different product, but at least you have some end-to-end starting point that already does something and you can start improving on.
AI assessment note: “We put out Canopy as really an example. It is an end-to-end kind of cookbook.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q example, you're, you're an HR company and you don't want different people's salaries to leak across. An LLM because you're using it as like a chat bot to help you with context regarding your own personal data in an enterprise or things like that. Can you talk a bit more about how embeddings can provide personalization and in some cases potentially other features that may be attractive to, to enterprises?
A Yeah. So that, that's a very common and reasonable thing to be concerned about. Data leakage can happen in, in two main ways. A, if you use a service for your foundational model that, that frankly, uh, retrains their models with your data or records it, right? Or saves it in some way that is opaque to you, right? That is a huge problem, and I think a lot of people are, a lot of people are struggling with that. The second is, if you're building an application in-house, whatever it might be, and you fine tune your models on added data, that added data might end up popping where it shouldn't in answers to, you know, other people's questions or whatever. What people do with vector databases is actually incredibly simple, right? You don't fine tune your model on your own proprietary data, at which point you know for a fact it doesn't contain any proprietary data, because it's never seen any of it, ok? And then at retrieval time, or at, you know, whenever you, you apply the, uh, the chat or the agent, you retrieve the right information from the database, Give it as context to the model, but only do inference. You don't actually retrain, you don't save that interaction, at which point that data doesn't exist anywhere. It's like an ephemeral thing. And the added benefit to that is, by the way, that you can be GDPR compliant. You can actually delete data. So if, if, you know, so, you kn…
AI assessment note: “at retrieval time... you retrieve the right information from the database, Give it as context”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So, you have often said that, ah, Notion is less a productivity company than an application building company. How do you think about the initial use case, and like, how, what makes you believe people want to build more applications?
A I don't think people want to build more applications. What got me started in Notion, got us started in Notion, it's, um, last year at college, I read a paper by, ah, one of the computing pioneers, Douglas Engelbar, ah, He talked about his papers named Augmenting Human Intellect. So every day we use software today, very much like application. When you go into one application, do one thing. But for that generation of computing people in the sixties, seventies, eighties, computers are a lot more, software are a lot more malleable. You can actually tinker and modify, right? Small talk. You can go into it and change how the operating system works on the fly. Um, that really inspired me. It's like today, people's software is so rigid, can we create a new breed of software that people can modify, can change and customize, and bring back some original ethos of those early computing pioneers? That's why we started Notion. Um, the hard lesson for us is like, like you mentioned, most people don't want to create software. They don't wake up and say, hey, I want to create my perfect project management tool, my project, perfect knowledge base. They both ask for something, they just have to get that work done. Right? Um, so the, in some sense, our learning and pivot is instead of giving people those, um, Software building tool, we have to package the software building blocks together as ready…
AI assessment note: “I don't think people want to build more applications.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q things in ML in terms of traditional ML. You know, I think fraud detection and the fraud detection API that you all have is one example of that. But you were actually quite early in terms of adopting LLMs and sort of early generative AI models. Could you tell us a little bit more about how that came about, how the interest was sparked, and how adoption really took off?
A I mean, I think it's fair to say that Stripe isn't first and foremost an AI company. As you know, fintechs, including Stripe, have long used traditional ML in many contexts, including sort of fraud and risk. But first and foremost, we're building financial infrastructure for the internet. So Stripe got started by enabling first really digitally native startups to accept online payments. And then over time, millions of companies started relying on Stripe's financial infrastructure for A bunch of different needs, whether that's reducing fraud or managing money flows or unifying online or offline commerce, all the way to launching embedded financial offerings. Um, and so as not a kind of first and foremost AI company, uh, we probably like a lot of people listening to this podcast had kind of our, like, hey, what the heck are we gonna do moment, uh, a year or so ago when LLMs really broke through the zeitgeist. And we were looking at the technical breakthroughs and the product launches. All over the ecosystem, um, with awe, but also honestly a little bit of overwhelm. The sense of, well, there's very clearly a real opportunity here to better serve our users, but what is it exactly? And how do we get it off the ground quickly and safely? So it starts with a story of three engineers who hacked together in three weeks an internal beta for an LLM Explorer. And the basic idea of LLM Exp…
AI assessment note: “starts with a story of three engineers who hacked together in three weeks an internal beta”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. That makes a ton of sense. As discussed earlier, Stripe has this amazing vantage point into all sorts of different online businesses and how they're evolving over time. What are the differences between some of these AI centric sort of next gen companies that Stripe serves as customers versus what you've seen traditionally in the e-commerce or SaaS or other areas?
A It's a great question. I mean, we've worked hand in hand with the builders of a bunch of different technology waves to make sure they have the financial infrastructure they need. Some of the earliest waves were marketplaces, infra platforms, social media, um, think kind of the young DoorDash or Instacart or Postmates or Twilio. And, you know, those were up to become some of the, the largest companies today. We've, we've grown up with them. There was also, as you noted, kind of the SaaS wave and the current wave is AI. And in terms of The unique needs of AI startups, um, probably four notable differences versus the prior waves. You know, the first is just at a basic level, unlike a bunch of the past generations of software startups, we're seeing AI startups have substantial compute costs right out of the gate, and that that's putting a bunch of pressure to build monetization engines faster. Um, the second thing we're seeing is a lot of these startups are seeing global demand for their products straight out of the gate, right? They're making Digital art, or music, or all sorts of borderless things, and they want to get that across borders from day one. Third, I would say, is a lot of subscription businesses, and obviously we see subscription businesses in a bunch of different contexts, but especially sort of the AI startups that are consumer-facing, heavily skewed towards subscri…
AI assessment note: “probably four notable differences versus the prior waves. You know, the first is”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Lots of lots of good stuff. What's a code red look like?
A A code red is when Toby sends an email to the entire company saying this thing is the number one priority, and he means it. From an actual operational perspective, what that really means is if the team that's working on this code red asks you to help, please drop what you're doing and help. This is the number one priority. So that's how they work. But usually a code red is a symptom of Some other much more systemic problem that's gotten you to that point. At least in my case, checkout code red in 2020 was, hey, the checkout is failing in, I don't know, three, four, five different ways. And the checkout is a pretty important part of Shopify, obviously. So let's go fix that. So we, me and, you know, somewhere between, I think at its peak, it was probably two, 300 people working on various parts of the problem. So we all like, Scrambled for a year and like did what it took to, um, fix those issues. But then at the end of that year, Toby took a step back and said, okay, well, why did that happen? Like, how did we get to the point where those problems even happened in the first place? And then that led to some of the reorganization of the company around Less, like there used to be 12 to 15 of these kind of fairly small fractured business units, and now there is actually only like three or four, which is actually truer to what the product is, but of course, each of those units is big…
AI assessment note: “A code red is when Toby sends an email to the entire company”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q that you and Toby and the team has done is actually shipped a lot of that quickly when a lot of very large organizations are like, oh, that makes sense. The outputs are non-deterministic, um, by nature. We have a process for measuring and evaluating quality of our products. Generally that process does not apply. Like, How did you get to this is good enough and we can ship it?
A Well, one of the principles that we currently hold, this might change in the future if confidence intervals go high enough, but one of the principles that we have for all of the Shopify magic features today is that they're allowed to propose changes, but not commit changes, right? So they can like generate text, but they're not going to save it without you actually reading it and saving it. We might suggest a reply for the user that's, you know, someone writes in like, Hey, what's your shipping policy? We can like suggest the reply based on us knowing what's in store and like running that through an LLM, but you have to hit enter, right? And so one of the things that like human is in the loop essentially right now is one of the, um, one of the ways that we're kind of mitigating risk here. And obviously human in the loop is great because it gives you the feedback cycle. You actually get a three part signal. You get which suggestions are accepted clean, which suggestions were accepted, but then with minor edits and then which suggestions were outright rejected. And that's An amazing loop to be able to improve things through. So we're always getting better, but that sort of human in the loop human must click save thing is a, is a big part of the strategy.
AI assessment note: “human in the loop essentially right now is one of the ways that we're kind of mitigating risk”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Is there anything external to Shopify? Um, that you think is especially interesting right now in the AI world, be it startups or things that people are working on or projects or.
A This is literally every single person has probably said this, but I think the, like the rabbit thing at CES was pretty interesting. Um, I mean, I'm, I'm literally wearing the shirt right now. Like I'm a teenage engineering nerd. So like I was, I was hyped on the hardware, but I think one of the interesting problems with like LLM based Applications and agents in particular is like, what's the actual interface they're interacting with, right? Like in the case of Sidekick, right? There's a couple different places. Like, like what is Sidekick using? Right? Is Sidekick actually using the admin API under the hood, or is Sidekick actually reaching into the pixels of the web app and clicking around in the web app and doing stuff, right? Like, that's a pretty important question. Like, what is the interface that the agent is actually learning and interacting with? And, um, I thought the really interesting part of the rabbit presentation was that they decided to treat the actual GUI as the interface and to try and like have a model that became very good at interacting with essentially web apps. If it actually works, the strategic brilliance of that is they instantly have the world's largest app store because the world's largest app store is just the web, right?
AI assessment note: “I think the, like the rabbit thing at CES was pretty interesting.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q That makes sense. You were at places that are great places to do research. Why did you decide to start a commercial company?
A It's a really good question. Um, I mean, there are a lot of companies that are funded by prior PhDs that are kind of the classic journey of there's a technology that was built in a lab environment, and it got to enough at level of maturity that, oh, we should start to commercialize it in the real world. That was kind of not the journey of CoVariant. Like when we started CoVariant, there was not AI that was good enough to make robots do Useful things commercially. Uh, and so it was not a classic journey of technology developed in academia and then transition to a commercial landscape. The key insight that we had at that time when we left OpenAI in, to start CoVariant was the future of AI is going to be the future of foundation models. These models that are truly multi-task, learn from large amount of data, And as such be more generalizable. They can solve new tasks more easily, and are also more capable at every single one of the tasks because of the transfer that you get across tasks. We just had early conviction that there was the path to build AI, and that is also going to be true for the physical world, for robotics. But there's one big problem, which is you have no data set to build robotics foundation model. Like there's no data set that you can build This AI that understands the physical world and take actions in the physical world. Um, and so in order to build this found…
AI assessment note: “in order to build this foundation models for robotics, you really have to build a company”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Then I think the right way to start actually be to ground the conversation and kind of like the application landscape. Can you walk us through the sort of limitations of robotics in warehousing and manufacturing that are commonplace right now and how much intelligence these robots have?
A Robots are extremely common nowadays. Like, so what we typically work on are robotic arms. So think of these as six axes, seven axes, um, robotic arms that can do very flexible movements. They are super precise. They're super fast and super doable and very cheap. Lots of factories around the world have robots. Um, but the challenge is like, 99 plus percent of the robots that are deployed in the world, Are dumb robots. Like these robots are pre-programmed to do the same thing again and again, and they don't really have any kinds of intelligence that can adapt to new circumstances, communicate with people and change what they do on the flight. And so think of robotics that exist today are extremely rigid. And so really the problem that we are solving is we're not trying to make the existing Dumb robot use cases better, right? Like we're not trying to say, oh, instead of, ah, manually programming this robot, you could just have an AI that, that program that robot. We're not talking about that. Like we're really talking about like opening up a couple orders of magnitude more use cases where the robots actually need to be smart. Like they need to adapt what they do based on the scenario that is presented to them, right? So like the w a good way to visualize this is on one hand, like think about A robot, for example, in a Tesla factory that is handling a car body. Like, okay, this is…
AI assessment note: “99 plus percent of the robots that are deployed in the world, Are dumb robots”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q One, one question for you before we zoom out from the, from some of the technical stuff. Does the offering of small models, like the seven or eight by seven B size that are quite capable, I think surprised a lot of people, um, from, from Mistral, like do small models that show higher level reasoning change your point of view at all or how you guys approach this?
A Uh, we're very bullish on small models, so we, we've actually integrated Mixtrol into Kodi. You can use Mixtrol, uh, as one of the models in Kodi chat, uh, as of, uh, last week, and it's just amazing to see the progress on that side. Uh, I mean, there's a lot to like about small models. They're cheaper and faster, and if you can make them, uh, approach the quality of the larger models for your specific use case, then, you know, there's, it's a no-brainer, uh, to use them. Um, I think we also like them in the context of completions. The primary model that Cody uses for inline completions right now is, uh, StarCoder, uh, seven billion. Um, and with the benefit of context, uh, that actually matches, uh, the performance of, uh, you know, larger proprietary models. And we're just scratching the tip, uh, of what's possible there with context fetching right now. So I think, uh, we're very bullish on, uh, Pushing, pushing that boundary up even further. And again, with a smaller model, inference goes much faster. It's also much, much cheaper, which means we can provide a faster, cheaper, uh, product to our users. What's not to like there? Um, I think there is a question, uh, with the smaller models, uh, specifically in the context of, uh, RAG, uh, because I think there's been some research that shows that the kind of like in context learning ability of large language models is, is a lit…
AI assessment note: “we're very bullish on small models, so we, we've actually integrated Mixtrol into Kodi.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q models. You know, I had very early, uh, in hindsight now forays into AI, things like Einstein and other things. And I know that's evolved into, you know, there's Austin Copilot and I send GPT and other things like that as well. Um, how much of the model development that you folks do now is internal versus using external sort of model sources, be they open source or closed source?
A We're taking really an open architecture approach because we have, we serve such a diverse set of customers. Some of our customers are large enterprises. They have their own models or they want to fine tune their own. Um, others are all the way down to SMBs who don't want to have anything to do with model selection and just want us to, to figure everything out for them. And so we're kind of taking the best of what's out there and we're, we're offering customers choice. And then there's a set of customers who have kind of asked us to take it on, right? They want us to figure out based on the data and the feedback that we're getting and given cost performance and latency objectives, they want us to choose the right model for the right task. So it's really a combination of using Whether it's, um, Cogen from, from our research team, which is the, which powers Apex, um, Cogen GPT that we have in, in our developer GPT, where you also fine tuning versions of that for domain specific models in customer service and for sales and for specific industries like healthcare and financial services. Uh, whether it's those in-house models, um, or it's working with our customers to allow them to very easily Spin up and fine tune their own models using the data that they have within Salesforce Data Cloud, or it's offering the choice of external third party models, be it Anthropic and Cohere, which…
AI assessment note: “it's really a combination of using Whether it's, um, Cogen”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q example, you do have reasonably specialized systems or all neural networks via specialized systems for the visual cortex versus, you know, um, areas of higher thought, areas for empathy or other sort of aspects of everything from personality to processing. Do you think that the transformer architectures are the main thing that will just keep going and get us there? Or do you think we'll need other architectures over time?
A So I have two, I understand precisely what you're saying, and I have two answers to this question. The first is that in my opinion, the best way to think about the question of architecture is not in terms of a binary, is it enough, but how much Effort. How much, what will be the cost of using this particular architecture? Like at this point, I don't think anyone doubts that the transformer architecture can do amazing things, but maybe something else, maybe some modification could have some compute efficiency benefits. So it's better to think about it in terms of compute efficiency rather than in terms of, can it get there at all? I think at this point, the answer is obviously yes. To the question about, well, what about the human brain and with its brain regions? I actually think that the situation there is Subtle and deceptive for the following reasons. So what I believe you alluded to is the fact that the human brain has known regions. It has like, it has a speech perception region. It has a speech production region. It has an image region. It has a face region. It's like all these regions. And it looks like it's specialized. But you know what's interesting? Sometimes there are cases where very young children have severe cases of epilepsy. At a young age. And the only way they figured out how to treat such children is by removing half of their brain. Because it happened at su…
AI assessment note: “can it get there at all? I think at this point, the answer is obviously yes.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Well, one thing that's stuck through many chapters has been the inbound conference, a conference that over a 100,000 people go to. You can give me the stats, like people like Barack Obama speak at it. It's sort of a pilgrimage for people who work in marketing. Like, uh, how did this happen? Why did you do it?
A So in the early days of HubSpot, the way we described ourselves to the venture capitalists like yourself was, At the time, Salesforce just said SFA. Salesforce.com is the sales. That's what's about is the marketing. And the good thing about that is it's stuck on VC's heads. It's like, oh, that sounds like a good thing. Like, we would go to Dreamforce every year. And by the way, we'd sit at Dreamforce and we'd sit there being like, God, I hope they don't announce a marketing product this year. Literally, the two of us would sit there like, fuck, please don't announce a marketing product. Anyway, we would go and we said, that seems like a decent idea, but it's more user conference. Let's create like a community thing that people will come, even if they're not HubSpot customers. So we said, let's create this inbound thing. And we just gave it a go. We, and we had a room of 400 people in the Marriott and Kendall Square. We invited Seth Godin and David Meerman Scott and Brian Solis. These folks are still floating around where, who are early sort of acolytes. They may or may not call it inbound, but basically the same ideas were coming out of their mouths. And we sold out in like five minutes and it was really good. And so then we sort of got on the treadmill and just kind of kept doing it every year. I really enjoy it. Like we had my, I'll tell you, we had some amazing speakers like…
AI assessment note: “Let's create like a community thing that people will come, even if they're not”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q of my favorite HubSpot people. So JD Sherman, um, your COO for a while, at certain points managed, you know, large parts, entire organization at HubSpot, um, and was there for a long time. In terms of all the founders that listened to the podcast, many at some point think about hiring a COO. How did you guys decide to do this? What advice would you have for, for people?
A We hired. So after we closed our Sequoia round, we bought performable and we stopped like tracking expenses. We just started spending like drunken soldiers and we just got distracted. It was a really good round. We had Sequoia, Google and Salesforce in the round and We just kind of took our eye off the ball. We had never missed a quarter, and we missed that quarter on the revenue side, and missed the expense side by a country mile. We had a couple long-time board members, Larry and David Larry from Journal of Catalyst, and they were like, you know what? You could use some help. I was like, what do you mean today? We think you should hire a COO, and I pushed back. I was like, nah, I got this, I got this. Like, not really. And so, They're like, go see if you can find one. And so I interviewed a bunch of people and I met JD and JD was a long time IBM person, which is like, oof. And, and then he was at Akamai for a long time. He was a CFO at Akamai.
AI assessment note: “We think you should hire a COO, and I pushed back.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You have made a huge bet on AI as a company. Can you just talk about Ghostwriter and how it came about and the investment in this area?
A Yeah, throughout my career working on code that handles code, right? Whether it's at Codecademy for teaching programming, whether it's my own project, whether it's React and building the runtime around React Native, I always felt like our tools that were handling code, whether it's like compiling it, parsing it, minifying it, all that stuff. We're very kind of rigid and very laborious. Um, in a lot of ways you're building sort of a classical intelligence system, very algorithmic, um, and a lot of heuristics and, and all of that. And I always thought that you can probably apply machine learning to it. And I started reading around whether anyone had done it. There was this seminal paper in 2012 called on the naturalness of software. And basically a bunch of researchers try to apply NLP to code. And what they found is that actually code Can be modeled like any language. That's why they call it naturalness is because hey, code is kind of repetitive, like language. You can do things like Ingram, basic Ingram model can actually start to generate code that is compilable. That had a huge impact on me. And every year or two while starting Replit, I would go look at the state of the art on ML on code and nothing ever really worked all that well up until GPT-II. And you can like take GPT-II and fine tune it and like try to make it write code and was like kind of okay. But obviously GPT-II…
AI assessment note: “And that's when we started building what became Ghostwriter.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q code autocomplete with local context has become as widely adopted as it has. Like your tool and co-pilot are sort of a handful of tools that people like very much accept have changed workflow. Maybe if we just project out a little bit, where do you think that the like next big leaps in productivity in software development come from in terms of AI? What else is going to happen?
A It's an amazing tool, but I think it's still very primitive, and we haven't fully explored the full capabilities of even GPT, but also transformer models applied to code in general. There's a lot to do. I think there's a lot to build on the sort of the layer just around the models. So they're basically how to give models better context, how to give them tools, the ability for the model to actually Be able to go out and read a file, install a package, evaluate code, the ability for these models to be more agentic and be able to kind of write entire features, uh, by themselves. And so that that's all the work just around the model itself with the model in the center. And then there's a lot of work on the models themselves that haven't been, you know, fully explored. The way we train models today is we just give it a large corpus of, of, of code. But you can imagine different ways of training models, and we've played around with different techniques at Replit. There are ways to make the models, give models more intuition about how code execution happens as opposed to just static code, because you're training them on static code. They understand the structure and syntax of code, but they don't really understand the semantics of code. So you can imagine a way to train Models where you're not only feeding them code, but you're also feeding them the results of the evaluation of that c…
AI assessment note: “how to give models better context, how to give them tools”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can you talk about the decision to, um, train models versus use existing APIs? Like, was it just a latency thing? How did that happen?
A So when Microsoft did it, like, they took GPT-III, they distilled it down, and they hosted it on their own infrastructure. They did a bunch of, uh, caching, a lot of things like that. If you're someone who's just using the OpenAI API, you couldn't do any of that. And it was very expensive. It was very slow. Um, and even now, you know, there's no completion models. Now all the models are chat models. You know, they're not going to be releasing any completion model going forward. And so if you want to build a completion based product, it's actually fairly difficult to do it using commercial APIs. And we felt like we can, we can make a model that's both cheap and fast and good enough for that autocomplete use case. And we found that the three billion parameter size is kind of the nice sweet spot where you get enough raw IQ from it, because like one billion felt kind of too dumb in a lot of ways. And it was going to be cheap enough and fast enough to host and be able to do a fast inference from. Um, and it was a time after Yeah, the chinchilla paper first came out and LLAMA had been announced and the idea of just training them longer had just been in the air and, and we're like, okay, what if we, we apply all these open source tokens on a smaller model? And then we applied some more tokens from replets data that gave us a 50% improvement over the model that we open source. So we op…
AI assessment note: “we felt like we can, we can make a model that's both cheap and fast”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you feel like, uh, you have customers that are already working for or planning for this or thinking about how to handle it, especially if they're more content oriented companies?
A Yeah. So on the, uh, bot mitigation and abuse prevention thing, every single customer that's deployed AI at scale, At any scale, a product that actually works has already faced this challenge. And of course we're, we're continuing sort of, in some cases you're playing cat and mouse. In some cases you're just advising the customer on how to implement better protections and better tools and finding that balance of, you know, how do I actually deliver a good experience for everybody while also protecting my business? On the SEO side, I think mostly I'm just hearing a lot of questions from people, right? Like, Is, is Google still king? Are the rules of SEO still the ones that, that apply to me? So I think those are the main ones. But again, my perception is there's a lot more people entering the crawling game and, and doing this retrieval process. Whereas before it felt like you had to delegate all of that to like Bing or, or the Google search API. And I think creating protocols to negotiate content and to make it more accessible and more distributable, It really depends on your business model to a great extent, right? For us, I would love if every single AI gets the most recent Next.js APIs to be correct, which is not the case right now. Uh, if you ask ChatGPT how to solve a problem with Next.js, it tells you the solution for 2020. And I would love for that to be the solution for …
AI assessment note: “every single customer that's deployed AI at scale... has already faced this challenge”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You've said that, um, we're on the cusp of the biggest positive transformation that education has ever seen. How important do you think AI is relative to the broader set of access that you have through things through YouTube and online coursework and all the rest of it? Is this a complete game changer? Is it an add-on? Like, what's the relative degree of importance of this shift right now?
A I think in the very short term, it is going to be a meaningful add on. I think if you go three to five years in the future, it will be a game changer. And the reason I say that is going back to Benjamin Bloom, but I think this even predates Ben, the gold standard was always to have a personal tutor. You go back to, if you were a prince in most of, if you were Alexander the Great, 2300 years ago, you had Aristotle as your personal tutor and Aristotle would speed up, slow down, motivate you when you're feeling down, like do all of these things with you. And that was always the gold standard. Two, 300 years ago, utopian idea of mass public education, but in order to do that economically, we had to make compromises, one of which is you don't get a personal tutor. You don't get a one-on-one teacher. We're going to batch you into groups of 30. We're going to move you at a set time or pace. We're going to apply some lectures and standards and homework, et cetera. On the test, some of you are going to get a hundred percent. Some of you are going to get 80. Some of you are going to flunk it. Too bad. The, the batch needs to move forward. And somehow we expect those of you who didn't know 20 or 30 or 40% of the material on the simpler stuff to understand the more advanced stuff. And what, what, what happens, those gaps accumulate and kids start falling off and then you eventually put kid…
AI assessment note: “in the very short term, it is going to be a meaningful add on”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Have you ever read the, the Neil Stevenson book, The Diamond Age?
A Not only have I read it, I used to give it away to people. Every employee at Khan Academy, uh, used to get the diamond age. And, you know, you're referring to it because, you know, this, it takes place in this like Neo Victorian, not too far off future in like China. And this member of this like Neo nobility gets this AI tablet app for his granddaughter to educate her, the young lady's illustrated primer. And it gets bootlegged and it gets in the hands of 200,000 orphan girls who live in barges, and then they essentially just take over. And so, uh, I've always used young lady, young ladies illustrated primer to our team at Khan Academy. It's like, this is what we're hoping to build in the long run. And I never thought we were going to fully build it. I thought we were going to be able to approximate it, but already conmigo can do some things that are maybe even beyond what the young ladies illustrated primer did, where you can talk to Don Quixote. In any language that you want, uh, you can get into debates with it, et cetera, et cetera. And I think in the next three to five years, almost everything that Neil Stevenson imagined, I think he wrote the book in 1994, um, I think is actually going to be a reality, which I didn't think was going to happen in my lifetime.
AI assessment note: “Not only have I read it, I used to give it away to people.”