Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q come on the show, um, is around, you know, how they think about running their company management style. Uh, given that you come from a different area and type of organization, we were just curious, you know, what was the most unexpected thing about running a government agency? And also how would you typify your management style or how has that evolved over time as you run, um, the FTC?
A So one of the biggest adjustments for me was, um, you know, in prior jobs, I really had just the privilege of going really, really deep on like one thing or a couple of things and, and feeling like I totally had it mastered inside out new kind of every last detail. Um, and in this job, you have to just have more capacity for breadth over depth and you're, you know, relying so much more on the expertise of, of other people. And we have a really fantastic team. And so Making that transition from, you know, greater, uh, capacity for focusing on, on breadth rather than depth and just being able to switch contexts, uh, very abruptly. I mean, you know, the types of things that go into running an agency range from, you know, really substantive decisions about what investigations you're doing, what cases you're bringing, how to think about what theory you're including versus not to, you know, much more administrative, Stuff relating to the budget and how are you allocating, you know, your resources and, uh, what is the future of workplace flexibilities look like at your agency? And so, uh, you know, really adapting to that full range. You know, for me personally, it's been important to figure out, um, what is my comparative advantage in this job and how do I Protect my time to make sure I'm spending as much of that time on that area of comparative advantage and build a team around me. …
AI assessment note: “one of the biggest adjustments for me was... more capacity for breadth over depth”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I want to talk about risks too, but what is the aspiration for how this changes performance management?
A For me, the, the sort of motivating factor here, honestly, it's the, it's the bad manager. If you're an employee and you are working in the bowels of the organization on hard problems, your manager, a little lazy, doesn't sort of recognize the quality of your contributions, shows up at that calibration meeting with a better vibe on somebody else and they get the promotion. Talent signal walks into that environment and slams your work product down on the table and says like, what about this? I can give you a concrete example of an underrepresented profile at Rippling. When we were building this product, she was an engineer in India who was working on one of our toughest problems, and she was singled out as a high potential employee. And she was in fact pretty early in her tenure at the company. And we paid attention to that. We talked to the manager about it and it was sort of an eyebrow raising moment where she was kind of lifted from obscurity by the model that was like, I don't know what your vibe is on this person, but like, man, they seem to be contributing at a high level and here are concrete examples of how they've done so. So the lazy manager who doesn't like represent things the way they ought to is held accountable by their manager when they look at the total organization through this tool and it does a better job of representing the employee. And obviously I can talk…
AI assessment note: “the sort of motivating factor here, honestly, it's the, it's the bad manager.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can we go back to, um, you said goals and guardrails for a minute? Um, like, um, as you described, we're going away from, you know, complex rules engines as business software. That's like a pretty big mindset shift for your customers to ingest. How do you work with them on, um, I guess, evaluation of, like, how well Sierra agents work and get people comfortable with that?
A Yeah. So a couple of things I'll describe technically and then talk a little bit more operationally as well. So technically we work a lot with our customers to actually formalize and define their processes, you know, and, um, sometimes our customers come in with really well defined processes. Sometimes they don't. We like to say an agent's made up of not only the factual knowledge, but the procedural knowledge, you know, with the process follows and it gets into the, the, the integrations with systems. And we spend a lot of time talking about where do you want guardrails? Where do you want Creativity. And where do you want agency? And then we do a lot of experimentation, you know, uh, in a proof of concept, have it live and actually through sort of, uh, this technology encountering the cold, hard reality of actual people, you know, did this actually meet the expectations you thought you had? And then with that, we've developed a lot of tools for customer experience teams. Uh, so we think that AI should not be the domain of Technology teams exclusively. Um, you know, the team that owns your customer experience at your company, maybe it's in the office of the chief digital officer, maybe it's a formal customer experience team. They should be the ones with their hands on the steering wheel of these experiences. So we built a lot of tools and platforms where those teams can audit a…
AI assessment note: “we do a lot of experimentation, you know, uh, in a proof of concept”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Usually when people talk about one-on-one learning, they talk about the adaptive aspect of it, where you're challenging person at the level that they're at. Do you think you can do that with AI today? Or is that something for the future? And it's more today, it's about reach and Multiple languages.
A I think the low-hanging fruit is things like, for example, different languages. Super low-hanging fruit. I think the current models are actually really good at translation, basically, and can target the material and translate it like at the spot. So I think a lot of things are low-hanging fruit. This adaptability to a person's background, I think, is like not at the low-hanging fruit, but I don't think it's like too high up or too much away. But that is something you definitely want because not everyone is coming in with the, with, um, with the same background. And also what's really helpful is like if you're familiar with some other Disciplines in the past, then it's really useful to make analogies to the things you know, and that's extremely powerful in education, so that's definitely a dimension you want to take advantage of, but I think that starts to get to the point where it's not obvious and needs some work. I think, like, the easy version of it is not too far, where you can imagine just prompting model. It's like, oh, hey, I know physics, or I know this, and you'd probably get something, but I guess what I'm talking about is something that actually works, not something that, like, you can demo and work sometimes, so I just mean, like, it actually really works in the way a person would.
AI assessment note: “This adaptability to a person's background, I think, is like not at the low-hanging fruit”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q sort of twofold. Um, on the one hand, you're doing a large scale, uh, custom model Um, specifically in part focus on code, and then you're also building sort of at the product suite that can really help, um, address, uh, coding and working, uh, on the, the software development side. How did you decide to start Magic, and why, why focus on that, um, versus other aspects of AI?
A It sort of came from a place of working backwards from AGL. If you, uh, your end goal is to have a system that can do everything, uh, you can reduce that to building a system that can build that system. And so that minimal system is a system that writes code and comes up with ideas and can validate those, uh, by writing code and running experiments, which is still, like, in the same order of complexity as the full thing, but at least we don't have to train Zora. Uh, and, you know, we don't have to think about ten billion other use cases that everyone building general domain products has to think about. We only have to think about code. So it's a lot simpler on all aspects except compute and slightly simpler on the aspect and slightly cheaper on the aspect of compute. I, uh, I think it's not a lot cheaper. I probably overestimated how much cheaper it would get, um, on the compute site, uh, at the beginning. Uh, but the other things are simpler, I think.
AI assessment note: “It sort of came from a place of working backwards from AGL.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q of the other open questions that people have wondered about is, um, does all of the value creation in the ecosystem go to, um, Um, your compute vendor, and eventually a big piece of it over to Jensen and NVIDIA, or to the model vendor, and I, I think the, the answer, at least to like right now, is clearly not, right? I think there's different, there's capture different levels.
A There's probably enough for everybody. Uh, today, most of it does go to NVIDIA, I think. That's, that's a lot of it, but I just think that's because it's where it is early in the cycle. Uh, you know, I think, um, uh, and they've built some incredible technology that's enabling some really cool stuff, so I think that that's, it's, it's, Um, it's fine. And at some point it's going to be the, the companies that find out how you actually go solve real problems and deliver real value to enterprises and to customers and other things like that. And that's going to be that, you know, I, I see a lot of, if I, if I take a step back and see who's implementing AI out there, It's a lot of enterprises that are doing proof of concepts and, and a lot, and sometimes they'll find one that really works well and it'll go to production. And I think if you can have a startup that can make that part easier, that says, look, this is a real value, right? It's not a chat bot on your website, but it's something that helps you go faster, make sales better, innovate more rapidly, um, you know, do something you were never able to do before, uh, improve manufacturing efficiency, whatever that is the startup is focused on. Or the company is focused on for that matter. It's, it's that, it's going to be an application level, right? It's most, most people don't, um, build a CRM from scratch. They go use a Salesf…
AI assessment note: “today, most of it does go to NVIDIA, I think.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q If we were to abstract out a level and, uh, you know, ask what you, what your vision is, or how are you thinking about the next three to five years of AWS more generally as a business? What are the key things or areas of focus for you?
A Well, this is one. I think, I mean, I'm, I'm just as excited about generative AI and, and AI broadly as, as you all are. Um, I do think that it's an enormous opportunity for us and for our customers to, and, and I think it actually, in many ways, it has a positive flywheel effect and is, can be a tailwind to some of that first stuff that we were talking about a little while ago about helping customers move to the cloud. You know, I think if we think about where can generative AI help, some of that can be like, how do you make that go faster, right? How do you, Take some of that more, you know, our original AWS thesis was we take care of the muck so you don't have to. There's still a lot of that that customers have to do today that I think generative AI can help with. And so over the next three to five years, there's a big investment for us in both building that tool set, building that whole platform that we're talking about so that customers don't feel like they have to go manage a bunch of these pieces. They don't, and they don't have to think about, you know, GPUs, or they don't have to think about how do they think about kind of tying these clusters together or whatever. All of that can be abstracted away. If you think about the start of what Bedrock is, if you go use Bedrock models today, You never interact with the GPU, right? You just, you send it tokens, you get tokens b…
AI assessment note: “over the next three to five years, there's a big investment for us”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And Jared, what, uh, for, for anybody who, uh, hasn't heard of Foundry yet, what is the product offering?
A Yeah. So Foundry, we're essentially a public cloud built specifically for AI. And what we've tried to do is really re-imagine All of the systems undergirding what we call the cloud, end to end, from first principles for AI workloads. And I'm sure I do this in a bit of a new way. I think the AI offerings From the existing major public clouds and kind of some new GPU clouds haven't really re-envisioned things. And by thinking about a lot of these things a bit anew, we've been able to improve the economics by, in my case, it's 12 to 20 X, um, over lower tech GPU clouds and existing public clouds. And, you know, we'll, partially based on some of these products that we'll talk about today, um, that we're releasing and a lot of new things that we're working on, um, we think we can push that quite a bit further as well. Um, and so, you know, our, our primary products are essentially infrastructure as a service, so our customers come to us for elastic and really economically viable access to state-of-the-art systems, um, and also a lot of tools to make leveraging those systems really seamless and easy, and we've invested quite a bit in things like reliability, security, elasticity, and just the core price performance.
AI assessment note: “we're essentially a public cloud built specifically for AI”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q one other thing that'd be great to cover, I know we only have a couple minutes left of your time, um, is, uh, the recent paper that you authored, um, which I thought was super interesting around compound AI system design and sort of, uh, related topics to that. So would you mind telling us a little bit about that paper and sort of what Uh, what you all showed?
A So I think it's kind of in this regime that we were just talking about where more and more often to kind of go beyond the capabilities on Frontier accessible to today's state-of-the-art models and kind of get GPT-V or GPT-VI early. Practitioners are starting to do these things oftentimes implicitly where they'll call the current state-of-the-art model many, many times. There are many scenarios where maybe you're willing to expend a bit of a higher budget. Maybe it's code or something, and if I said that I can give you a 10% better model, you know, for code, many developers might pay 10 X for that, access to that. Instead of 20 dollars a month, they might be very willing to pay 200 dollars a month, right, for obvious reasons. Um, and so there's a question of what do you do in that setting. And so people are, you know, if you're willing to call the model many times, you can compose those many calls into almost a network of network calls. Right? And, you know, I guess one of the questions is, how then should you compose these networks of networks? What principles should guide their architecture? We kind of know how to construct neural networks, but we haven't yet elucidated the principles for how to construct networks of networks, so to speak. Um, these compound data systems, so to speak, where you have many, many calls, maybe external components. And so one principle that we Star…
AI assessment note: “one principle that we Start to explore was... you can look at how verifiable the problem is”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q about are like, as you said, for organizations. And I remember talking to you guys when you were just getting started, you were like kind of discovering that there was all this value in businesses and how businesses spend instead of like your first company was more consumer oriented. Like, how'd you decide to like make that shift and like how, you know, learning about that audience, investing in it?
A I'd say, I'd say a lot of the like really ideation phase of, of ramp. We were, um, Talking to a lot of our friends, people in our community, and it just so happens that a lot of them were either starting early stage companies or joining early stage companies. Um, and a lot of the, the, the problems that we're facing at a larger scale were some of the ones that we're trying to solve for consumers first. Um, and the, the, the funny thing with, with, uh, businesses is the, the, the better they got, and the larger they got, they actually, the more wasteful they would get, and the less they would know about their Not only is it, like, a more interesting, uh, uh, and in some cases, like, bigger problem to solve, it just, like, scales with success in some ways, so, like, the better companies were even more interesting opportunities for us, so, uh we,, we went after that, and there was, I'd say, like, another realization we had early, um, around, uh, like, you look at the user experiences of, of different products out there, and the ones we use As consumers on a day-to-day, like, they obsess over every single interaction, every single flow, like, Instagram's amazing. Robinhood did that very, very well for trading stocks and investing, and then those same people who use these apps in their daily lives show up at work and are expected to use tools that were built in the eighties and are …
AI assessment note: “we saw an opportunity to really bring a lot of consumer thinking around like UI UX”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Where did the first, um, savings thesis come from? Like what, you know, if, if the idea is like spend less money and have less waste for consumers on like, there's a price change, like you go deal with this. That sounds amazing. Like what was the first hook for a ramp?
A First was just a basic business problem of, I think most startups, their biggest battle is people not giving a shit. Like you work on is so hard on the software and there's a million other pieces, you know, apps, tools, cards, How are you going to stand out? And we realized there was this gap in the market. Most of the credit card ecosystem was really designed around kind of ego and access and use my card and get lounge access and points and you can fly to exotic places. And it just was so different from what business owners and CFOs and people I know were trying to build enduring profitable businesses. So it felt very at odds. And so some of it was just, um, thought there was a large unmet need.
AI assessment note: “we realized there was this gap in the market. Most of the credit card ecosystem”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can you tell us a little bit more about how that's impacting Functions across the, the company today?
A A hundred percent. I mean, so a lot of people know outside in, sort of, said the ramp's one of the fastest growing FinTech of all time, one of the fastest growing software companies. Some of this is good positioning and timing in a good market. Some of this is like very, very early adoption of, of AI in, into augmenting the capabilities of, uh, our sales team, of our marketing teams, of our underwriting teams, all the, all the way throughout. And I'll, I'll, I'll zoom in on, on one. I think that there's a lot of startups now starting to think about Um, you know, AI sales and automations, and, you know, internally years ago, we had built a functionally outbound automation team, uh, and what we had noticed is, this is back years ago, we had very little resources, but there was one sales rep who was incredible. You know, he could go out and book far more meetings than anyone else, and we were trying to figure out, we're like, wow, if only we had two of him, this would be great, or three, and, and we want to understand, like, what is making this person so productive? Uh, and so one of the unusual things we did is we, uh, had a few engineers sit down with him and just track what he did during the day. And he'd follow kind of his calendar and attorney wake up in the morning and there'd be a new set of companies that raise funds or the people that he was in touch with who moved compan…
AI assessment note: “augmenting the capabilities of, uh, our sales team, of our marketing teams”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Robin Hood, and others were, you know, um, that ECHOI structure as well, but I, I, I find it striking that in Europe it seems like a more and more common pattern. So it's just, it's kind of an interesting aside. Could you tell us a little bit about how Brex is starting to think about AI and some of, some of the innovations and approaches that you're taking there?
A There's three big areas for us. One is, um, the obvious, which is product. So how we can improve the product and, and, and make the experience of essentially Expense management better, right? And, and within product, there's two areas that we spend a lot of time. Um, accounting is one, and the other one is essentially expense assistant and expense management, basically what EAs would typically do. Um, second big bucket besides product is, uh, go to market and operations. So things that are very ops intensive internally, prospecting, you know, KYC, underwriting, compliance, uh, lots of use cases there. Uh, and the third broad bucket is, like, developer productivity. So how we can help engineers be more successful. Uh, and there, I don't think we've done anything particularly remarkable in this third one because, uh, you know, we're using the same tools that folks have used, co-pilot, and things like that. Uh, we're experimenting with new tools, but I would say buckets one and two is where we spend the majority of time so far.
AI assessment note: “There's three big areas for us. One is, um, the obvious, which is product.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Maybe even one step back from that, like, at what point did you say, um, internally at Brex, like, I'm gonna get up to speed personally on this, or we're going to make this part of the product?
A Probably 18 months ago, I think right after ChatGPT launched, we, we started to just play with, um, you know, ChatGPT online and, and I had played with the GPT three, uh, through the APIs before ChatGPT was out. And, and obviously it was, it was impressive, but ChatGPT was that moment that everyone started to think, okay, what does that mean for my business now? The mental model that I don't think is particularly unique that we, we have is, you know, if we were to think about, you know, humans that are free, what would we do? Uh, it turns out there's a lot of work in expense management and accounting that people would automate, right? And the one that was really obvious to us early on is, you know, if we think about what is the best customer experience when it comes to expense management and corporate parts is essentially what executives have, which is there is no experience. You just swipe and it's done, right? There's an EA in the background that will figure out how to get a receipt for that, how to categorize his expense, how to get it approved, right? I prototyped something on the weekend with, you know, in, in Python with GPT-PT-Port five on, can we just get someone's calendar and use that context to generate a memo, categorize an expense, and potentially find a receipt, uh, on someone's email. Uh, and the results were like surprisingly good. Uh, and it took me, you know, …
AI assessment note: “Probably 18 months ago, I think right after ChatGPT launched”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q That's really impressive velocity, especially in an area where, um, you know, I, like, from other fintechs to sort of traditional financial services companies, a lot of them feel like they cannot use AI because it is probabilistic and there's, like, risk and reliability issues. Like, how do you, how do you handle this or get something into production relatively quickly?
A So for us, um, that is sort of the holy grail of AI and fintech, I would say, is, like, how do you build this Degree of conviction that what you're, what you're suggesting is correct. Right. And, and, and, and maybe, maybe it's interesting to talk about accounting, which is one that people are fired if the results are wrong. But basically the way we thought of it is, is, um, is, is twofold. One is, um, how do we, uh, expose ambiguity to the user? Right. So instead of saying, Hey, let me just try to predict something and put it in front of you and say, Hey, Um, you know, this is the generate, this is what we generated, you know, good luck. We thought it would be a much better customer experience if instead of having, for example, like, you know, chatbots are particularly bad at this, where, you know, a chatbot gives you an answer. There's no affordance. Uh, you can't understand other potential options that are generated. You can't understand context. So a lot of it is just like building ambiguity into the UIs and into the flows. And for example, like our expense assistant, when we're generating a memo, Um, if we have really high conviction in a suggestion, we can go and say, hey, this is what we strongly believe is the answer. Uh, and we automatically apply it for you. If we're not that confident, we show suggestions in the bottom of like a field. And if we're not confident at a…
AI assessment note: “A lot of it is just like building ambiguity into the UIs and into the flows.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q If you were going to look at sort of outside of the product to, as you said, go to market and the operations intensive parts of the business, do you have like a rank ordering of like where you think it's going to have the most impact?
A In general, scaling marketing is, is the biggest one. And the way, the way I frame it internally to, to folks is like, you know, if you look at, again, the same framework applies everywhere, right? Which is just like, what would we do with infant humans? And back in the day, if I think about marketing 10 years or 15 years ago, Um, you, and you, and you were at like Salesforce and you were trying to close, you know, Coca-Cola or a really large enterprise customer. You literally had an account based marketing team where you literally have a PMM, a product marketing manager working alongside with a sales rep to market to that account, right? They were literally creating a pitch deck. They're creating materials. They're creating, they're going to Coca-Cola and meeting the executives of like very specific pitch of like the value that Salesforce can provide and so and so on. And, and really the way I think about it is like, there is a world now where you can generate, uh, effectively account based marketing for any accounts because the cost of doing that is, is marginal, right? So, um, the way we think about this is like, how can we prospect across accounts that never got the level of personalization, uh, and that level of care, uh, with a lot of these tools and, and, and really in a really specific way, Provide value to these accounts in ways that we couldn't provide before. So, you…
AI assessment note: “In general, scaling marketing is, is the biggest one.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q very commodity, um, and not very different from, like, one overworked SDR sending a spam campaign themselves, right? But what you're talking about requires, like, much more insight into, like, What is a demand signal, a value signal, like, specific data that Brexit itself has to collect because it understands that signal and wants to meet it now? And so I, I think, like, that should be much more powerful.
A A hundred percent. There's this book, like, Crossing the Chasm, that is, like, a kind of classic go-to-market starter book, and it talks about this idea of, like, what is the definition of a market, right? And, and a very important concept is the idea of customers that can reference each other. So, like, it's not, you're not really in the same market if this customer can't Go talk to this other person about your product and hear something, right? You know, like, do they know someone that uses it? And, and, and effectively, like, I think now there is this ability of creating, like, almost an infinite number of markets where effectively, like, you know, I want to create a market of, like, construction companies in Missouri that, you know, uh, are high spenders on card or, you know, they have complicated accounting needs that may leverage breaks. And, you know, you, you can essentially, like, Outscale your ability of doing this with humans in a way that you probably wouldn't be able to do before. Uh, and, and to your point on the, a lot of the tools doing that, um, the writing the email is the easiest part. That's actually not that hard. The part that is the hardest is aggregating all the data and essentially building this, um, this database of your whole team. Like who is every account in the market that could buy Brex potentially. And what do we know about them? And how do we en…
AI assessment note: “The part that is the hardest is aggregating all the data and essentially building”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I mean, I'm thinking of things like the older, like MedPalm II models and things like that, where they outperformed human physicians in terms of output. Relative to the physician expert panels. And then obviously you could then do some post training with physician experts as the, as the key, but at some point the machine will be better than that. And so how do you keep scaling reward functions?
A Traditionally, right? You, you just get it's supervised learning, right? Reward function means good or bad. So we can scale that process as much as we, we have so far. Um, I mean, obviously many, many players are, are realizing the power of Human annotation. And in fact, deep learning comes thanks to amazing like Fei Fei Li and lab effort to, to label a data set of a million examples, right? So, so that way of scaling is one. Uh, but then I strongly believe that there might be a bootstrapping effect of the models that become better at judging their own outputs, right? And so maybe, and that really is probably maybe even the main hypothesis of reinforcement learning as I see it. I mean, I'm not Huge expert in RL, but If checking that something is correct is easier than creating the solution, then we're in business because the language models will be able to evaluate their own samples more accurately than to generate them. And then we have a sort of reinforcement learning loop because we can reinforce the ones that seem more promising and then the model gets better, right? So that using the model itself as a reward, um, which Incidentally uses language, which is already fuzzy, is one area that, I mean, I'm excited about. There's a leaderboard of reward models. Um, some of them are, I think the name that they use is maybe generative reward model. I think that area goes beyond this…
AI assessment note: “a bootstrapping effect of the models that become better at judging their own outputs”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q some, like, obviously work on real, um, workflow, structured recurring processes, as you said, um, you know, did you, did you figure anything out about how to give other people, like, you've been looking at this for a long time, any frameworks for how to give people intuition for what today's models can do, either like in Airtable, because you guys have to have the expertise, or in your customer?
A The short answer is we've been trying, uh, to do a number of things to kind of codify that and scale it beyond like one-on-one bespoke, you know, kind of, uh, interactions, right? Cause we can't be like Palantir and go really, really, you know, forward deployed for every customer. Um, so one is we actually now run this AI workshop program. It's a lot lighter touch than like the Palantir AI bootcamp. Um, but, you know, for instance, we just have one in LA. We have like, you know, probably 60 people from all kinds of companies, a lot of media companies, some like retail, big like retail companies. Uh, et cetera. And it's a full day kind of master class in first, you know, really just teaching people like, what are these transformer models? Why, why have they gotten so much better recently? I mean, looking at literally, uh, this slide of, you know, parameter count of these models from like five years ago to now from GP one to GP four, and obviously parameter count is not the NLBL and now, you know, smaller models are actually doing really well, but I think it just kind of illustrates to people like, What is this thing that now everybody's talking about and why now? Is it just a fad or is there like a real foundational kind of technology, uh, you know, kind of improvement, sort of like with the 8086 processor that has made this the time to actually pay attention and care and like t…
AI assessment note: “we actually now run this AI workshop program. It's a lot lighter touch”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q One, one thing that you said at the beginning is that like kind of the no code, um, enterprise app platform category is the thing that came before Airtable and to some degree you're that, um, how does it change your thinking about Airtable now that code is becoming easier to generate?
A Well, um, you know, well CodeGen basically replaced the need for vertical software and will replace the need for no code because now like even code is so easy to generate. And I, I have a very specific point of view on this, which is, you know, I think, um, you know, sure you can generate small snippets of code very easily, and maybe that's getting better and better with the more advanced models. Um, you know, I think code, uh, is obviously one of the core capabilities of all these LLMs, and I think it's, you know, has some nice properties of being, um, you know, simulable, so you can, you know, you can actually do a good job with synthetic, uh, data and training, uh, on it, and there's just so much on the corpus of, Code out there that, that, uh, you know, there's some really interesting things you can do with training at making the models better. Um, and there's some really interesting innovations happening out there, right? With, with not just the big companies, but like the startups, like the magics and the pool sides and so on of the world. Um, that being said, and you know, this may come down to as much a religious debate as how close are we to AGI? I think it's going to be fundamentally hard, like really, really hard to generate Really sophisticated end to end process automation type apps. So if you think about like a bespoke solution for content production, I mean, it's…
AI assessment note: “fundamentally hard, like really, really hard to generate Really sophisticated end to end process”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Basically, you're moving from a human-to-human collaboration company to a human-to-AI collaboration company over time in some sense, because, you know, what you're describing seems like a really interesting way to have co-pilots augment humanity or augment creativity. Are there other ways that you've thought about the substantiation of that sort of creativity augmentation or how AI really interacts with human creative potential?
A Well, and these are just examples of things that I have seen or, or thought about that I think could be cool in the creative space because you asked about, but I think in the design context, um, one thing that really matters a lot is the iterative loop and being able to keep going back and forth, uh, to an agent and give more instructions over time. If you just kind of like go to first principles here, there's so much that you're not able to communicate via a prompt. Like, if you think about great design, it often captures something about the culture, the ethos of the moment. It captures, uh, something about the temporal aspect of the sequence of interaction someone's having or the, the context they will have mentally. Something about affordances, what people are used to in terms of the language of design, uh, which is sometimes similar and dependent on the platform, you know, but oftentimes there's something about emotional state too. Uh, there's, you know, videos that the designers probably watched or, or in-person research interviews they've conducted. And so I, I think like fitting all that plus the product requirements, um, plus visual style into a prompt that's hard. Even if you could just get unblocked by an AI helping you brainstorm and thinking through problems, you know, that's your first sort of draft. And from there you can keep iterating from there. You can Keep ev…
AI assessment note: “in the design context, um, one thing that really matters a lot is the iterative loop”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q to take on the dull and dangerous jobs that humans shouldn't be doing. In this clip, we talk to Brett about how he runs a team with Velocity to drive hardware, software, and AI into reality. Big question, but can you describe, like, if you want to run a hardware project, a hardware and software project like this, with this complexity at velocity, like how do you manage product development?
A From like a thesis perspective, I, I strongly believe in like an iterative design approach. We really don't believe on spending a lot of time, like just, just doing research and analyzing. We spend a lot of time on just testing, building the testing here. And, um, that helps us really shake out all the problems. It helps us learn, helps us recursively add it into a continuum of product that's coming down, uh, coming out. And, um, so first that's our strategy. We, um, we want to be continuously updating the hardware and software forever. It'll, I don't think it will ever be good enough for us. Um, So we have a whole process built around building a robot from a, like a basically hardware and software design that we run here. We first set out with understanding who are the customers? Like, what does the robot need to do? From there, we, uh, we basically set requirements like, okay, we need the robot to lift this much pounds. It needs to run this long. It needs to charge here, the safety requirements so that it can't Battery can't burn down the building. There's a bunch of stuff we have to, um, the environment on IP rating has to be done on level actuators. There's just a bunch of requirements that come from there. From there, we look at those requirements and we, we do engineering design and we have basically Like three big phases. Uh, we have a conceptual and preliminary and crit…
AI assessment note: “we have basically Like three big phases. Uh, we have a conceptual and preliminary”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q foundry for AI. Here, we talk about what's next for Scale as models, uh, approach and go beyond human abilities. With great power comes great responsibility. Um, if, uh, you know, if these AI systems are what we think they are in terms of societal impact, like, trust in those systems is a crucial question. Like, how do you guys think about this as part of your work at Scale?
A A lot of what we think about is how do we utilize, how does the data foundry, um, enhance the entire AI lifecycle, right? And that lifecycle goes from, you know, A, ensuring that there's data abundance, as well as data quality going into the systems, but also being able to measure the AI systems, which builds confidence in, in AI, and also enables for further development and further adoption of the technology. And this is, this is the fundamental loop that I think every AI company goes through, you know, They, they get a bunch of data, or they generate a bunch of data, they train their models, they evaluate those systems, and they sort of, you know, uh, go again in the loop. And so evaluation and measurement of the AI systems is a critical component of the life cycle, but also a critical component, I think, of, of society being able to build trust in these systems. You know, how are governments gonna know that these AI systems are, are safe and secure and fit for, uh, you know, broader adoption within their countries? How do, how are enterprises going to know that when they deploy an AI agent or an AI system that it's actually going to be good for the consumers, and that it's not going to create greater risk for them? How do, um, how are labs going to be able to consistently measure what are the intelligences of my, of the AI systems that we build, and how are we going to, you …
AI assessment note: “evaluation and measurement of the AI systems is a critical component... of society being able to build trust”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Uh, okay. We're still in asynchronous video creation land, um, in 2024. How do people use Hey Jen today? Like, what are the favorite use cases you have?
A I would categorize the use, uh, the use case of HeyJet into three. Create, localize, and personalized. And, you know, people can select, uh, um, you know, the cast from a library from our avatar, or create their own digital twin, and just like select a template or type the script and generate a video. This works the best for product explainer, how-to videos, learning development, and some self-enablement training content. We can also Take existing video that localized that into a hundred, more than a 175 different languages and dialysis. And, you know, in this way, we can help customers to really localize their content into local languages. And last but not least, people can also use Hadrian to personalize the video messaging at scale. So I think there was a many, many very creative use case on Hadrian today. We are a very horizontal platform. I would say one of my favorite, um, use case is probably the recent launch with Madonna, and Madonna's launched a sweet campaign where they can allow people to send a message to a family member in the different languages.
AI assessment note: “I would say one of my favorite, um, use case is probably the recent launch”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What, what are you most excited about in terms of like what's next or new releases you guys have coming?
A I think there's many things very excited going on in our technology and product roadmap. I think particularly I'm very excited for the full body generation of the avatar. Historically, all the other technology has been focused on the upper body. It's really hard to generate the gesture and the emote, the body motion, but a lot of academic research has proven that this is very possible now. And we just need to, like, basically take that into the last mile. And another thing I would say, um, something I'm very excited about, the streaming avatar, uh, especially with the latest release on GPD for all, really, really helped to improve the performance of the real-time interaction with text and voice, and HN avatar could become a visualization layer for all those applications.
AI assessment note: “I think particularly I'm very excited for the full body generation of the avatar.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q application areas. Um, I, I guess, uh, related to that, what is the technology that you folks are using today? You mentioned some things like GPT four Oh, but you've also built your own models in house. Like how, how do you think about the technology stack that you're using and how does that have to evolve in order to be able to do full body or other new things?
A There's a three more doubt, right? Voice and video. Um, so we work with, uh, um, OpenAI, ChatGPT on the text generation side. Obviously also serves like the brain of the orchestration engine that we build internally. And we work with, uh, you know, OpenAI and Event Lab on the voice engine, but we build the entire video stack in-house, including after creation, video rendering, and B-roll generation. So I think over time, I think the whole technology trend has been moving towards to a direction. A lot of all these things will be trained together. The multi-model model, multimedia, all get into one single model. Um, one of the challenges I want to call out for the full body generation is actually, how do you actually connect that voice into, together with the, you know, gesture motion. And that's actually something will be unlocked by actually getting the voice model and the, um, video model training together. So that it can sort of like build a connection underlying the model as well. And that has been historically really, really hard because we have to like train the TTS model on one hand and then feed that TTS model outcome into a video model. And that's, it's pretty hard to build that connection, but with multi-model model training, that's very possible.
AI assessment note: “we build the entire video stack in-house, including after creation, video rendering”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q value is the fact that it is generalizable in different ways that didn't exist before. And it does open up the aperture in terms of one system that's kind of trained broadly, but then can do a lot of very specific subtasks. So could you explain more about why you don't feel that just scalability of LLM sort of leads in this direction eventually or scalability of some multimodal model?
A You know, the, the, the sort of claim goes like this, uh, effectively what large language models do today is they are high dimensional memorization systems, right? They are trained on lots of training data. They're able to find and generalize patterns off of the training data that they're trained on and then apply those in, in new contexts. And memorization is a form of intelligence, I would claim. Um, but it's not a form of general intelligence, right? We need something. There's something more that we need in order to be able to go discover and invent alongside us. You know, this is the things that I care about, like with AGI. This is why I want to build AGI. I think like, if we want to pull forward the future and actually have AI systems that are able to, you know, discover new branches of physics or pull forward our understanding of the universe, um, pull forward like new therapeutics. The answers to those don't show up in high dimensional patterns from our existing training data, because. Like the answer is, is literally unknown, right? The pattern is unknown. In fact, you might be able to find some sub patterns that can apply in like similar reasoning chains. And that's actually how current sort of AI agent systems work, right? If the reasoning chain that you need an agent to follow is simple enough such that the reasoning chain shows up in an abstract way in the training …
AI assessment note: “effectively what large language models do today is they are high dimensional memorization systems”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you know, you've, you've now established this arc prize, which I think is super exciting. It's a million dollar prize towards, um, you know, an open source model that, you know, meets certain criteria against your metrics of artificial general intelligence. Why do it as a prize versus investing in companies or, you know, taking a funding, uh, more traditional funding of startups or efforts model versus a prize model?
A I think outsiders are needed. Um, you know, there, there was 300 teams that actually competed in the ARC, like small version of the contest last year in 2023. And if you go look at all the teams that competed, you know, these are like one or two person teams. They are outsiders to the industry. They're not working in AI startups. Many of them don't even live in like the Bay area or Silicon Valley or California. It's a very globally distributed set of people with new ideas that are working on this stuff. I am more confident actually that, uh, or I guess I would bet that the solution arc probably comes from an outsider. Um, I think it's probably gonna come from somebody who's sort of not indoctrinated in the current way of thinking about language models and scale. Arguably like the solution arc doesn't even require that much scale. Um, you know, the, the cool thing about the puzzle, the RQGI values, it, it, it's like kind of a minimal reproduction of general intelligence. Uh, it fits into a two by two game board. That's like at max, like 15 by 15 squares big. Like it's, it's so small and reproducible. The data fits into such a small, uh, small set that, um, it's quite likely actually that the solution, um, it, it can be like written in like 10,000 lines of code or less. Uh, and it's not gonna require these like, you know, gigantic You know, two hundred billion large parameter mod…
AI assessment note: “I think outsiders are needed.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q just, um, zooming out, like, you know, you started by saying in order to have the debate about RAG versus alternative architectures for working on proprietary data, you need to predict forward, right? Um, any predictions for how these systems change as LLMs improve dramatically, right? If we look at the next generations of open AI and, um, or GPT and Claude and the Mistral models and Lama and such.
A Yeah. So my prediction is that the, the, the system will be simpler and simpler. Maybe this is, uh, my biased view. Um, so, or at least this is something that we are working towards. Um, so the idea of what would be that, um, it's a very, very simple system. So you just do, you just have three components like large English model, um, vector database and embedding models, and maybe four components, another ranker, um, uh, which refine the retrieved results. Um, and you connect all of this and each of the new artworks does everything else. Uh, you don't have to worry anything about trunking, multi-modality, changing the data format, um, because new artworks can do most of them, right? So seven years ago, if you talk to any of the so-called language models, seven years ago, you have to turn the format into a very, very clean format. Um, and now you talk to GPT-IV, you can have typos, you can have all kind of like a Weird formats. You can even dump JSON files to it. Um, right. So the same thing would happen for embedding models as well. So my vision is that in the future, AI will just be that, uh, a very simple software engineering layer on top of, of, of a few, um, uh, very strong neural network components.
AI assessment note: “So my prediction is that the, the, the system will be simpler and simpler.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q think, um, uh, respected for like the scrappiness culture, right? Hacker culture. And, uh, you know, one question is like, do you expect companies in YC to do pre-training and do their own foundation models in any way? Because the, you know, there is a, um, compute driven narrative that you need, A hundred million, and then a billion in next generation, closer to 10, in order to compete there.
A Totally. Um, I think people are starting to do it, and then I guess what, you know, it's kind of like when Cruz came through YC, here's this giant, like, mega research project, and you need a hundred million dollars to do it, and then you work backwards from that. Like, what are the milestones that we can get to so that, you know, we can draw a line from, you know, we have nothing, to we have something, to we really have something. You know, that's sort of the question always. So we absolutely have companies that are building foundational models. Um, you know, Diffuse Bio is doing over on the bio side. Uh, you know, uh, there, you know, there's a company working on, like, robotics foundational models. There's so many things that you could do. And then, uh, the half a million dollars, you know, during the three or four months, it's sort of what, like, Kyle did for Cruise. He had to figure out, well, how do I make a 10,000 dollar, um, automated driving Attachment to an Audi A four. And how do I get people, get people to pay for it? How do I, you know, all he did was actually make a demo that, uh, funny, funny enough resembles full self-driving in any Tesla car today, but this was in 2013. You could drive up and down one on one. It didn't do streets. It didn't do cities. It only did, um, uh, you know, driving up and down the highway. And that was enough to sort of Show that, and t…
AI assessment note: “we absolutely have companies that are building foundational models”