Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I severely underestimated how difficult this would be, and I haven't done anything. And so this is like one of these things that really puzzles me. It's a computer program. It has access to computing technology, right? It can do these calculations. Why do models sometimes exhibit this behavior where they cheat at a goal that you give them, or they pretend like they're doing something, and they don't? Monty?
A Yeah. So there are a few, a few different reasons for that, but generally, you know, it's going to come back to, um, something about the way the model was trained. And so in the example that Evan gave, which I think maybe related to the root cause of the, the, you know, the anecdote you just shared, um, as Evan said, when we're training models to be good at writing software, we have to evaluate along the way, whether they're You're doing the, you're doing a good job of, of the writing the, the program that we asked them to write. And that's actually quite difficult to do in a way that is completely foolproof. Right. It's, it's, it's hard to write a specification for like exactly how do you, you know, cover every possible edge case and make sure the model has done exactly what it was supposed to do. And so during training models try all kinds of different approaches, right? Like that's kind of what we want them to do. We want them to say, well, today I'm going to try it this way. Maybe I'll try this other way. Some of those approaches involve, involve cheats, right? And in, in, in the, The Claude, 3.7 model card, , we actually reported some of this behavior that we'd seen during a real training run where models got, uh, you know, sort of developed a propensity to hard code test results, right? And this is sort of what Evan was alluding to. So sometimes the model can kind of figu…
AI assessment note: “it's going to come back to, um, something about the way the model was trained.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q often you see people who are either part of these houses or, you know, in their orbit, They attend certain programs, right? Programs like Landmark, for instance, or programs like Hoffman, which you have also spent some time reporting on. So talk a little bit about, you know, what these, their self-actualization programs are, basically programs built for the type of people that gravitate to Silicon Valley. Am I right?
A Yeah. I mean, I, I want to, you know, be careful to caveat with like, you know, not everyone falls into all these categories, but I think if we're speaking broadly, like you're right. People who gravitate to Silicon Valley, they're looking to make a difference. They tend to be very ambitious. Um, they want to do things. You know, from first principles and or slightly differently than maybe how they perceive others in the past have done it, and I think that type of person is drawn to often these programs that I would call personal development or personal transformation. There's kind of a whole industry of them, and some of them you might recognize, you know, Landmark actually has a decades-long history associated with this, like, predecessor group called EST. Um, Tony Robbins is kind of like a classic example of these, um, but some of the ones that I know are popular among tech types right now include the Hoffman process, which is, um, yeah, something that, like, I am curious about and, like, definitely want to, like, potentially report more on, but that's, like, a week-long intensive Therapy-ish retreat, um, that kind of helps you process some of the, um, stories and narratives in your life. Oftentimes these programs will help you like reassess the stories in your life with the idea that it can help you unlock like a new level of performance. Um, I know a lot of people who have…
AI assessment note: “some of the ones that I know are popular among tech types right now”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I mean, what is it? It's interesting that that has to be said out loud. Um, like what would be the alternative be that you sort of go with the flow and yeah.
A Or that you feel like you're, you're kind of like a victim of your circumstances. Um, what's interesting to me about agency being popular in my view, among a certain like circle of tech people is that. Agency is a different way of saying this idea of like, oh, you should consider yourself like radically responsible for your life experience and the things that are What's happening in your life? And that's very much an idea that comes from these personal transformation workshops and lineages. So that is something that you would find at landmark. That's something you'd find at Tony Robbins for sure. That's something that comes up a lot at one taste, um, in conscious leadership group and also this. And so it's sort of been reframed as agency, but it's all getting at the same thing. Um, and I would say the opposite is yeah, someone, you know, you can imagine someone who laments The circumstances of their life without thinking about how they might change it, you know, and, and they, they're like, oh, I was just born, you know, without the ability to like charm people. It's like, well, guess what? If you were more agentic, you would think of yourself and your personality as more malleable. You would think that you could learn skills that you might otherwise like dismiss. So it's people, it's drawing a distinction between people who kind of are like throw up their hands and say like, w…
AI assessment note: “Or that you feel like you're, you're kind of like a victim of your circumstances.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And the people that came in, didn't they call them something like?
A Oh, marks. Yeah. Sometimes they would, um, jokingly, but also it's obviously somewhat serious. They would, um, refer to these potential customers as marks, um, which is a suggestion of, you know, one of the allegations leveled at this company by, by many of its former members is that its sales practices were very predatory. So again, the way this, the way that The way that One Taste made money was by selling courses, and it wasn't just courses on orgasmic meditation. If you got deeper, it would be courses about, like, how to live your life in alignment with the philosophies of orgasmic meditation, and these courses could cost upwards of 20 or 30,000 dollars, and they might be, like, two-week intensives or this kind of thing, and in that way, again, those group, those transformational, In a way, those transformational packages and courses are similar to like what you might find at like, you know, intensives at like Tony Robbins or landmark or that kind of thing. So again, yeah. So the typical person like comes in, they, they go to these intro evenings where again, everyone's close day on and you just play communication games where people like talk openly and vulnerable, vulnerably about their feelings. And maybe there's like a little bit of, Suggestive or sexual like undertones, but it's not like a sexual experience, but then like the people who work at one taste are so friendly…
AI assessment note: “Oh, marks. Yeah. Sometimes they would... refer to these potential customers as marks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Ok, so then, um, just stick to the details of the court case, but where do things go wrong?
A Well, um, basically the company was doing fairly well and, like, had a lot of, like, mainstream success. Again, Nicole, the founder, was speaking on stage at the Goop Health Conference in 2017. Um, they had these, like, endorsements. They were making money. And, um, in 2018, I wrote this big investigation for, um, Bloomberg Business Week where it was the first time that, um, people Um, in the company or former members of the, the group talked to me at length about these allegations that the group was a cult. Um, they basically said that they'd been exploited financially by being pressured to take on debt in order to buy more of these expensive courses. They said that they had been, um, exploited sexually by being pressured and sort of taught these lessons that, um, you know, pressured them to have sex that they didn't want to have in order to like further the company's business. So I wrote this big story, and the company kind of went into hibernation in response to that, and around the same time, and in all likelihood spurred by the story, the FBI started investigating, and then that led to many years of, like, the FBI looking into whether a crime had happened here, and then in twenty-twenty-three, federal prosecutors charged Nicole, the founder, and Rachel Cherowitz, who was kind of her, like, Second in command, the woman who had been head of sales for a long time, charged bot…
AI assessment note: “federal prosecutors charged Nicole, the founder, and Rachel Cherowitz... with, um, forced labor conspiracy”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q know, I've lived in, I mean, I'm using Silicon Valley as a broader term for San Francisco and, and, um, the Valley itself. But, um, when you're out there, it seems like a lot of people just want to believe in something, right? You're there to do something. You want to believe in something. And so once you get going and you believe you have agency, it's tough to stop.
A Totally. And, and it also, once you've dedicated a lot of your life toward a belief, To change your mind and say, actually that belief might have been misguided, or it shouldn't have been implied in the circumstance. There is a lot of cognitive dissonance or sunk cost that comes up in a situation like that, where like, there were lots of people at one taste who had tough experiences and they had a moment where they could have thought to themselves, actually fuck this place. Like, I don't want, you know, this place is, is hurting me. But in order to say that they would have had to had, they would have had to grapple with the Cognitive dissonance of like, well, I also had invested five or six years of my life into this group. It's a really hard thing to admit to yourself and also sunk cost. It's like, I've already dedicated so much time of time, money, energy, social connection in this group, like to leave it feels very costly. It feels very painful. Um, and so you're totally right that like ideology has this stickiness to it, which is the more that you Orient your life around a belief. If it becomes part of your public or professional persona, if you know, if you are like, yeah, the startup founder who is fighting for, you know, X, Y, Z, the more you put yourself in that direction, the costlier it is for you to change course.
AI assessment note: “Totally. And, and it also, once you've dedicated a lot of your life toward a belief”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You can take rings of data and drop it in ChatGPT and then in natural language, you can query it and it will give you answers today. Um, But if you're working on the line in a factory, or if you're managing operations in some industrial settings, it's hard for me to fully picture how Generative AI can then be applicable in that setting. So how do you do it?
A A hundred percent. Well, think about an example that we're working on. Real example, real use case we're working on in development with Boston Dynamics, with Eversource Energy, and with Anthropic, ok? Three parties orchestrated as one to make manhole, Duct inspections operate at a different level, ok? So what happens in practice? We send a Boston Dynamics Spot robotic along a five kilometer manhole duct. The Eversource are required by Boston law, in Boston law, Massachusetts, that by law they have to inspect that manhole duct on a, or a periodic basis, ok? Now that robotic dog is picking up LiDAR, picking up video, image, gas sensors, heat, Temperature, pressure, all the way through that manhole duct, ok? They're spotting issues and fractures and, and problems in that environment that the human often will miss, or they might not get with the same level of accuracy. Something spotted a stress fracture on a transformer, immediately captured. GPS coordinates immediately triggering a work order, looking for the spare part, dispatching a crew. That's all happening pretty much instantaneously. Now you think about the alternative when the human's doing that, often they don't want to do that work, it's difficult to resource it, and they might not spot that, and a catastrophic failure might arise, and therefore you see issues in uptime, you see issues in transmission, all sorts of probl…
AI assessment note: “Real example, real use case we're working on in development with Boston Dynamics”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q there are arguments to say, okay, fund it with debt. Who cares? You'll pay it back. Everything's growing well. Um, but I, I actually will, we'll turn it to you and, and, and hear your perspective on why debt is such an issue here. We, we've just started to see debt make its way into this conversation. So how, why is it a problem and how concerned should we be?
A So we have to go back to finance one-on-one, right? There's certain things we, we finance through equity, through ownership, and there's certain things we finance through debt, through an obligation to pay down interest over time. And as a society, for the longest time, we've had those two pieces in their right place, right? So debt is when I have a predictable cash flow and an asset or and or an asset that can back that loan, and then it makes sense for me to exchange capital now for future cash flows to the lender. So again, the conditions are an asset that is long standing that can back the loan, And or predictable cash flows to support the loan payments, right? That's why we have a mortgage. A mortgage is an example of both, right? A mortgage is, wait a second. If I, if I stop paying my mortgage payments, the bank owns the house. And since they only lend me 80% of the value of the house, even if the value of the house goes down a little bit, they'll be fine. And they have an access to my income, which is relatively predictable, even on Wall Street. And so they know that I'll pay my, my mortgage payment. That's, that's a loan that should be there, right? We use equity for investing in more speculative things for when we want to grow, and we want to own that growth, but we're not sure about what the cash flow is going to be. That's, that's how a normal economy functions. When…
AI assessment note: “When you start confusing the two, you get yourself in trouble.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You can't trust, you can't trust them because you just don't, how do they know?
A Exactly. So, so one, AI may turn out as well as we expect, but it may not. And two, OpenAI is not in a vacuum. They're competing. Part of the reason they're overpromising and creating this too big to fail and, and fake it till you make it and getting everybody else to have skin in the game. Part of the reason they're doing that It's because they know they're competing with Meta, and with Google, and with Elon, people that have a lot more resources than they do. So for them to say, oh, we're gonna have a hundred billion dollars of revenue by 2027, which Sam Altman just did, is completely disingenuous. He has no idea. He's competing against much bigger, more powerful companies that have technology that's at least as good as his. So, so lending money based on that is, is dangerous because again, these GPUs, you're building a data center, you're renting out GPUs, and right now maybe you're renting out a GPU for four dollars an hour, and maybe that way the business makes sense. But these GPUs keep getting so much better every year that that same GPU in just three years may be only renting out at 40 cents an hour, at which point the data center is literally worthless. Because that won't be enough to cover the expense of operating the data center. So this is where we get in trouble. When somebody underwriting a JP Morgan, a US bank or Mitsubishi bank ignores that. And to answer your q…
AI assessment note: “He has no idea. He's competing against much bigger, more powerful companies”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q these companies say that these chips will depreciate over five or six years, but like you said, if the, if the NVIDIA chips get that much better, that much more quickly, um, we can have a much more accelerated depreciation making the data centers that they're investing billions in today. Worthless, as you put it. How do you evaluate Burry's critique of the situation? Sounds like you agree with him.
A He's spot on. He's spot on. By the way, Big Short is the story of how he was spot on, but he almost didn't make it, right? A lot of the movie is about how long it takes to play out, and you can be right, but if you're right too early, you don't make it. And the story is about him and the handful of people that did make it. There were a lot of people that were short the market for a long time and lost everything because they couldn't wait Long enough. He was just in a position to wait and he's spot on right now. And look, depreciation gets wonky. So let me just hit it at a high level because it is really important to this conversation, right? Depreciation is based on an accounting standard that helps companies say, well, how I have an asset. How long is it useful? How long can that asset generate revenue for me? And if it, if it can generate revenue for me over five years, Then I should take the cost of acquiring that asset and spread it over five years as an expense for accounting purposes, right? That's, that's what accounts are there to do. And these accounts spent time three to five years ago with companies like Microsoft and Amazon and said, you know what, based on where the technology is now, we're looking at these chips and it looks to us like they can generate revenue for you for about five or six years. And that's why we're going to allow you to depreciate that to exten…
AI assessment note: “He's spot on. He's spot on.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q is there, because these companies are writing the chip depreciation at six years, or maybe, you know, seven years, um, but they're not going to actually be doing anything, um, that, uh, effectively they're going to be overstating their, their profits. Uh, and, you know, smart, smart investors. Is he saying that smart investors will catch on and then think their valuations because of it or what's the risk there?
A That's exactly it. Is that all we're one account conversation away from having all these companies have to report much lower profits. And, you know, we use, uh, we use profit multiples to value companies. So if a company has to depreciate most of its assets, Over three years instead of five years, that means their profitability is going to go down proportionally. And, and we could have in, in, in, in a stylized case, the value of a company declined by 40% because an accountant said, you have to depreciate this over three years instead of five years. That's why this is very real. Sounds wonky, but this is very real. If you release a company's profitability by 40%, their value will go down by 40%.
AI assessment note: “That's exactly it. Is that all we're one account conversation away”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q companies financing in debt. So they're, they're basically what they're, I mean, what they're saying is, You, you, uh, this is not acting like a rational, uh, uh, technology because A, everybody wants market share and B, um, yeah, they're, they're willing to spend, uh, to get there. And so what do you think about that? Is it, is this gonna be a persistent issue? How should we view this?
A Yes. Yes. Cause so here, here's the thing. You, who are the players in this game theory, right? It's meta, it's Google, it's Microsoft, it's Amazon. It's companies that are used to win or take all markets. And they think of all markets as being win or take all market, meaning if I don't win this market, I'll get none of it, or at least not enough of it that will be meaningful to me. So they are willing to do anything to win, which to the point of that means they'll be willing to lose money for a long period of time, so they have a chance to win. And what happens then is it's only the biggest, most deepest pocket player that can win because they, they can wait it out or at least Communicate to everybody else that they're willing to wait it out. So that's, that's where a company like OpenAI has no chance, because they can't make it through another year or two at this level of spending, so they certainly won't be able to outlast Google Meta and Microsoft in this game, right? So, and it explains a lot about Mr. Zuckerberg's behavior. Again, he's not just spending the money, he's telling us he's willing to spend anything to win. He's signaling to all the other players is I will not lose. So you can keep throwing money at this. I'll keep throwing money at it longer. And that is exactly where we're at, which is why we may have persistent losses for a while here, because these companie…
AI assessment note: “which is why we may have persistent losses for a while here”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I'm not saying that's what's gonna happen with you. Um, and maybe it makes sense just to, you know, buy off the shelf or, uh, use open source. And in fact, that seemed like that was the strategy that you had for a long time. It seems logical. Um, and so I'm curious, like why you would disagree with that? Why is it so important to build your own models?
A I mean, we, we're going through a foundational platform shift, um, you know, in software, um, from the operating system to apps, from browsers, search engines, mobile, social. This is the next major platform, and it's going to be bigger than all of the other platforms put together. So the idea that a three trillion dollar company with three hundred billion dollars of revenue and 80% of the S&P 500 on our Azure stack and M three five stack Um, you know, could, could depend on a third party. This it's, you know, just in perpetuity. It doesn't make sense. So we, we, you know, this is a company that's been around for 50 years, uh, and navigated many of the past platform shifts incredibly well. And that's the, that's the journey that we're on. We have to be AI self-sufficient. There's an important mission that, uh, Satya set last year. And I think that we're, we're now on a path to be able to do that.
AI assessment note: “The idea that a three trillion dollar company... could depend on a third party”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q talked about, I think, on all sides of the political spectrum, the idea that America's fallen behind in, in critical infrastructure. So if that is the case, and we are buying into this story that AI is going to be the battleground of the next century. Does it seem okay that they're asking for federal guarantees or a backstop in terms of all the debt financing that's being taken out?
A I mean, I think that's my point here. I think it's personally, it's perfectly reasonable for OpenAI to ask. Um, I don't think that the US taxpayers should be backstopping the company's debt though, because, and Sam Altman in a follow-up made this clear, if they fail, you'll still have Google, you'll still have Anthropic, you'll have many others that are going to be building this. Um, so in, in other words, I do think that the United States is, is gonna, is going to be in a good position. And I also don't really think that the government should be picking winners, uh, and giving open AI these guarantees, assuming that open AI would be the only one to get them. So to me, it's just like, it's a personal, it's a perfectly reasonable ask, um, and it's a perfectly reasonable, uh, no from the US government and David Sachs, the AI czar, basically said no to bailouts.
AI assessment note: “I don't think that the US taxpayers should be backstopping the company's debt though”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I cannot think about from a multiple perspective. I mean, that's already just insane, but I'm sure we've seen that at some point From a, like, loss perspective, certainly not, because I, I can't imagine at this scale any company ever losing that much money and being able to, like, confidently even think about an IPO. So do you think it happened? When do you think they go for it?
A I think it's going to be a while. I mean, again, and just some, so, so Altman, uh, did tweet out, he did this live stream this week. He treated, he, um, shared some perspective there and then tweeted out some notes about it. I mean, on the live stream, he said, OpenAI has 1.4 trillion worth of financial obligation. Uh, and it's made commitments to use 30 gigawatts of data center capacity. Uh, but OpenAI's revenue is going to be thirteen billion this year, and we already have eight hundred million people using ChatGPT. I just don't see how you go from, and I could be wrong, but I don't see how you go from OpenAI's base of revenue now to, uh, 1.4 trillion dollars that it can spend and, and ever have a economically viable company. I mean, we could go back to this and I'll look like an idiot when OpenAI pulls it off. Um, but to me, the financial picture just seems absurd.
AI assessment note: “I think it's going to be a while.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q bots to go in and find some information for you, uh, if I'm reading you right, if I'm hearing you right, the answer is that they just are dealing with a con, you know, just an unbelievable amount of data. They're struggling to pull it out. Um, so how do you, how do you then go up and, and put that infrastructure into limit and then find the right stuff?
A From a data perspective, you know, by the way, we acquired a company a few years ago for enterprise search and the infrastructure and platform that they built was allowing us to synchronize any application in real time and store that data in a universal database. So a data model that is consistent across the board, regardless of what application it is. We can plug into HubSpot and Salesforce and Slack Uh, and your task management platform if it was outside of ClickUp. And we can synchronize that into one database in real time with permissions and privacy aware. And this is what we optimize for AI retrieval. So our, our agents understand this, this singular data model. It can do semantic search in one place. It doesn't have to go kick off 10 APIs and then try and pull that back together. Uh, and then do re-ranking. It's very slow. Um, and it just doesn't work at scale. Use it using APIs to do that. Uh, so certainly this, this is where the world will be headed is actually unifying that data into like a single database. And I do think that's part of the reason why, you know, Salesforce started putting up those, those walls, um, is because they, they do want to be that provider. They want those, those, that database, that full universal data model inside of their ecosystem, not somebody else's.
AI assessment note: “the infrastructure and platform that they built was allowing us to synchronize any application”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And look, I know we're 15 minutes in, so I think we should probably take a minute to talk about the concrete things that you've improved in Claude between four and 4.5. Do you want to just give us briefly a little bit of a list of the things that get better with the new model?
A Yeah, I think the, the ones that I think are, are highlights, um, maybe I'll, I'll, I'll bucket into three. One is from a, uh, price performance, uh, perspective. So, 4.5 sine 4.5 basically outdoes Opus, our largest model in effectively every category, but does so while running faster and at a fifth the cost. So if you think about where we were in May at, you know, code with Claude, we were announcing, announcing Opus four. We now have a model that is better than that, and even its successor, Opus 4.1, but does so to fit the cost, which is very like, you know, opens up a whole new set of, of use cases for that kind of intelligence. That's, that's one on the price performance piece. Um, the second one is, um, on its ability, not just to, to code for, for longer, but just execute agentically for longer. We talked a little bit about agents, but, um, what we saw was, um, and actually I put a fun video of this, uh, on my X account, which is we asked every Claude from Claude one to Claude, you know, 4.5 to recreate Claude.ai. So like our flagship AI products. And, um, 4.5 was really the first one that was able to do it end to end and actually produce something of, you know, quality. It actually works. You can log in, you can use an API key, all of those things as well. Um, and so that ability to like execute agentically work for long time horizons. We had one customer had it work for…
AI assessment note: “I'll bucket into three. One is from a price performance perspective.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I mean, do you know what I want though? I want a super brain every morning when I wake up writing me a morning brief, a personalized report while I'm sleeping five to 10 briefs a day that get me up to speed on the day and encourage me to check, chat, chat, GPT first thing in the morning. Do you have anything like that for me?
A Well, I think you are touching on what may be the beginning of OpenAI's ability or the product that might spark OpenAI's ability to pay all this money back. It's called ChatGPT Pulse, and we'll talk about it right after this. And we're back here on Big Technology Podcast Friday edition. In the first half, you might have heard us talk with varying degrees of trepidation about the economics of the AI business. Why don't we talk about one way that OpenAI might be able to justify its valuation, and that is continuing to innovate on the product front. This is from TechCrunch. OpenAI launches ChatGPT Pulse to proactively write you morning briefs. OpenAI is launching a new feature inside of ChatGPT called Pulse, Which generates personalized reports for users while they sleep. Pulse offers users five to 10 briefs that can get them up to speed on their day and is aimed at encouraging users to check ChatGPT first thing in the morning, much like they would check social media or a news app. Pulse is part of a broader shift in OpenAI's consumer products, which are being designed to work for users asynchronously instead of responding to questions. What do you think about Pulse, Ronjan? You excited about it? Is this, is this the killer app?
AI assessment note: “It's called ChatGPT Pulse, and we'll talk about it right after this.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q who are Building these things practically. I want to get a little bit of insight from you in terms of how this is happening. Obviously you work at Okta. Okta is helping companies set up agents. So how exactly is this process taking place of you working with companies to be able to handle some of these tricky things we talked about in the beginning and actually set up agents?
A Yeah. So we, we do four things, uh, currently that really help, um, our customers and the developers that are, that are building this agent experience, right? So the first one is pretty simple, which is we verify both the agent, uh, and the user. So making sure that You are who you say you are. You're Alex, and that the agent that you, that you have essentially, uh, consented this agent to go do stuff, um, on your behalf. The second thing that we do is we provide capabilities for, um, our customers essentially, um, uh, secure APIs. So this capability called Token Vault, because in this world, You know, agents are going to be talking to lots of systems, and it's really cumbersome to go system by system or API by API and figure out how to handle their security. So we do this in a scalable way and make it super easy for a developer to use our product to essentially make sure that all of the API and agent communication is secure. Then the third one is agents will always need humans in the loop, or at least At the moment, right? And like just this example I had shared about, um, the travel example, right? I want to go to Japan in, in November and give a bunch of criteria to the agent to go find me the best itinerary. But before the agent purchase attorney, I probably want to review it, right? So there are always tasks that you will want to review. And so we call this having human in…
AI assessment note: “we do four things, uh, currently that really help, um, our customers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q which is that it's a very promising technology, but companies trying to put, because if bad guys are companies, they're businesses, companies trying to put it into action have seen mixed results. And so actually what I'm getting from you is same thing happens happening with the bad actors. But then the question is, if these models get much more intelligent, uh, does that open us up to bigger risks?
A I believe there are areas that Yes. If they become increasingly better, let's take vulnerability research as an example. Vulnerability research is one area. Let's explain what's a vulnerability. Okay. Vulnerability by definition is the ability to move from one trust level to another trust level in a way that is not permitted. For instance, if I can run remote code, then I'm moving from the outside to the inside, and this is the worst that can happen because I'm literally running code from remotely, external, In your internal environment, right? So that's a vulnerability. A vulnerability can be un, like, um, unauthenticated access, so authentication bypass. I'm logging in and I have this trick that I'm giving False password, and I'm still able to log in, ok? That's authentication bypass. I was able to walk into an, like, a higher trust without the permission to do so. So that's a vulnerability. The ability to research and find vulnerabilities, this is kind of the bottleneck of the security space, ok? Because vulnerabilities are what allow threat actors to move from, you know, lower trust to higher trust environments, and The ability to automate research by, ah, ah, you know, AI of vulnerabilities can open up maybe a race where you can find many vulnerabilities and unable to patch them at the same pace. There are solutions today that are already leveraging AI to detect vulnerabil…
AI assessment note: “I believe there are areas that Yes. If they become increasingly better”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Right. And so the week we're talking, you at Box are releasing a number of different agents. Um, let me start this discussion by just asking you, what is an agent? Because it does seem like it's an overused term and, and even myself who I'm, I'm in this all the time. I don't fully have clarity on what that word actually means.
A Um, I, I think the, ah, I think we should anticipate that it's fully overused. It, it is now the new term of art for talking to a, an AI system that is doing work for you. So just, we will hear, this will be the main term that we use going forward as an industry. And not because it's a buzzword, but actually it's a, it's a useful term. It's a, it's a definable object that is doing automated work for you. That could be in some cases as simple as answering a question. Um, but I think most people in, in the tech industry would generally argue that it should be doing some degree of, of work and looping through the AI model multiple times, um, uh, to do that work, and so, uh, that could be everything from, you know, very clearly something like Claude Code, or Cursor has an agent, or Replit has an agent, where you give it a task like, build me a website that has these qualities, and it will go off and do, you know, weeks worth of human work, In 10 minutes, and that's an agent that is managing that whole process, looping through the model multiple times, keeping track of what it's doing, updating its memory in the process, and that's effectively an agent. So that's an agent in coding, and we're going to see that same kind of agent architecture emerge in law, in healthcare, in finance, in education, where you can deploy agents to go off and do work for you. And, um, and, and there'll b…
AI assessment note: “a definable object that is doing automated work for you”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q we know that the other devices have been left in. Nolan is kind of moving about the world with a Neuralink, uh, in his brain right now. Uh, but this is something that has only been used in circumstances of like when the brain is already open to understand Uh, to try it out or understand a little bit more about surgery. So why haven't you left it? That's right.
A We, we've, we've taken, we will be leaving it in is the, is the short answer to your question. And the, the reason is that we've taken a slightly different approach to, um, to development, which is, um, to emphasize, to make sure that what we've developed is, uh, safe and highly functional before beginning permanent implants. And so, uh, for us, it's been incredibly important Uh, to ensure that the interface works and delivers a level of functionality that we think is, um, is essential, uh, to guarantee to patients before we start leaving the devices in. So we, we therefore pursued a, uh, a strategy that was sort of a phased development approach. Um, and so in our first 40 patients, um, those were temporary implants, uh, that were designed to Validate that the quality of the signals and the ability to decode those signals in real time. We then, um, you know, we now we're actually the first, um, modern brain computer interface company to have a FDA clearance. So the, uh, version of the electrode that you were holding in your hand actually has now, um, FDA clearance. So among the current leading BCI companies, we're the only company, uh, that has a full clearance from the FDA. And, uh, as part of leveraging that clearance, um, We're moving to a next phase in our clinical studies that will allow the system to be left in place for up to 30 days, and that phase will help us further …
AI assessment note: “to make sure that what we've developed is, uh, safe and highly functional”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Okay. So we will have to do another hour then is what you're saying. Uh, and then what about this idea of, of AI and human brains merging? Any thoughts on that?
A It's already happening, right? I mean, in some ways that's, that's exactly what we're developing. And we, we see the brain computer interfaces today as, as you alluded to, and as Michael mentioned, as kind of like the, um, in some ways the foundational layer of, um, you know, a merger between The brain and artificial intelligence right now, it has some very practical manifestations, which is effectively to become a different kind of user interface. You know, as Michael mentioned, we have a ways today that we've become accustomed to of how we interact with the digital world. And it's usually with voice or hand control. Um, but the technology that we're building is to enable direct brain to Digital inter digital world control. And right now, actually what we're doing almost of necessity, uh, because so much technology is just built around, um, voice and gestural and, um, you know, hand motor control is kind of a two-step bridge between neural intent and a conversion to what would be, for example, Typing on a keyboard or moving a cursor or speaking some commands to a computer, but that's just kind of a, an artifact of the way the user interfaces of today are built. We already know actually, um, that the latency between your brain and your hand and the ability to think something and type it is around 25 milliseconds. So that actually puts a, Biological hard limit on how fast you ca…
AI assessment note: “It's already happening, right? I mean, in some ways that's, that's exactly what we're developing.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q from the standpoint of a consumer, could this be good? I mean, it's pretty annoying to, uh, type this question into Google. When was Cloudfair, uh, uh, founded? And then have to click to Cloudflare's website to get the answer that Google could just service for you, and so much of the web has sort of become effectively the service of Google queries, where websites don't really need to exist.
A Well, you know, so absolutely this has been happening for a while, and if you look at up until six months ago, the ratio 10 years ago of crawls from Google to clicks was two crawls, one click. Six months ago, it was up to six crawls one click, and that's all because of the answer box. Um, what the AI overviews, which they've rolled out over that time, have done is they've taken it now to 18 crawls to one click. So yes, it, it is a situation of, you know, the frog boiling in, in water, but that's, it has gotten progressively worse. And I think across the media industry, it's gotten harder and harder to actually survive as a, as a publisher. And so what I worry about is, Yeah, you know, the publishers are struggling at, they were struggling at six to one. They're struggling at 18 to one. I think they're dead at 250 or 1500 to one that we're seeing with OpenAI and completely dead at 60,000 to one that we're seeing with something like Anthropic. And so that is the direction that things are going, and that's a challenge. I think you're exactly right on the other point as well, that the, that the challenge here is that this is actually a better user experience. That's why more of the web is going to turn to AI. Um, it is great that you can type something in and you can get back an actual response as opposed to having to hunt for it yourself. That's a better user interface. And so tha…
AI assessment note: “this is actually a better user experience. That's why more of the web is going”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And is, are these vibe coding programs or these bespoke programs that people are building with prompts? Are they in production or are they mostly hobbies that people fool around with?
A Depends on, um, first bucket is, is more hobby, personal life. Second bucket entrepreneurs. As you know, most startups die, so most startup ideas don't make it to fruition. The 10% of startups that are small businesses that get off the ground, they get the most value out of Replit. Uh, and some of them are in production now. Um, you know, I've, I've talked about a lot of these stories, but, you know, for example, we, we have this, uh, creator, his name is John Chaney. Uh, he's a serial entrepreneur. Used to take him many months and hundreds of thousands of dollars to build applications, and now he can spin up a business. And get to million dollar run rates in, in a matter of, of weeks. Obviously he has experience. Like he, he knows the formula of what it means to be an entrepreneur, but people can learn that over time. And in terms of the, um, enterprise, um, you know, we have, for example, Zillow. The CEO of Zillow recently on New York Times Dealbook talked about how everyone at Zillow is using Replit to accelerate product innovation, because product innovation no longer depends on engineers. You can have product managers do the entire iteration, getting user feedback, even without going to the engineers. So it just like increases, we have Duolingo, Um, a bunch of these customers that are really focused on innovating, building their second, third product, uh, that are now usin…
AI assessment note: “first bucket is, is more hobby... And some of them are in production now.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q they said founders in public, AI is writing 99% of our code. In six months, we won't need any engineers. Founders in the DMs, uh, does anyone know a good React developer? 30,000 dollar bonus. And I will name my firstborn son after you. So Can you explain that disconnect between this view that engineers are going away and this still like very intense demand for engineers in the market?
A I never made the point that engineers would go away. I make the point that entrepreneurs can start businesses without needing engineers and that we already see that. We already see, you know, I meet YC companies and, uh, Y Combinator is the most, uh, prestigious startup accelerator in the, in the world, Bay Area. And In the past Y Combinator would encourage you to go get a technical co-founder. But like we said, there's so many people with amazing ideas that don't have a technical co-founder. And so they're starting to get into YC. And what they tell us is we're just going to build this thing on replet. We're going to see how far we can get. And they often got, get really, really far. Now, if you're building a venture scale company and you want to like get to hundreds of millions of dollars of revenue and you want to, you know, become 1,000,000,010 billion, a hundred billion dollar company, you're going to have to hire engineers. But if you're trying to build, um, a company that creates a really great living for you, even, you know, you can, Potentially get rich from it. You, I think we're almost there where you can do it on your own without any developers. And so when I'm talking, I'm talking to our audience.
AI assessment note: “if you're building a venture scale company... you're going to have to hire engineers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Okay, I do want to ask you about something that you didn't mention when you looked at the different factors for why prices might not be going down. There might be investor pressure, uh, there might have been this equilibrium reached, or is it possible that these models have just gotten so big and expensive to run that the fundamental economics of AI are just not working? So explain why.
A Um, uh, You can surmise the bigness of the models based on speed, token, token throughput. It's not perfect, but, but if you remember GPT, 4.5, GPT, 4.5, uh, was an experimental model from OpenAI. It was the idea, let's train a trillion parameter dense model, meaning it is not sparse, meaning all the token, all the neurons are activated on every request. And it was so slow. It's really hard to run these things. The new models, even when they're big, they're sparse models. They're called MOE, mixture of experts. So in every request, there's a router layer that takes it to the expert part of the circuit in order to answer that question. So, you know, there are models with trillion parameters But any given request is thirty-two billion active, and that's like a kind of small model. Um, and what we're seeing based on speed and things like that, it's actually probably the models are getting more efficient. I mean, Deep Seek showed that the models are getting more efficient, and if, you know, Deep Seek open source was able to make it, you better believe that the labs are also getting more efficient.
AI assessment note: “The new models, even when they're big, they're sparse models.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q have seen is just a general increase in capabilities across the board. There were no caveats of like, um, and maybe there's a reason for those caveats, but there were no caveats of, you know, there's, uh, intelligence increases in this place and that place. It was We trained a bigger model, I'm pretty sure this is what it was, and it's better across the board. So, have things changed?
A They've changed, yeah, from a technical perspective. I think when you go from GPT two to GPT three, three to four, these were really just, uh, exploits of what was, uh, and is the scaling paradigm of training larger, pre-training bigger and bigger models, training larger models. Um, it's kind of one vector of training, uh, and you get a better model that, uh, as a, as a result. Um, and that continues to hold true, but we now have this kind of other category of, of, of training, which is post-training, Uh, and being able to use test time compute in more interesting ways than we used to as almost kind of a second stage of training. And so we think that that actually gives us a little bit of a boost, um, a force multiplier on our ability to push the model toward new intelligence levels, um, and also be able to train into it a lot of the things that you want an intelligent model to be able to do. Um, so using tools, for example, is something that really thinks really important, uh, for overall intelligence, GPT two and three, Um, couldn't really do that as well. GPT-IV could do it in a more nascent way. Um, and now GPT-V, you get that baked in, uh, with the benefit of, of these kind of multi, multi-step and, and longer horizon reasoning processes. So, um, yeah, we, we want to abstract that from users. Obviously, we don't think that you as a ChatGPT user should have to stop and thin…
AI assessment note: “They've changed, yeah, from a technical perspective.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, it is interesting to know that you've been, you have been already working with many companies, uh, and letting them use GPT-V already. So has there been a sort of unified, we couldn't do this with the previous models, but we can do it now with GPT-V or is it sort of spread out in terms of the capabilities that it's now enabling?
A Um, I would say it's, it's been, uh, you know, rising tide across the board. So everyone who's kind of benchmarking and all the companies that we work with typically now are, are pretty accustomed to, to evaluating and benchmarking performance across all the models that they use. But, um, everyone has kind of reported, you know, much higher, kind of consistently higher performance on those evals. There are a few areas in particular we've seen spikes. So one is coding for sure. Um, I mentioned companies like Cursor, JetBrains, Windsurf, Uh, you know, Cognition and others that we work with who, um, anecdotally are all, uh, you know, have, have all said that GPT-V now feels like the most capable coding model, whether that's in an interactive coding environment or more of an agentic coding environment. Um, and then also one of the things that we see consistently now is its ability to reason and problem solve in very technical domains, uh, is significantly improved. And so, um, Harvey's a great example of that where, uh, you've got, you know, Harvey AI working with legal firms and law firms, Uh, is, you know, very, very reliant on its ability to, uh, reliably, accurately, um, and, uh, and consistently portray, uh, you know, uh, cases that, that, that it's looking at, legal analysis, um, to provide that kind of level of structured thinking you want when you're doing legal analysis, a…
AI assessment note: “I would say it's, it's been, uh, you know, rising tide across the board.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But so many folks in the AI industry are talking about diminishing returns from scaling now. That really doesn't fit with the vision you just laid out. Are they wrong?
A Yeah, I, I, from what we've seen, I can only speak in terms of the models at Anthropic. Um, but what I've seen in terms of the models at Anthropic, if we look at, you know, let's take coding. Coding is one area where, you know, I think Anthropic models have advanced very quickly. Adoption has been very quick. We're not just a coding company. We're planning to expand to many areas. But if you, if you look at, if you look at coding, um, you know, every, you know, we released 3.5 Sonnet, a model we call 3.5 Sonnet V two. Um, uh, which I, you know, let's call it 3.6 on it now, um, 3.7 sonnet, uh, and then four point O sonnet and four point O opus. And, you know, that series of four or five models, each one got substantially better at coding than, than the last. If you want to look at benchmarks, you can look at, you know, sweet bench growing from, uh, you know, I think 18 months ago is that like three percent or something, um, growing all the way to, you know, 72 to 80%, depending on how you, how you measure it. And, and the real usage has grown up, grown exponentially as well, where we're heading more and more towards autonomously. You can just use these models. I think the actual majority of code, um, uh, uh, that's written, written at Anthropic, uh, is, you know, at this, at this point, uh, written by, or at least with the involvement of one, you know, one of the quad models, um…
AI assessment note: “we see the progress as being very fast, and the exponential is continuing”