The Wisdom Wall
1,824 quotable lessons, heuristics and mental models. Every one is playable at the moment it was said. No fortune cookies allowed. Showing the 400 best of this view.
“And compared to a few years ago, You do not have the need to go and hire people that you do not trust, that you do not trust to have high agency or the right skills, because you can use an agent to do those things.”
“And I see some other companies that are around our size starting to go hire a PM or a marketer or something like that. And that just feels like the old way of building a software business. And ultimately that's gonna lead to something where 10% of the people are thinking about how to make a great product. And the other…”
“Someone, someone, uh, someone could argue that if, if a country currently doesn't work on their own nuclear program, They're doing a disservice to their country and the government should be fired. Like, because it seems like from the recent world history that, that, that is like the only way to actually provide…”
“It literally becomes an issue of like raise capital, turn that directly into growth, use that to raise three times more. And if you can keep doing that, you literally can outspend any company that's built. Not any company. You can outspend the aggregate of companies on top of you, and therefore you'll necessarily take…”
“And now we've discovered that it is for a larger and larger and larger class of Piece of bodies of code. It is better to just start over and rewrite it from scratch than it is to try to fix it. The LLM will do a better job.”
“they tried to solve this on the latent space level. Um, I think I've It's shown every single time that that doesn't work.”
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to take it out. If it's really hard to put it in, it's really hard…”
“our take is that it is very unlikely that the optimal UI or the optimal interaction pattern for this new software development where humans spend much less time writing code. I think it's very unlikely that that optimal interaction pattern will be found by iterating from the optimal pattern when you wrote a hundred…”
“A lot of, like, agents that I see are, are, are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper bound of, like, complexity of problem that the…”
“with machine learning, first of all, you see that the performance of the models follows an S-curve. So it's not like it just goes off to infinity, right? And the, the S curve, it kind of plateaus around human level performance.”
“I think with agent frameworks in general, they can certainly save you some like boilerplate, but I think there's actually this like downside of making agents too easy, where you end up very quickly, like building a much more complex system than you need. And suddenly, you know, instead of having one prompt, you have…”
“If you're optimizing for cost, absolutely be remote. If you're optimizing for creativity, which I think that software and product building is a creative endeavor, if you're optimizing for creativity, it's kind of like composing an album. You can't do it on the cheap. You want the very best album that you can make, and…”
“you cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, like, you know, like, eight pm, I'm not going to, my job can die, and then the chips can stay idle for,…”
“this emergent, uh, behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
“If you ask the LLM to translate a bit of Python into a little bit of C, and it's performing this task, obviously it is understanding in the sense that it has a, The causal, functional model it implements.”
“Claude exists only as a pattern. It's something that is a pattern in the activation of the transistors. And even transistors don't actually exist. They are A pattern in the atoms that we are able to see as an invariance because we tune the atoms in a particular way, right? So we look at invariant patterns that we use…”
“The issue really is the fact that these observability companies isn't actually doing observability for the system, it's just doing the LLM thing. Like I still end up using like Datadog, right? Or like, you know, Sentry to do, like, latency. And so I just have those systems handle it, and then the, like, prompt in,…”
“so much of it is changing that if you give control of these systems away too early, you end up ultimately wanting them back. Like many companies I know that I reach out or ones were like, oh, we're going off of the frameworks because now that we know what the business outcomes we're trying to optimize for, these…”
“Like, I think we can, you know, 10 x the amount of developers, and still, you know, have a lot of people making a lot of money, you know, building amazing software, and also being, while at the same time being more productive. Like, I, I never understood this, like, you know, AI is gonna, you know, replace engineers.…”
“higher valuation, given the same company quality, is always worse.”
“RAG is basically just, just a hack, but it turns out it's a very good hack, uh, because Uh, what is RAG? RAG is you keep the model fixed, and you just figure out a good way to, like, stuff stuff into the prompt of, of the language model. Everything that we're doing nowadays in terms of, like, stuffing stuff into the…”
“Quantization is a poor man's compression. Um, I think we're only talking really here about, like, maybe a couple gigabytes, right? And then if you have, like, a couple gigabytes of true information of yourself up there, cool man. Like, what does it mean for me to live forever? Like, that's me.”
“I think it's actually not a question of whether the computer is aligned with the company who owns the computer. It's a question of whether that company's aligned with you or that government's aligned with you. And the answer is no. And that's how you end up dead.”
“really an agent is the ultimate settings screen for any software and code is the ultimate setting screen for any software.”
“On the other hand, if you think about using transformer architectures that have worked so well for language, that just wouldn't be able to support a five trillion context length. No matter all the compute in the world is thrown at it. So that kind of quadratic complexity is infeasible and also unnecessary because the…”
“you can use those PDEs both for numerical simulators to generate data, but if you're clever about it, you can even use them as a training signal. You can check how well is my model actually doing on the PDEs themselves. And use that as an additional training signal where you're suddenly in a place where you can push…”
“And the difference there is compared to language where self-improvement needs something like human feedback or other reward signals that are very sparse. They just tell you yes or no, thumbs up or down. We have dense feedback because the physics laws, there's so multiple of them, and you can decide again a curriculum…”
“My intuition behind the actual, when do you train or even post train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train is it already has the physics. We trust the…”
“This is the reason why you want to run a simulation. You spend five years, forty million dollars on this one study and have one finding. But if you can run simulation many, many times instantly, then that's the value.”
“You have some outer system that's optimizing some inner system. What if you want to optimize the way you're doing? You're optimizing, then you need some outer, outer loop, and it's this infinite recursion out. And the only way I think out of that is to collapse that loop down and make it so that the system itself is…”
“there are drug modalities that you just can't discover with immunization. Uh, like you're not gonna design your like crazy multi-specific Warheaded, super intense formats. Um, these are really things where you kind of have to design these from first principles. Uh, even just, uh, with bi-specifics in particular, like…”
“It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in is that you can predict which layers are going to have quantization errors that will…”
“And with training, it's more of like a math, like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning, like GPU kernel auto-tuning... you shadow the same traffic, like real traffic, and you just see which configuration gives you the best TPM, TPS, and you just use…”
“if you have a, a trillion dollar or five hundred billion dollar training run, then take fifty billion of that and make an ASIC. Like, it's fine. Like, like you will get more than 10% efficiency from From the ASIC. And like, that makes sense.”
“If you were to somehow be able to, in like this theoretical dreamland, have extremely fast NICs, you could, in theory, spare that HBM and you could just transfer KVCache trans, like directly from one node to another. This would give you like almost a hundred X speed up when you're doing this aggregated serving between…”
“Weights are a binary. Let's call them what they are. Yes, we can modify it and we can change them, but like, Giving someone the weights does not allow them ultimately to recreate what you're doing.”
“We are going to be able to squeeze so much more out of smaller models than I think we had imagined in the industry. Because yes, there's intelligence and larger models are more intelligent. Like no doubt about it. We should continue to scale up. Uh, but the behaviors Of being really persistent, of being able to…”
“the training run is not the expensive part. The training run is a very anticlimactic event, right?”
“And, and RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the entire web. But for RL, you have, you know, X millions of…”
“overall, if we have to give an order, my order would be the quality among scale of the datasets, and then the architecture, and then the prior knowledge.”
“science is as an infinite token generator to train models at scale.”
“It actually thinks in latent space, it emits tokens. So, like, the chain of thought is often an unreliable narrator for what the model, the computation of the model is actually doing.”
“So like we have just seen incredible lift from showing the model that even if we're at like a parameter disadvantage relative to the frontier models, just showing it an experimentally verified reasoning, reasoning trace, you, you see just immediate lift when, uh, when we do that.”
“And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole. Every, every person in every enterprise has different data, and you can really very easily pick and choose what goes in and what goes out. So the holy grail is have the model learn for itself, have…”
“The point for me is, the principle is, any kind of, like, efficiency and intelligence, they cannot really be decoupled. Sometimes people think if you're building something that's more efficient, that can save you dollars, therefore you're not in the premium category, you're in the, ah, you know, you're making the…”
“the only way to counter that, and again, this is throughout human history, throughout human evolution, is to increase the diversity of ideas, the size of the ideas, the size of the population, and the interconnectivity of the ideas. So rather than having individual monolithic models that we're all in, or, or…”
“people talk a lot about, we made these kernels faster and whatnot, but improving kernel only give you like a few percentage points of improvement and, uh, increasing except length, uh, literally is a multiplicative decrease.”
“the reality is that the translation of high throughput screens, whether it's Dell, DNA encoded libraries, or more traditional screens, the R-squared of those predictions to, like, the actual business of resynthesizing a molecule de novo and doing a Low throughput, high fidelity experiment is a shockingly low R squared.…”
“the kinds of chemistries that today can be automated are fairly constrained. And so your ability to search chemical space broadly for those really top top Pareto optimal compounds is, is actually very limited. So the benefits you get in speed have a very harsh trade-off with the actual quality And novelty of the…”
“There's no retrieval-based search, semantic search, keyword matching search that's going to be able to give your agent that information.”
“a lot of coding agents today have very basic things like you can tell me which tool patterns I'll allow or disallow or whatever. It's like yes or no, but that puts you in a very tough spot. So just as an example, like, should my agent be able to read, you know, some confidential documents or, or let's say, should it be…”
“This is sort of the holy grail of database engineering is, why not build a single system that can do both of this? But it ends up just being a lot of compromises. And one, I think one of the first issue is that, hey, each, they say Postgres has a massive ecosystem, right? You want to be using the tools that's built for…”
“HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage and just have a single storage layer.”
“I think the industry has a sense of, hey, maybe if you overfit to like one or two customers, it's going to be really bad for you. But I think the, uh, downside overfitting is much smaller than the, the upside itself. And if you sort of try to be too ambitious and boil the ocean, it's a much bigger problem. Because you…”
“it turned out that, you know, it's easier to, to go from that bod thing that's really good at the scale and ingesting and, and super low cost and create versions in it that have the speed and features of the, you know, super easy to use, like smaller data for, uh, uh, business users thing.”
“traditionally this has been an area where both in terms of safety models don't get better by just being bigger, unlike most other areas where models do get better by being bigger. Safety has not been like that traditionally. You know, you have to train them explicitly To be safe or they won't do that. But on the flip…”
“It is clearly a different form of intelligence than people. It's some alien intelligence that is vastly different, and that difference is actually often brought out to a large degree by things like adversarial attacks and red teaming, because there are certain things that fool humans that would never fool an AI, but…”
“It's like we could kind of run experiments on the brain, observe every neuron in it, reset its state to prior states, and run counterfactuals, none of which we can do with humans, and yet we still understand neither very well. Even with that, all that ability, we still don't understand AI, you know, on some fundamental…”
“one of the most effective ways of doing capability elicitation is actually through some amount of, Of, of what you would call red teaming, right? So if a model refuses a task because it thinks it's being evaluated, but it knows how to complete that task, getting it to complete that task is arguably actually a…”
“if you parse some untrusted content and there is like a prompt injection, you know, something that's clearly trying to get the model to do a bad thing, you might be interested in knowing about that, but you don't necessarily like want your, your cloud code that you were hoping was gonna run for like the next three…”
“The problem remains then and now is regulatory because you actually can't shift the burden Of the wrong clinical diagnoses from the physician to the AI system.”
“And I think teams who can raise too much money too fast, too early, who don't have to define what the P zero is, because that's the only thing when you have scarce resources, you gotta, you gotta invest in. Those cultures end up being the most fragile and brittle, and they almost don't even make it to take off.”
“We're not compute constraint in the materials industry. We're experiment constraint.”
“there's only one type of agent, and that is a coding agent. It can do it all, right?”
“So I feel like this always ends up being a tool call, a hardness issue. Then, you know, an actual model issue.”
“if you run any coding agent with permissions on, The models are actually number. And if you run them without, you know, the complete bypass of permissions, they do much better. Even if you like sit through those yes, yes, yes, accept or whatnot, you will see that, you know, the model ends up getting steered in the…”
“feels like you can fix 90% of a design slop, which is not a capability gap. It's more like a contract gap in what your harness is telling an LLM to do versus what your user is saying.”
“even when you're not at a hundred, I think a lot of these evals have a lot of problems in them. So, like, actually, it's, like, if you get to, like, 92 or something like that, many of them, it's, like, then there's, like, there's no, really no difference between 92 and 93, because the, the eval itself is, Problematic…”
“my hypothesis is that like deep down, they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other, um, then like, Basically the context fills up with them rather than the…”
“if you have more structured and formal data, it's going to be a lot more horizontal than the specific vertical we are tackling.”
“it is not about, like, formal verification or verified AI to us. It's not just about handling or, like, kicking out the lousiness, the hallucinations, the mistakes. It's about scaling brilliance. It's about super intelligence.”
“I think that we found, um, scaling inference to have almost no wall, um, recursively decomposing, uh, you know, approved goal into many sub goals and then learning to backtrack as well.”
“if you want proof to be informal math, It's, it's very annoying, because then that's, like, just makes objective function. Um, your code is something like Python, your proof is, say, natural language, math proof. Um, you will not have very strong RL kind of performance, right? But if you have proof as Lean, and you…”
“coding has worked so well that we now have to rebuild the IDE, right? I mean, it's kind of nuts to see what we launched is like, oh my God, I have these hundred agent sessions. I, the cognitive load, it transfers back to me as a human is so excessive that now I need a new UI. Uh, oh, by the way, I like the, the chat as…”
“So that's why I would say every company having private evals may be the biggest IP, right? I think about it like what's that private eval that you can then use even a frontier model to hill climb on and not leak the traces. Maybe one of the biggest drivers Uh, of IP.”
“So I think the challenge of the SAS business model is we packaged one way. We now have to learn how to unbundle these things and rebundle in new ways and discover new business models, right?”
“most people love outcomes until they have an outcome. Because once you have an outcome, it's like giving away royalty, right? I mean, I've talked to customers who love, you know, outcome-based pricing, and I say, I'm all in until they, oh my God, like, what are you talking about? You're sharing in my outcome? No, no,…”
“I think the reason why there's not a single answer is ultimately we're trying to codify trust. We're trying to say like, okay, if Sean reviews this, I'm going to trust it because you're Sean or you're the senior dev or you're the whatever. And right now when we are working in a flow where an agent writes code and…”
“And often I find that this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the, in the model training pipeline. Those gave, those gave the biggest boost to the model quality.”
“So surprisingly video models is like the, the cost is very, is comparable to language models. And obviously the largest scale is language model. Maybe like a medium scale language models.”
“The, uh, the, the visual intelligence are actually mostly coming from language. Like, these video models, especially from now, since the diffusion model technology is more mature, the, like, every time you see there, there's some improvement on these models, I would say mostly, uh, this again comes from language model,…”
“I don't want to compete for like 20 dollars a seat. I think that that is just a really difficult business. I think it's very easy to copy the main pieces of it. I mean, again, like I built this fairly quickly, and I think because you are not owning, I guess, the entire stack, it's hard to monetize. You have money being…”
“The thing we found is, so like MCPs, um, obviously it has been like, it's like really big explosion of, oh, you can go like integrate it with all these different things. Uh, but to actually get the integration right and get the right experience, oftentimes we found that we had to go build our own ad hoc things.”
“We've actually given Devon an MCP so they can just go arbitrarily message other Devons and create new Devons, et cetera. But I guess like it somehow creates like a really chaotic world in that sense. And so we, we've still found that most practical use on a day-to-day basis has been one single. Figure out how to…”
“the meme that I have is that your code base regresses to your worst engineer because that engineer who is, you know, very gung ho about AI and is not auditing their code, their pattern starts cementing into the code. And now the AI is referencing their patterns, and so now their if-else block that, you know, is 20…”
“what is, you know, what is the cure to disease look like, right? It's not, it's not a pill, right? It's not a medicine in the conventional sense. You know, it's, it's going to have to be, um, a system that is capable of modeling and understanding, you know, the underlying physiology of disease in a way that's…”
“Most people thought it was the same infrastructure for humans and agents. We understood a quarter ago. It's not, we just didn't know what was the right primitive.”
“agents will be like humans in the sense of you don't want your laptop to be shut down until you're done with work. Like, and you want to close the lid and open the lid. It's the same state. So agents would want that like pause and come back. They want those two things, but also agents really, really want speed, right?”
“What we saw from our customers was that they were all trying to figure out how to do, uh, versioning. Everyone is doing it in different ways. There were some really weird ways where people were doing that. And the reason was that GitHub as is Was an overhead. Like it wasn't fast enough what they needed. It didn't solve…”
“things that were prohibitively annoying to humans are not actually prohibitively annoying to agents. They're really, really nice, right? And so, if, if I wanted to hand you a CLI and I said, hey, guess what? The CLI has 40 arguments and. Right. 600 flags. You'd be like, wow, that's crazy. Like, I'm never gonna use all…”
“Temporal was always like really, really, really great in theory. And it was really, really great when you got it working the way that you wanted to in production. It's just, it required you to like model that entire journey in your head. And if you didn't have the entire journey in your head, you could put yourself in…”
“I think you can move towards having pets so long as, uh, you have a, and this is gonna be a jump, so long as you have a cloning machine for your pets. If you can snapshot every single thing at every frame, then, like, It actually doesn't matter if, you know, that thing gets obliterated because you have some sort of,…”
“And this is why I was saying like two is the worst number of co-founders is because you have no tie break, right? You, you basically are like, well, I disagree on this thing and I disagree on this thing, right? I was like, well, how do you resolve that, right?”
“when you're in a situation where you're in a forest, uh, in front of a wolf, you know, you first gonna deal with the wolf that wants to eat you, and then you're gonna go consult Greenpeace. So that's kind of situation that Ukraine is in.”
“such full autonomy increases the capabilities of an FPV drone, which is already, like, three orders more powerful than an artillery shell. Full autonomy increases its capabilities by four orders of magnitude, because now you can have a hundred times as many people who can use it, because you don't need to train those…”
“It's, it's no longer about an aircraft carrier that costs whatever, fourteen billion dollars, and takes forever to build. It's, it's about mass, ah, that is, you can iterate on very quickly, you can upgrade it, everyone can operate it, and then that mass, when it is combined, or the technologies, when they're, you…”
“if the harness is good enough, I would say, of course, it depends on the industry, but on average, you can get to 80, 90% of, uh, resolutions for, like, customer product issues, but to get there, it's the job of the harness”
“people keep saying like the cost of switching between models is zero. It's cheap, but it's not zero because sometimes you like spend, you know, three, four months, like fine tuning exactly how the model should be, or like how the, the, Product should be. And then once you like change the model, it, it, it, it's bad.”
“And I think in that you, you start to see more of a bimodal distribution of engineers, right? You, you start to see like, wow, there's, there's this subset of, of people that they, they really get it. Like they're, They're all in and they've, they've clearly invested the, the hours needed to learn these tools and, and…”
“in the physical AI world, we're not really constrained right now by like the intelligence of the models. It's actually what Peter's talking about is actually deploying them. ... On the hardware you give you. And so, and there's just a reality is of safety critical systems. So those end up being your, your, your…”
“At PR review time, you want to run the largest models. That means codex or, or cloud code is not going to cut it. You need to have pro-level models if you really want to stem the tide of bugs from going into production”
“it actually, in terms of the overall time to deploy, it's total time savings if you spend more time on a longer model, like, thinking for an hour, because then, then you, you don't have to spend all that time During testing and rolling, you know, rolling back the deployment.”
“I think one thing that's becoming more clear is I think like, like coding agents are the kernel VGI, sort of everything is a coding agent.”
“every software engineer in Notion this summer went through like this, um, sheer, um, one of our engineering leads at the company called it like, every software engineer is going through the, the identity crisis that every manager goes through, where all of a sudden they realize their ability to write code is less…”
“There's a critical difference To being a manager, which is that like, it is actually very deeply technical. The problem of, you know, humans are very like, like, like fuzzy and you can't like treat a team of humans like a, like a rigorous system where like, you know, PRs like, like flew through and can be in like a…”
“I'd actually say we don't try and make it as easy as possible to use because the more we do that, the more we abstract away that interpretability that Simon's talking about that basically nerfs the model or nerfs the agent from being super capable.”
“And actually, 99% of the time, it's a bug in one of the tools. Right. And so, just fix the bug.”
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have the model, the harness be the whole box and give it a bunch…”
“if you can figure out how to collapse a product that you're trying to build a user journey that you're trying to solve into code, it's pretty natural to use the codex harness to solve that problem for you.”
“The level of complexity of the dependencies that we can internalize is I would say low medium right now, right? Just based on model capability. What is medium? I would say like a, a couple thousand line dependency is a thing that we could in house no problem in an afternoon of time.”
“The structure of the repository is like, 500 NPM packages. It's like architecture to the access for what you would consider, I think, normal for a seven person team. But if every person is actually, like, 10 to 50. Then the, like, numbers on, like, being super, super deep into decomposition and sharding and, like,…”
“no humans in the loop here. So like my, Own personal ability to write or not write Elixir doesn't really have to bias us away from using the right tool for the job, which is just wild.”
“if we want to actually, like, make it see the layout, it's almost easier to rasterize that image to ASCII arc and feed it in to the agent.”
“If we were building an entire separate Ross scaffold around Codex to restrict its output, that I think would be like additional harness that would be prone to being scrapped. But yeah, if instead we can build all the guardrails in a way that's just native to the output that Codex is already producing, which is code, I…”
“So it turns out what we now know is an agent is the following. It's, it's a language model. And then above that, it's a bash. It's a bash shell. So it's a Unix shell. And then the agent has access to the shell and, you know, hopefully in a sandbox, maybe in a sandbox. So it's the model. It's the shell. And then it's a…”
“The, the, the bots can pass the Turing test. And if the bots can pass the Turing test, then you can't, you can't screen for bot. You can't have proof of not a bot, but what you can have is you can have proof of human.”
“Both the AI utopians and the AI doomers are far too optimistic. You see what I'm saying? Because they believe that because the technology makes something possible, that eight billion people all of a sudden are going to change how they behave, and it's just like, nope. So much of how the existing economy works, it's…”
“on our way to, let's call it embodied general intelligence, Models need to learn the consequences behind their actions, which means that they need interactive data.”
“believing that there can be a really rich connection between a more symbolic layer of abstracted understanding of visual domains, which aren't in the mainstream vision models, which are still trying to operate on the surface level of pixels.”
“you only actually have a world model if you can predict, given some action is taken, what is going to change in the world because of that, and in particular that becomes hard over longer time scales, so if you're simply, you know, trying to predict the next video frame, that's not so difficult, but what you actually…”
“if there are ways in which you can work with five orders of magnitude, less data than people working purely from pixels, you're going to be able to make a lot more progress, a lot more quickly, and that's the bet here.”
“That's why we can actually use models audio, but also like OCRs that are like really, really good at that, and that will be much more cost effective than a, than a, than a general model. That will contain a lot of capabilities you don't really need.”
“For instance, for, for audio, for audio here, if you want to, to do transcription, I think it makes no sense to use a model as this large. If you just want to transcribe tech, it would be very inefficient. Like, uh, if you want to do audio, you probably just want to do the one B or a three D model. Performance would be…”
“GRPO, for instance, it doesn't really work with any bit of policy, which was okay initially, because you are solving math problems that can be solved in like a few thousand tokens, so the model can actually generate them pretty quickly, so when you do your update, the model is never too far off. It's never too far off,…”
“most people who actually work on Getting materials to the device scale, say something that would, um, be in your television or something like that, is they will tell you that it's not just the material, it's the process. And I think we're at, we're at ground zero. We're nowhere when it comes to, like, well, how do we…”
“one of the funniest things I think we noticed is that you can get the temperature at which a material will break down two ways. One, you can get it from the graph, and two, you can get it from what the authors say about how they interpret the graph. And those two things do not line up.”
“TypeScript is an amazing language for AI because there's tons of training data. In the models, um, and it's strongly typed. And actually at the company, we built most of the stack in TypeScript, and we have this amazing property, which is we have type safety all the way from the database to the front end. And there's…”
“very early on, we were putting lots of facts into a vector database and, uh, and doing embeddings and pulling them back out, um, using, you know, reverse look of embeddings. Rag that actually worked, but turned out to be much more complexity than was actually required. So, you know, today we've replaced it with a…”
“It turns out being an engineering manager, as long as you stay very close to the code and are able to continue to craft it yourself, is actually a great skill profile for being able to make agents work for you and for your team in this, uh, in this age.”
“I think that entity needs to have access to all the same tools you have access to. Otherwise it's going to be hamstrung, like all these complex ways.”
“I don't think we need to wait for like a hundred percent model alignment. We can rely on the same Swiss cheese model we've used in the industry for a long time, but I do think we need to like universally, maybe eventually invest more. And that's what we're doing. We need to invest more in systems where we can say you…”
“My take is that there are four, four things where if you can, if you can cover all four of these things, LLMs are not like three X faster or five X faster. They're like a hundred X faster.”
“we can take all of the world's knowledge, all of the exabytes and exabytes of data that there is, and we can use those tokens to train a model, but we can't compress all of that into a few terabytes of weights, right? We can compress into a few terabytes of weights, how to reason with the world, how to make sense of…”
“the only real downside to that is that if you go all in on object storage, every write will take a couple hundred milliseconds of latency, but from there, it's really all upside, right? You do the first query, it takes half a second”
“the way OpenClaw has with QMD or with whatever memory plugin you use, it inherently relies on tools to search through these memory.md files that it prepares. So, you know, like, what did I decide about the API? Then agent will decide to search, and sometimes it won't decide to search, which is probably the biggest…”
“agents can do three things. They can access your files, they can access the internet, and then now they can write custom code, uh, and execute it. And you really only let an agent do two of those three things. If you can access your files and you can write custom code, you don't want internet access because that's one…”
“I think that in pre-training, there's just an enormous amount of command line data. Like even let's ignore, let's like, let's ignore RL. Like you're doing no harness, you're doing no harness post training. Just the amount of like CLI versus API documentation for just like navigating this world of the CLI in your file…”
“the way the top one percent build is so different from the bottom 99%. You have all your, like, indie hackers and, like, open source frameworks that are very, very popular. And, like, everyone uses, and that's what you see in public, and then you go see how real people are building stuff and getting reliability, and…”
“What's happening is we are changing our work to make the agents effective in that model. The agent didn't really adapt to how we work. We basically adapted to how the agent works. All of the economy has to go through that exact same evolution. The rest of the economy is going to have to update its workflows to make…”
“if a, if a really, really smart human could not do that task in five or 10 minutes for a search retrieval type task, you know, your agent's not gonna be able to do it any better.”
“A few of the insights is, like, everyone, frontier model is not good at search. Humans have this natural explore-exploit trade-off, where we kind of understand, like, when to stop doing something. Also, humans are pretty good at, like, forgetting, actually, like, pruning their own context, whereas agents are not. And…”
“The suggestion in this, in this paper is that if you think that algorithmic progress, you know, that, that is coming up with the transformer, coming up with RLHF, you know, MOEs, all of, all of this stuff, better learning rate schedules is, is, um, uh, is itself a function of compute because, you know, you, you need to…”
“Within model generation, it's valuable, and across model generations, it's not so valuable.”
“Google, Amazon, they get to vertically integrate, integrate, and vertical integration always saves tons of money. Um, so he has to be better than everyone. Um, by, not just like a little bit, by, by a ton. To justify his margins. Otherwise, the vertical integration of, of his competitors will win out.”
“I think moats are as shallow as they've ever been, right? Uh, because how fast things are moving, the, the, the size of the numbers that are being thrown around now, right? It's hundreds of billions of dollars for each, uh, major hyperscaler. The size of the numbers are so large that you can just go and justify hiring…”
“I think coding is a little harder, if I'm being honest with you, and you're telling me the hard one got automated. Why can't the easy one get automated? So I started to ask myself, how much can we do? And the answer is, it feels like a skill issue.”
“if you didn't pay any like human cognition to get there, I don't think you're going to be a great reviewer. One of the reasons why, you know, what, what makes that, that human feet, that loop well is because once upon a time you did that and you could make the three like, yeah, yeah, idiot. You're not thinking about…”
“Another thing that is very different this time than in the history of computer science is, is in the past, if you raised money, then you basically had to wait for engineering to catch up, which famously doesn't scale. Like, the mythical man must take a very long time, but like, that's not the case here. Like, A model…”
“when it comes to hardware, um, most companies will end up verticalizing. Like, if you're If you're investing in a robot company for an app, for agriculture, you're investing in an ag company, because that's the competition and that's the pricing and that's the supply chain. And if you're doing it for mining, that's…”
“No, no, no, a billion, a billion dollar training run, a one billion dollar training run, it makes sense to actually do a custom ASIC if you can do it in time. The question now is timeline. Not money. Cause just, just, just rough math. If it's a billion-dollar training run, then the inference for that model has to be…”
“And I think one of the conclusions is, is like there's no such thing as a coding model. You know, like, that's not a thing. Like, you're talking to another human being, and it's, it's good at coding, but like, it's got to be good at everything.”
“I think, um, what I've been calling agent labs, which are people who build on top of all the other models. We'll probably have a better time with the margins because they, they price against the end user hours spent or like human labor. Whereas models get commodity price per token. And so margin wise, we know inference…”
“I mean, I think there's still a, there's also sort of the more exotic things like analog based, uh, uh, computing substrates as opposed to digital ones. Uh, I'm, you know, I think those are super interesting cause they can be potentially low power. Uh, but I think you often end up Wanting to interface that with digital…”
“Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state, and that I don't think we've made that much progress on, but the idea of, like, yeah, going straight to the answer, we've become pretty good at.”
“if there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
“he actually has a paper that, as well as some, you know, others from the, from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience Bayesian framework, but basically how, you can…”
“And so where the unstructured data is such a large part of it and conversation data is such a large part, most likely the context graph will actually consist of small models. The context graph itself is a model layer. Like you use that data to train a set of models, but then become part of your context graph.”
“We don't actually have to hold their hands so much anymore, or like they don't actually need to necessarily have an automated lab. They can like write an email to a CRO or something, or they can like tell you what experiment to do. And you can take a video of you doing it and then show it to the model and they can be…”
“Basically the, the best model, um, whatever, Opus seven or GPT 10, like, um, it really can only propose the first experiment, maybe slightly more clever, but at a certain point you just need information, right? Like some little calculations you can do that, like, there's more atoms in the brain than you could ever…”
“we learned a lot about how bad our LHF is with people, just like people pay really attention to, to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what people didn't really pay attention to is like, I don't…”
“I have a lot more faith in these, like, um, verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test, or whatever, you're going and running the experiment. Anything like that, I think, is going to give you a higher signal than the sort of vagaries of…”
“in cosmos, we basically, we had all the pieces sitting around. We working on world models, we working on data analysis agent, working on literature agent. Um, and then we're working on, you know, we built a platform for scientific agents. So we had things that can write a law tech report. We had things that can make…”
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the tokens. I think it's more about the learning algorithm itself.”
“the bitter lesson gets used too much in, like, too conveniently used around, but actually there's also a little bit of a, not a bit, there's also a sweet lesson where it's like, ideas matter.”
“I think of like the MCP and tool usage as being like the interface to all of our conventional, uh, imperative systems, not at the AI space.”
“I view agentic development as being something that amplifies all the, all the good, just as much as it amplifies all the bad. And it amplifies, uh, uh, sloppiness, poor architectural thinking, um, Uh, misunderstanding of, of the requirements. Like there are, for all of the, the acceleration of good outcomes, it also…”
“my intuition has been that trying to craft LLMs into deterministic workflows and DAGs is, is kind of underselling like the power that they have to actually plan and execute more in a more sophisticated, like fluid way.”
“once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the next couple of years.”
“I think where we're getting to is that these models have gotten smart enough, they've gotten better, better tools that they can perform better when just given a minimalist set of tools and, and let them run, let the model Control the, the agentic workflow rather than using another framework that's a bit more built out…”
“if you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's, it's the equivalent of using GPT-V as a judge, but it's,…”
“I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective. This objective doesn't actually use rewards in it, and so there's another word in the title reinforcement learning that also might be a…”
“I think it's because we're fundamentally shifting the burden of learning from something like, like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're trying to classify whether future states Is along the…”
“really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
“as you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know, you find the answer to a math problem, it's a lot less…”
“one of the pitfalls of academia is that, um, it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas. Uh, those mathier ideas also give you these, like, kind of implicit knobs to tune, uh, that allow you to, like, overfit. While, you know, the, the things…”
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. Um, it generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like it kind of, it can, it can kill the training…”
“a big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, see it. And then you use RL on top of that.”
“I think human feedback is kind of like a, a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you, you take the model, and you, like, elicit it to be a little bit better, uh, in terms of personality”
“Yeah, so you, all you do is leave IntelliJ running, but you shouldn't look in it. It's a tool for the AI now, right?”
“as soon as you get to the point where like every developer is 10 times as productive, merging their code becomes this incredibly complicated problem because I, you and I work at the same time for two or three hours. We make, you know, 30,000 line change each. Mine makes it in first, ah, and it gets merged and then you…”
“This is one of the coolest things about, like, model training is literally, like, they develop habits. It's just like a person does. Like, if, if you're, like, working on some podcasting tool, right, you're really good at editing, and then somebody makes you use a different one, it's gonna slow you down, you're gonna…”
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. You need to go to this RLHF domain that human really just…”
“I do believe that AI will do only one thing. It will Separate faster the good engineers from the bad engineers. If you're a good engineer, and you're using AI well, you will be an amazing engineer. If you're a poor, lazy engineer, and you don't want to understand things that you're doing, AI will make you even worse”
“in the world where the technical moat is not that a moat anymore, because Like startups in two weeks, they can build something that is close to what you're building. The difference is like the, how you think about the user or the flow and all of that.”
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's, it's good marketing.”
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they, do they agree and kind of iterate On them over time.”
“in many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
“I don't think you can value it unless you actually model it yourself and see what the capabilities are.”
“I would recommend, if you're gonna do large data deals, like, just try to get like a large chunk of equity in the company that you're doing it with, um, if you can.”
“actually making every single clip on the platform playable, um, at billions of clips scale is how we go from imitation learning to RL.”
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose, you lose the line breaks, you lose sort of the two D arrangement on the page. Um, and for a lot of cases, for a lot of things, maybe that doesn't matter.…”
“I wouldn't be surprised that given enough celestial movement data, an LLM would actually predict pretty accurate movement trajectories... I wouldn't be surprised, but F equals MA or, or, you know, action equals reaction. That's just a whole different abstraction level. That's beyond just the today's LL.”
“And also you're perversely incentivized because you leveraging AI in your work as you operating faster, but your incentive, just like a lawyer or just like any hour, uh, pay based knowledge worker is to rack up as many hours as possible.”
“Even if you have some equity, even if you deeply care about the mission, you're not deeply incentivized day in and day out to try new AI tools and push yourself to work better and faster and smarter.”
“Selling productivity tools to enterprises are challenging because as no matter what ROI argument you make, people aren't actually buying tools for ROI. People buy productivity tools because their users like using them.”
“furthering the intelligence of models and chat GPT, a consumer product does not lead to more users or more retention. It only is really applicable to us the thin slice of users who care about very smart type queries, right?”
“But the, the interesting about Anthropic is if you look at coding, that's probably never going to be the case. Like there's always an increasing frontier of how you good you could be at a task like that. And we're nowhere close to that frontier. So it's more possible to underwrite the quality of the future models…”
“It is far easier for Anthropic to try to go into one of the spaces of the apps than an app to try to go into the space of Anthropic, which makes me feel like one is more defensible than the other, um, all else equal.”
“I think go so far as to say for in any SaaS market, if there can be a PLG motion, the PLG motion will win.”
“Left to right reasoning for code doesn't actually really make sense, because in code, we don't, like, we might sometimes write code left to right, but after you write code, you go up and down and figure out, hey, is this variable set? Did I do this? There are many bi-directional dependencies in code, so there's a…”
“what was really, um, liberating for us was actually the focus on just one stack, or one framework, whatever else was trying to do, General purpose coding agent. We were like, no, we're just going to focus on Next.js front-end and Chad CNN. And that was really, that allowed the team to focus.”
“I think one of the things we've seen is the scaffolds get simpler and simpler over time as the models get better. And in some ways, this, the scaffolding is almost a crutch for, for, uh, things the, the, the model struggles with. For example, like, you know, really complicated sub-agent systems.”
“part of what we're trying to unlock here with Biohub is the idea of what actually, what happens If you do frontier biology and frontier AI in sync together and you're designing the tools on the frontier biology side in order to specifically collect and be able to learn types of data that you then want to feed into…”
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates kernels that no one has to learn CUDA again. There's always a…”
“And I do suggest that people not try and use the model to offload thinking. It should enhance your thinking or else like you'll, you'll get worse over time and you won't know when the model is quality, whether it's outputting quality or not. If you don't know if it's outputting quality, then what do you like? You're…”
“I don't really care if someone knows CUDA kernel writing. If someone does string theory and is really good at understanding complex problems, they will be productive faster than someone who knows CUDA and doesn't really care about trying to get better at it.”
“the biggest cost to very high-income people consuming is not the dollar cost, it's the time cost of consumption, and so I'm very interested in how agentic commerce can Open the aperture for spending by high-income people because it's removing the most costly or binding constraint, which was their time. And we're…”
“When I think about the optimal adoption of these AI tools, it seems wrong to focus on in-year ROI, and it seems right to focus on two-year, three-year ROI. Now, inherently hard to know what two or three years is gonna look like, but if we look at the history, model's getting much better, much more quickly. Model's…”
“the bet here that we made was that GUI based computer use was going to be much slower to come into its own. Than terminal based use. And there's a few reasons for this. I mean, as I mentioned earlier, text is just the modality that works best with these models. And so having a text based interface that allows the model…”
“And if you look at the difference between the same model and the biggest spread across agent frameworks versus the same agent framework and the biggest spread across models, It's much larger across models, which I think goes to show that like improving models and model selection is going to be your, your biggest lever…”
“I think that there is today, like. 10 times as much AI inference that could exist than is existing right now, just Purely with projects that are like sitting in the proof of concept stage and have not been deployed because there's like huge bucket of those. And it's, it's all about this kind of like reliability issue”
“the autoregressive nature of language models means that if they make a mistake, and you correct it, and then say, no, that was a mistake, please do it this way instead. The more often you do that, the worse the dialogue answers get. Because it's in the training data, when it's seen some mistakes, it tends to be…”
“That really is like, to me, the number one thing I've learned from this and all forms of vibe coding and all the things that I've, I've played with is that it's just extremely easy to get into this kind of like soporific state where the AI is doing the thinking for you. And you start to think that all the information…”
“the nice thing is that with the dialogue engineering we discussed, the longer your dialogue is, the better the AI gets, which is the opposite to what we're used to, right? Because you can edit the outputs that aren't great.”
“It feels to me that, um, the better code generation gets, the more design matters, and the more that actually the human pushing on design matters too, because even if you have a good starting spot from your design system, from, you know, AI generation, whether it be code or image, you, I think, need to push design…”
“that's not like a comment on, uh, oh, front-end engineering is dead. There's so much more to front-end engineering. Than just that translation step. That's kind of like the most mechanical part, and actually thinking through all the states, all the intended behavior, um, and how to make the design truly come to life.…”
“And people are now also waking up to the fact that it's not just the model, it's the system prompt, it's the tools, it's the harness. Um, the scaffolding around the model. So I can give you the choice to use Gemini 2.5 in, in AMP, but without the system prompting you into it without, you know, what I call before, like…”
“the failure mode of outsourcing the thinking, but not the typing, which I think it should be the opposite. You still have to know engineering. You still have to know how to program. You still have to know your application and its architecture, how it's deployed. And then basically use the agent to do the work that you…”
“if I give it this other custom-made MCP that we built internally, and our processes don't map to anything that OpenAI and Anthropic have seen or trained for, it won't be used, and you won't get good results.”
“What we're seeing now with agents is as soon as somebody has seen what it can do, they have such a multiplying effect or this brings so much value that people are willing to adopt the code base for this. Like the first time in how many decades where people are like, maybe our code base is wrong. Like maybe, maybe we…”
“people assume that coding agents need to meet the bar of writing the exact same kinds of software to the exact same standard, and that is not necessarily an assumption that end users, consumers will apply. If they have software that's much faster, cheaper, much more personalized, if they can conjure it up on their own,…”
“I could see a path to AI solving the integrations problem. I really couldn't see a path to AI solving this sort of messy data problem because it's more of a people process problem than it is like an AI software problem.”
“this sort of Palantir model really only works if you not, you do a lot of like, you have very low or even negative gross margins in the first year, but then once it's deployed into production, you take the head count off, the margins go up and you have this nice high gross margin recurring revenue stream ongoing. But…”
“And so I think that is what enables us to ultimately see a path to actually pretty high gross margins. Whereas the, the CLI based and the sort of engineering focused ones, I think you're in a tough position because everyone always wants the best model.”
“We should be adding structure necessary to get things to work today, but keeping an eye on improving models and keep, but keeping a close eye on models, improving rapidly and removing structure in order to un-bottleneck ourselves.”
“basically, you know, when you get to enough scale, inductive biases matter not at all. All that really matters is the learned posterior from the data distribution, and that's really what defines everything.”
“most of what we do in post-training is better were done in pre and mid training and earlier on in training in general.”
“if you filter the data after each point, that's now information injection, and that can break all of this, um, and I think can prevent model collapse.”
“because the model that's doing the rephrasing just needs to know how to rephrase. It doesn't need to know anything about the content itself. It doesn't need to understand it. It means you can use a pretty weak model. To do the rephrasing and have it generalize and generate data that can teach a model that's much better…”
“if there's only one thing that you should take away from this entire interview about what is good for, for data quality, it's diversity.”
“Repeating higher quality tokens is almost always better than seeing net new lower quality tokens. So, like, epoching over higher quality data, almost always better than getting the same amount of new data of an unknown quality, or of average quality, average in this case being like what you just get from an internet…”
“we actually found out that the problem was that the lottery ticket was actually data dependent. And that was where the fundamental problem came. That as soon as you change the data distribution a little bit, like the winning tickets changed in a really big way.”
“The performance of LLMs is not invariant to how many tokens you use. As you use more and more tokens, the model can pay attention to less, and then also can reason sort of less effectively.”
“everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, real bottlenecks.”
“If you are requiring a user or having yourself as the company needing to actually recompile a workload, that's already one step too far, even if you assume it works perfectly.”
“Another thing at a very practical level that we've thought about with open source models is that people building on our open source model are kind of building on our tech stack, right? If you are relying on us to help improve the model that you're relying on us to get the next breakthrough, then that means that you…”
“the productivity impacts of people being able to do more means we actually want more people, right? It's like we are so limited by the ability to produce software, so limited by the ability of our team to actually clean up tech debt and go and refactor things, and if we have tools that make that 10 x easier, we're…”
“the other thing I always think about from a finance point of view that, Maybe people don't even think, don't think about that closely, but obviously push back if you, if you disagree, is that it's basically a proxy for the gap between open models and closed models. Because if the open models do better, um, sort of the…”
“there's a sense that whenever you join a startup, like, you're gonna be compensated for the risk that you're taking on of joining, like, something that's unproven, and then now you're having these situations where, like, yeah, some people do get rewarded for taking on that risk, but then other people who have been…”
“any kind of fancy optimization you're doing on top of the LM eventually gets, like, you kind of lose the need to do it given how fast things are moving.”
“We don't think a product is good because it's open source. Usually a product is bad because it's open source. It tends to be like designed by committee and like just a smash of like a bunch of different competing priorities.”
“Open code itself will never be on monetize. The day you monetize your open source project is a day it starts to die.”
“in the same way that chatbot arena can never be saturated. RLHF can never be solved.”
“I definitely don't think the algorithm tends to be the most important thing.”
“it's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it stops or it gets it on the 81st. Is just the RL behavior that…”
“the model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don't write down our goals of the model in a constitution form.”
“talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on the top people.”
“I think the domain experts, like in our case, clinicians, they're really good at like debugging model outputs, meeting with users, distilling that feedback into something actionable, maybe annotating or doing evals, but they don't necessarily have like, you know, the right intuitions about what techniques to even try…”
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right?”
“the open source serving. Offerings are just, I will say not great in that, in that they aren't customized to transformers and these kinds of workloads where I have high latency and I want to like batch requests and I want to batch requests while keeping latency low. But one of the weird things about generation models…”
“One of the issues that ends up coming up with things like human eval is contamination, because a lot of these things that train models end up training on all of GitHub. GitHub itself has human eval. So they end up Training on that, and then the numbers are arbitrarily a lot higher.”
“And so at Codium, we believe that this embedding based retrieval is the heuristic. We should be planning for AI first products, throwing large models at these, at these problems so that AI is a first class citizen.”
“One of the cool kind of features of Devon, I'd say, is, you know, if, as an engineer, you're, you're working on, let's say, four different tasks today, you know, you just give one to Devon number one, you give the second one to Devon number two, the third one to Devon number three, you have four Devons that are all…”
“once it starts hitting the peak of these benchmarks, getting that last 10% actually probably is, like, counterintuitive to the actual goal of what the benchmark was. Like, you probably should find a new hill to climb, rather than sort of p-hacking or really optimizing for how you can get higher on the benchmark.”
“For a lot of the systems, we do believe embeddings work, but for complex questions, We don't believe embeddings can encapsulate all the granularity of a particular query. Like imagine, imagine I have a question on a code base of find me all quadratic time algorithms in this code base. Do we genuinely believe the…”
“I think that right now optimizing for making money off of individual developers is probably the wrong, actually, strategy. Largely because I think individual developers can switch off of products, like, very quickly, and unless we have, like, a very large lead trying to optimize for making a lot of profit off of…”
“Thinking about what's going on in AI as really replacing search is the wrong take, or maybe not that interesting. It's really more like AI is replacing web browsing. I think that's actually like a more meaningful, like more fundamental shift.”
“In chat search people, people don't really do that. People just say, solve the problem for me. What I think is really interesting about that is, like, it's actually the highest intent. Like, it's the most valuable thing you could be doing. Because you're going from being, like, I think I need to solve this problem, and…”
“And the problem with that is that we don't want to incentivize AI to derive the program that made the game. Right? And so if we continue to have humans make the game, then the AI is incentivized to try to reverse engineer the G inside of humans, and that's kind of the whole point of what we're trying to do here.”
“Because our hypothesis and our definition of AGI is as long as we can come up with problems that humans can do and AI cannot, then we do not have AGI. And then the flip side of that is also true, which is When us as ArcPrize, we're like, we consider ourselves, like, our job is to come up with problems that humans can…”
“The counterexample that I always have that I think is quite illustrative is in German, the verb is at the end of the sentence. So if you're trying to do real-time translation from German to English, as an example, you can't actually make any progress on the English until you hear the whole German sentence and you know…”
“So you should actually build your products ahead of where costs are so that by the time they are popular, actually your frontier. So it's actually, I think the efficiency thing trips up a lot of people in terms of Being cost and cost conscious and wasting a lot of time on that when actually the, the most cost…”
“I think, like, all of the things that I would consider paradigm shifts in the Kuhnian sense came from a new technique, but trained on new data, and I think the new data is super, super important”
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what, what is the best research report that you could generate? And yet these models are doing extremely well at this, at this domain. So I think that's like an existence proof. That these models can succeed…”
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better job than a human would.”
“basically, when you're playing, like, the zero-sum games, like, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's, there's room for collaboration, then Game Theory Optimal actually doesn't work…”
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer. In fact, if you look at O-three, it's thinking for longer…”
“language models are next token predictors is like a fact. Like that is what they do. They, they are trained to predict the next token. However, that does not mean that they myopically only consider the next token When they choose the next token, you can work on break the next token, but still like doing so in a way…”
“what we're witnessing is developers becoming platform engineers. They're, they have to learn how to enable others to be productive. These others, of course, are AIs. And to do that, they have to give them environments to work in.”
“If the answer is no, you're not, you're, you're solving part of the problem, but you're not fully solving the problem of standardizing dev environments for coding agents. It's not, not going to, um, stand the test of time because it will, it cannot be ubiquitous. It'll be a great commercial solution. You'll make lots…”
“So for example, Sonnet 3.7 clearly has, uh, it smells like cloud code, right? Same with codex. Uh, it very much, uh, impacted the way that those models want to write and edit code such that they seem to have a personality that wants to be in a CLI based tool.”
“Like if you do have billion token context window model, you throw it all in there. It's still gonna be more expensive. The reason why retrieval is so important for us is because even if there is a model that's going to have these larger context windows, and certainly over time, we're going to get larger context…”
“because if you look at a lot of, like, Sweebench passing, like, outputs from, like, an agent, they're not really, like, PRs that you would merge, because, like, the code style might be, like, different. Like, it works, but the code style is different.”
“if you can, like, build, do something very specific for, like, a specific purpose, actually, when you bring that and you bring it into the, into the generalized model, like, you might even get outsized returns on that. Because there's, like, transfer from all these different domains.”
“the existing paper people are writing about multi-turn RL are not actually incorporating this, and it kind of, like, breaks all the math.”
“you can have a model trained on code, and a model trained on math, and a model trained on Spanish, and you can literally average the weights, and it works.”
“The updates made to model weights are orthogonal enough for specialized tasks that this is actually like totally fine. Things are nice and linear in most cases, things are nice and orthogonal, and you can get away with a lot of, uh, async, uh, updates to models that are then merged even without full communication.”
“before I would write a big design doc, and I would think about a problem for a long time before I would build it sometimes for some set of problems, and now I'll just ask quad code to prototype, like, three versions of it, And I'll try the feature and see which one I like better. And then that informs me much better…”
“our thesis on consumer software now is, it's gonna be It's going to reflect less or look less like our operating consumer technology companies is going to look less like traditional tech companies and more like CPG companies that were like very much distribution first, always branding like at the forefront because…”
“For us, we've actually like figured that the like unit, the like The return on it was not actually that beneficial. What makes sense for us instead is, you know, like we use base models, um, for, for like the final, like response. What makes sense, what makes more sense for us is, can we like spend more time on that…”
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at test time, I think kind of the natural, like very big missed…”
“And I think that's another aspect of like what makes something agentic, like not having to have a user send an event to trigger the machine to turn on, just allowing these machines to run all the time.”
“the core issue issue is actually to build the knowledge graph, right? The, the entities, the relationships. So if you say graph, graph, You know, databases or graph rag is going to kill vector rag and all that discussion. I think the first issue is to actually build the knowledge graph the first place, right? And if…”
“So if you have a 10% margin increase because you have great software, um, on your billion dollars, the customers are that price sensitive. They will immediately switch off, um, if they can, because why wouldn't you? You would just take that hundred million dollars, you'd spend fifty million dollars on hiring a software…”
“So that means that the best way to make money in GPUs was to do basically exactly what CoreWeave did, um, which is go out and sign only long-term contracts, pretty much ignore the bottom end of the market completely, and then maximize your long-term contracts with customers who are, um, who don't have credit risk, who…”
“The GPU clouds are fantastic real estate businesses. If you treat them like real estate businesses, you will make a lot of money. Um, the, Cloud services you can make on that, all the software you want to make on that, you can do that fantastically. Um, if you don't own the underlying hardware, if you mix these…”
“What we found with SuiteBench is an interesting benchmark, very useful for doing some things like refining prompts, for example, and trying out different tools. Specifically for codebase understanding, in the beginning, we were hopeful That it will tell us more about how we're doing there. But looking at the actual…”
“I would rather under-engineer something than over-engineer it if I were gonna err on the side of something. And here's the reason is that when you under-engineer it, uh, yes, you take on tech debt, uh, but the interest rate is relatively known and payoff is very, very possible, right?”
“it's easy to conflate long running with CPU intensive, and these are kind of not the same thing, and actually a lot of Specifically agentic workflows. A lot of what you're doing is you're waiting. You're waiting on human in the loop. You're waiting on LLM. You're waiting on random external resources, right? Like that's…”
“if you're building a consumer app, you have to move beyond the chat box. Uh, people do not want to always type out what they want.”
“I put very little value in benchmarks, like general benchmarks. It's has some value, but you know, what you really need to do is like measure it in your domain and see if, if that, uh, alum as a judge is more aligned than like an off the shelf LLM. And what I've found is like, it's kind of, yeah, it doesn't seem like…”
“The tricky thing for RAG, it really works well because a lot of these things are doing like cosine distance, like a dot product kind of a thing, and that kind of gets challenging when your query side has multiple different attributes. Uh, the dot product doesn't really work as well. I would say, at least for me,…”
“I feel like it's still early days for us, like to try to platformatize or like try to build these, oh, there are these five horizontal pieces. And you can plug and play and build your own agent. My personal opinion is we are not there yet. In order to build a super engaging agent, I would, if I were to start thinking…”
“What we've learned is, like, doing the traditional, like, embedding and, and RAG is suboptimal. We, we kind of built our own using small models to do really massively parallel retrieval, which I think is going to be maybe more common in the future.”
“I tried R one, but R one is a bit under, uh, O one, uh, with small agents. And I think this is also a matter of formatting. Like sometimes the model struggles to just output them, the code snippets in the correct way that we expect.”
“few great things have been created by committee, you know? And so if, uh, engineering is an order taking organization for product, uh, you can sometimes make meaningful things, but rarely will you create extremely well-crafted breakthrough products. Um, those tend to be small teams who deeply understand the customer…”
“To go use that metaphor, we're sort of in the jQuery era of agents, not the React era”
“I actually think for every single agentic domain, whether it's customer service or legal or software engineering, that's essentially what the company building those agents is building is like the system through which you express the behaviors you want, esoteric and small as it might be. Anyway, I think that's a really…”
“Prompts are great, but it's not actually a complete specification for anything. It never can be.”
“It seems likely to me that, you know, at first and something that is AGI will be good in digital domains, um, you know, because it's software. So if you think about something like AI discovering a new, uh, say like pharmaceutical therapy, the barrier to that is probably less the discovery than the clinical trial. And,…”
“Certainly software engineering will be one of the disciplines most impacted. And I think that it's very, so like, I think if you're in this industry and you define yourself by the tools that you use, like how many characters you can type into Vim every day, that's probably not like a long-term stable place to be…”
“Agents, agent frameworks, graphs, all of this stuff is basically making up for the fact that right now the models are not that clever.”
“the reason that AI engineering can exist outside of the model labs is because the model labs release Models with capabilities that they don't even fully know because you never train specifically for it. It's emergent. And you can rely on basically crowdsourcing the search of that space or the behavior space to the rest…”
“sometimes I feel like a lot of researchers or, like, people in the AI community are, like, so into, like, yeah, agents, delegate everything, like, blah, blah. But, like, on the way towards that, I think, like, collaboration is actually one of the main roadblocks or milestones to get over. Because then you will learn…”
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's, it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pretty decent at like taking a long sequence of steps to…”
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by, um, I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still impossible. You basically always have to look really…”
“You know, if you want to get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is, but it's also, like, it's really important, like, to be able to move up six percent by doing something like that, like, the problems,…”
“there must exist a platform where a small team can produce an AI for a unique purpose and they can iterate and build the best thing for that.”
“You can try and do, like, a routing thing where you say, for a user, given user requests, we're going to try and predict which of these end models that users enjoy the most. That turns out to be pretty expensive and not a huge source of, of, like, edge or improvement.”
“How do you get the user an experience that is both smart and funny? Well, just 50% of the requests, you can serve them the smart model, 50% of the requests, you serve them the funny model... the eighty-twenty solution, if you just do that, you get, you get a pretty powerful effect out of the gate.”
“The models have proven to just be far worse at reasoning than people sort of thought, and I think whenever I hear people talk about LLMs as, as reasoning engines, I sort of cringe a bit. I don't think that's what they are. I think of them more as like a simulator.”
“The thing I'm just trying to share here is there's one surprising thing about humans is their preferences are pretty correlated. What you find funny and entertaining, I find funny and entertaining, and he finds funny and entertaining. There might be degrees of variation in it. I might find it super funny. You might…”
“The link prediction objective can be seen as like a neural page rank, because what you're doing is you're predicting the links people share. And so if everyone is sharing some Paul Graham essay about fundraising, then like our model is more likely to predict it. So like inherent in our training objective is this, uh, a…”
“you should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in more complicated things like process rewards models,…”
“I kind of call this the GPU smiling curve, uh, where the, the, the edges do well, cause you're either close to the machines and you're, you're like number one on the machines, or you're like close to the customers and you're number one on the customer side. And the people who are in the middle inflection, um, character…”
“And the method that we adopt in open hands instead is we provide these tools, but we provide them by just giving a coding agent the ability to call arbitrary Python code. Um, and in the arbitrary Python code, it can call these tools. We expose these tools as APIs that the model can call. And what that allows us to do…”
“for example, smaller than one was trained only on one trillion tokens, but this model is trained on 11 trillion tokens. And we saw that the performance kept improving. The models didn't really plateau me training. Which I think is really interesting. It shows that you can train such small models for very long and keep…”
“none of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like, it's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's our claim. But somehow when we jam both together,…”
“we basically built a whole library just around this basic idea, um, that all, uh, your basic compute primitives should not be a float, but it should be a matrix and everything should just be matrix compute.”
“one key advantage of, uh, this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million, uh, uh, context. Um, it is actually still stable. It's still, it's able to run. It's still be able to rationalize. It…”
“to give you a sense of, like, how I personally think about research budget, um, for each part of the, of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real estate of the art, you know, your DeepSeq minimum…”
“If you're interested in, um, you know, your, Open replication of what OpenAI's O-one is, um, you're gonna be on the 10 K spectrum of our GPUs.”
“Um, and then one disappointing learning that we found in our own portfolio is no one has the data we want. In many cases, right? So imagine you are trying to automate, um, a specific type of knowledge work, uh, and what you want is the reasoning trace, um, all of the inputs and the output decision. Um, like that sounds…”
“I think that right now optimizing for making money off of individual developers is probably the wrong, actually, strategy. Largely because I think individual developers can switch off of products, like, very quickly, and unless we have, like, a very large lead trying to optimize for making a lot of profit off of…”
“I think like, uh, if you ask me what my advice, I think you have three options. One is to focus on bolt. The other is to focus on the web container. The third is to raise one billion dollar and do them both. I'm, I'm serious. Like, I think like otherwise. You need to choose.”
“the model is like this sort of like 10 X multiplier. You're, how good the bottom line model is, huge, huge swing. And then kind of what you can do on top of that, you can squeeze out three, four X kind of more. And so that's kind of where the realm of, you know, prompt engineering and, you know, multi-agent approaches,…”
“And I think like the smarter the models are, the less you need that kind of extra scaffolding.”
“Like if you're trying to output a code in JSON, there's a lot of extra escaping that needs to be done. Um, and that actually hurts model performance across the board. Where versus like if you're in just a single XML tag, there's none of that sort of escaping that needs to happen.”
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
“the more you can put your agent on rails, one, the more reliable it's going to be, obviously, but two, it's also going to be easier to use for the user, because you can really, as a user, you get, instead of just getting this, like, big, giant, intimidating text field, and you type words in there, and you have no idea…”
“I think people fail to truly, and me included, they fail to truly internalize the bitter lesson. So for the listeners out there who don't know about it, it's basically like you just scale the model, like GPUs go brrr, it's all that matters. I think it also holds for the, the, the cognitive architecture. I used to be…”
“static benchmarks are intrinsically, to some extent, unable to measure generative model performance. And the reason is because you cannot Pre-annotate all the outputs of a generative model. You change the model. It's like the distribution of your data is changing.”
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
“The diff between that and something like Devin is all in like, sort of like the agent scaffold or the agent code, right? So that's, what's really exciting about this stuff. It shows off what you can do just from prompting and just from adding tools.”
“they actually require you to submit these things called trajectories that are proving, um, what you did and that you're not cheating. And a lot of this, a lot of this ends in controversy. A lot of this ends in people say, well, we don't want to submit at all because if we review what we're doing as intermediates, then…”
“the simplest way I think about it is like, this is sort of a rent, not buy phase. Cause you know, I wouldn't want to be, we were still so early in the maturity. I, you know, I wouldn't want to be buying like pallets of over, like of Two eighty-sixes at a five X markup when like the three 86 and four 86 and Pentium and…”
“The problem is that the challenge in deploying vector search has very little to do with vector search itself, and much more to do with the data adjacent to vector search. So, for example, if you are at Figma, the Vector search is not actually the hard problem. It is the permissions, and who has access to what design…”
“if you make assumptions about the capabilities of models, and you engineer around them, you're almost, like, guaranteed to be screwed.”
“And I think in terms of the actual prompting method to use for a particular problem, I'm, uh, I think we should all be in the minimum list kind of camp, right? You should try the minimum thing and see if it works and if it doesn't work and there's absolute reason to add something, then you add something, right? Like…”
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progress, you should just be able to talk to like a coworker. And…”
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
“obviously coding is the best application for agents because it's all the gradable. It's super important. You can make everything like API or code action, right?”
“I think making the tool good and reliable is probably like 90% of the whole agent. Once the tool is actually good, then the agent design can be much, much simpler. On the other hand, if the tool is bad, then no matter how much you put into the agent design planning or search or whatever, it's still gonna be trash.”
“the use of Python and PyTorch and everything else is just a crutch, because we humans are finite. We have finite knowledge, intelligence, and attention.”
“basically prompt injection is something that occurs when there is developer input, In the prompt, as well as user input in the prompt. So the developer instructions will say to do one thing, the user input will say to do something else. Jailbreaking is when it's just the user and the model. No developer instructions…”
“When you're using these models, if you're getting the answer you want, always, it means you're not asking them hard enough questions.”
“90% of this is not doing something new. Like, 90% of this is like doing things a million people have done before, and then a little bit of something that was new. There's a reason why we say we stand on the shoulders of giants. It's true. Almost everything that I do is something that's been done many, many times…”
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
“nothing about GPT-IV would be at all different if the field of, like the entire field of Everson machine learning disappeared. Like everything to do with Everson examples, like all of the, like for the most part, like GPT-IV would exist identically.”
“The way that we Decided we want to try to extract what happened in the past, like as forensically as possible, has been and is currently like one of the, the main things that we focus all our time on. Because doing that as we're getting as much signal out as possible and doing that as well as possible is the biggest…”
“if you're training for random weights, you better have a really good reason, you know, because it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the only situation in which I think starting from random weights…”
“So the point at which they're doing proper continued pre-training is the point at which that becomes a continuum rather than a phase. So the only difference with what I was describing last time is to say, like, oh, they should, you know, There's a, a function or whatever which is happening every batch”
“this didn't make sense to have like a so-called non-profit where then there are people working at a commercial company that's owned by or controlled nominally by the non-profit where the people in the company are being given the equivalent of stock options. Like everybody there was working there with expecting to make…”
“companies are sociopathic, like, by design. And so the alignment problem, as it relates to companies, has not been solved. Like, companies become huge, they devour their founders, they devour their communities, and they do things where even the CEOs, you know, often of big companies tell me, like, I, I wish our company…”
“Um, to explain, it's not that you shouldn't merge models, it's that you shouldn't be distributing a merged model. You should distribute it a merged adapter. Um, 99% of the time. And actually often, one of the best things happening in the model merging world is actually that often merging adapters works better. Uh, the…”
“most companies that are buying AI tooling, they want the AI to do some sort of labor for them. And that's why the picks and shovels kind of disinterest maybe comes from a little bit. Most companies do not want to buy tools to build AI. They want the AI, and they also do not want To pay a lot of money for something that…”
“And so, to be compute efficient at inference time, it's much better to train it much longer training time, even if it's an effort, an additional effort, than to have a bigger model. That's what I call, like, I refer to the chinchilla trap, Not that Chinchilla was wrong, but if you consider inference time, you need to…”
“Usually when, like, a provider, like, provisions new notes, or they would, like, uh, give us... Yeah, it's usually, like, bad, like, like, dog shit, like, like, at the start. Uh, and then, uh, it gets, like, better as you go through the process of, like, returning notes, like, and, and, you know, like, uh, like,…”
“Like, we do a lot of code generation, but we don't really do a lot on, like, code competition problems for the very, very hard ones, so that you can go very far down that route and make something like really good at those problems, but not actually that useful as, like, a programmer day-to-day.”
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. Um, so you can do some wacky things to your model, like quantize the hell out of the KV cache…”
“it's not clear that you get superhuman reasoning capabilities from human level demonstrations of skill. And by that, I mean the pre-training corpus, but then additionally, the fine tuning corpuses, I think you largely mimic the demonstrations that are present, uh, at that model training time. Um, but from a working…”
“Um, the other piece of this is that the financial modeling is often very, when we, when we talk to our users, it's very personal. So they have a specific view of how a company is structured. They have the, you know, one key driver of asset performance that they think is really, really important. Um, it's kind of like…”
“the way of how to get self-organizing computation to run on the brain that is producing representations of an agent that lives in the world is a simplification of the interests of that organism, so it can, the organism can be controlled. It's a difficult technical problem, but it's not a philosophically very hard…”
“It's not just the competition between organisms, as Darwin suggested, or the competition between genes, the way in which the software can be written down, and partially at least, but it's the competition between spirits, between software agents.”
“The singularity in this way is not an event in the physical universe, It's an event in our modeling universe, a model, a point where our models of reality break down, and we don't know what's happening.”
“instead of doing like a react type reasoning loop, I think my belief is that we should be using like workflows, right? If we do this, then we always have a request and a complete workflow. We can fine tune a model that has a better workflow. Whereas it's hard to think about, like, how do you fine tune a better react…”
“the importance of supervising the process of AI systems, not just the outcomes. And so a big part of how, then, like, how Elicit is built is, We're very intentional about not just throwing a ton of data into a model and training it and then saying, cool, here's like scientific output. Like that's not at all what we do.…”
“having deeper models of how, let's see, what are the underlying structures of different domains, how they're related or not related, I think will be an important ingredient for models actually being able to make novel contributions.”
“I think agents are just absolutely the correct long-term direction, right? You just go to find what AGI is, right? You're like, hey, like, Well, first off, actually, I don't love AGI definitions that involve human replacement because I don't think that's actually how it's going to happen. I think even this definition…”
“Like de novo RL is like a pretty terrible way to get there quickly. Why are we rediscovering all the knowledge about the world? Like years ago, I had a debate with a, with a, with a Berkeley professor as to like, like what will it actually take to build HCI? And his view is basically that you have to reproduce all the…”
“I actually think that being an augmentation company Forces you to go develop your core AI capabilities faster than someone who's saying, ah, ok, my job is to deliver you a lights off solution for X.”
“if you go itemize out the number of things you want to do on your computer for which every step has an API, um, those numbers of workflows add up pretty close to zero.”
“You can write a very simple framework, uh, but then you also should be willing to eat the long compile times of, like, searching for that optimal performance at runtime.”
“Outside of this, like, where we don't have good symbolic models, like, synthetic data obviously, like, doesn't make any sense. So synthetic data is not a magic wand where it'll work in all cases, in every case, you know, whatever. It's just where we as humans already have good symbolic models of, we can, we need to…”
“And this, this might be a little controversial, but like, I find a lot of arguments, ah, based on whether, like, closed source models are safer, or open source models are safer, very much related to whether, what kind of cultural, ah, culture they grew up in, what kind of, Society they grew up in. If they grew up in a…”
“the open source thing will be very much in line with, um, getting to AGI, because open source has that, like, uh, natural selection effect. Like, if a better open source model comes, Really no one says, huh, I don't want to use it because there are ecosystem effects, I'm logged into my ecosystem, or like, I don't know…”
“And so one of my frustrations has been like so many startups are like, in my opinion, like Kubernetes wrappers and like, you know, like, and not very like thick wrappers, like fairly thin wrappers. And I think, you know, every startup is a wrapper to some extent, but like, you need to be like a fat wrapper. You need to…”
“I think that's actually true in most startups. Actually, it's like most startups die in the early stages, not because, uh, you run out of money, but really because you run out of motivation.”
“The chat, again, will help you, like, 10 or 20%, but it's unlikely that you're going to replace an employee with chat. Like, you know, you're not going to be like, oh, I have a relationship manager at JPMorgan Chase, and I've replaced him with an, you know, AI chatbot.”
“most of the software engineers in the world and most of the code-written in the world actually goes towards these internal facing applications.”
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not, you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
“you could use like a hundred times smaller language model and do much better at filtering than RLHF”
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation, and RL is one tool used at one point in the time. So that's…”
“if you want to get to the point where you can actually be truly agentic or, or like multi-step automated, uh, a necessary part of that is like the single step has to be robust and reliable.”
“Multiplying tensors is kind of, you know, anyone can, there's a lot of people who've made good matrix multiply units, right? But it's about, like, getting good utilization out of those, and interfacing with the memory, and interfacing with other chips really efficiently makes designing these chips very hard. And most…”
“like, I think it's generally been shown that if you have the space to just put The raw files inside of a big context window. That is still better than chunking and retrieval. It just, it just is.”
“And models today, they're not optimized for reasoning. It turns out that there's not actually that much explicit reasoning data on the internet.”
“Code also is a big piece of improving reasoning. So yeah, uh, generated code is not That much worse than, like, regular human written code. You might even say it could be better in a lot of ways.”
“The second thing we learned is that, uh, reinforcement learning is not a good vehicle. Like, pure reinforcement learning is not a good vehicle for planning and reasoning.”
“So to your question of, like, what agents work well and what doesn't work well, like, most of the agents don't work well, and we're slowly making them work better by improving the underlying model and improving these.”
“Chat as an interface is skeuomorphic. So in the early days, when we made word processors on our computers, they had notepad lines because that's what we understood, uh, you know, these like objects to be chat. Like texting someone is something we understand. So texting our AI is something that we understand. But…”
“I don't think RAG is enough for that kind of thing. But RAG is certainly enough for, like, user preferences and, um, and things like that.”
“I thought, okay, so if I do this at a much bigger scale, using all of Wikipedia, what would it need to be able to do to finish a sentence in Wikipedia, uh, effectively, to do it quite accurately, quite often? I thought, geez, it would actually have to know a lot about the world. You know, it would have to know that…”
“To me, the right way to do this is to fine, fine-tune language models, is to actually throw away the idea of fine-tuning. There's no such thing. There's only continued pre-training.”
“ULM fit is the wrong approach, um, and that's why we're seeing a lot of these, uh, You know, so-called alignment tax, and this view of like, oh, a model can't both code and do other things. You know, I think it's actually because people are training them wrong.”
“There's, there's a whole lot of technical debt everywhere, you know, nobody's really figured this stuff out, um, because everybody's been so busy building what we know works as quickly as possible. So, yeah, I think there's a huge amount of opportunity to, you know, I, I think we'll find things can be made to work a…”
“there's always going to be some curve regardless of like the performance of the best performing models of like, um, cost versus performance. Um, and so Um, what RAG does is it does provide extra data points along that access because you kind of control the amount of context you actually wanted to retrieve.”
“I don't think the delta on, like, improving the vector store, like, embedding lookup algorithm is that high. I think this stuff has been mostly solved, um, or at least there's just a lot of other stuff you can do, um, to try to improve the overall performance.”