The Wisdom Wall
37 quotable lessons, heuristics and mental models. Every one is playable at the moment it was said. No fortune cookies allowed.
“if you have a, a trillion dollar or five hundred billion dollar training run, then take fifty billion of that and make an ASIC. Like, it's fine. Like, like you will get more than 10% efficiency from From the ASIC. And like, that makes sense.”
“I think, um, what I've been calling agent labs, which are people who build on top of all the other models. We'll probably have a better time with the margins because they, they price against the end user hours spent or like human labor. Whereas models get commodity price per token. And so margin wise, we know inference…”
“the other thing I always think about from a finance point of view that, Maybe people don't even think, don't think about that closely, but obviously push back if you, if you disagree, is that it's basically a proxy for the gap between open models and closed models. Because if the open models do better, um, sort of the…”
“So you should actually build your products ahead of where costs are so that by the time they are popular, actually your frontier. So it's actually, I think the efficiency thing trips up a lot of people in terms of Being cost and cost conscious and wasting a lot of time on that when actually the, the most cost…”
“the reason that AI engineering can exist outside of the model labs is because the model labs release Models with capabilities that they don't even fully know because you never train specifically for it. It's emergent. And you can rely on basically crowdsourcing the search of that space or the behavior space to the rest…”
“I kind of call this the GPU smiling curve, uh, where the, the, the edges do well, cause you're either close to the machines and you're, you're like number one on the machines, or you're like close to the customers and you're number one on the customer side. And the people who are in the middle inflection, um, character…”
“AI engineer exists in the white surface area between the peak capability and deploying it everywhere else, right? So the more model research peaks and spikes capabilities in one domain, but it's not evenly distributed in all products yet, that's where engineers have a job. Forever, basically.”
“Actually, don't pick the solution, pick the problem. And if you're like, okay, I'm like, whatever it is in AI, whatever the hot thing is, whatever the new trend is, whatever the new model is, I will be the AI guy for dentists, or for lawyers, or for finance people, uh, or for coders, whatever. Um, then you will just…”
“a full world model must also be recursive, meaning that the participant in the world model must also be aware that they have a world model. Uh, which is like this whole recursive thing down the, down the line. Um, but yes, and, and that the world model can be wrong, and that they need to update it”
“I think one, one failure mode of the post-technical, you know, like manager, which you are, which I am, is that you think you, you, oh, AI can do it. And then you like, you know, you vibe code something, you throw it to your, your employees, and then you expect that to just pick it up. No, so like, I, you know, I…”
“I think like almost like every agent should have its own wiki that it's updating and that's, Persistent memory. That is a very weak knowledge graph. And you could strengthen it if you want more structure, but you may not need it.”
“every company must either build or buy a media company. Right. Until you, unless you realize that you have to take it that seriously, that you are running a media business in your company. You will never be good at it.”
“a lot of people think about cost in terms of dollars per million tokens, right? And I think that that is actually amateur thinking. This only, only the kind of pricing you care about if you're a solo developer, but once you're in a large scale, like you guys, uh, and it's also something I learned about at Cognition,…”
“I think there's a perverse incentives for research agents to take longer and it be perceived to be better to people are like, oh, you're, you're searching like so many, so many websites for me, you know, but like 30 of them are irrelevant. You know, like, I feel like right now we're in kind of a honeymoon phase where…”
“Chinchilla paper is compute optimal training, but what is not stated in there is it's pre-trained compute optimal training. And once you start caring about inference, compute optimal training, you have a different scaling law and in a way that we did not know last year.”
“I had always had this strong view that, um, filtering ranking Sorting, all this stuff is, is kind of like the same layer in the API stack, um, and, and should be, you know, kind of done together. Um, and this is my number one problem with, um, with AI news right now, which is that all filtering is effectively a rexus,…”
“There are trade-offs. You cannot have it all. You cannot have safe and cutting edge sometimes. Because sometimes cutting edge means unsafe.”
“few shot is not necessarily better than zero shot is, which is counterintuitive because you're working harder.”
“AI is mid by design. Right? And you do not want to be mid as a thought leader, as a content creator. People don't want to subscribe to mid things. They want to subscribe to a level slightly ahead of where they are.”
“I had seen basically front-end engineering become its own professionalized fields with dedicated conferences, dedicated influencers, and, uh, tech stacks and all those things, and I've seen the same thing for cloud engineering, um, and data engineering. And all that. And I was like, it's very obviously going to happen…”
“every database company is also a compute company.”
“And I think introducing narrow waste in a code base Allows you to contain slop, right? You can say, I don't care what goes on in here, but what goes in is defined by me. What goes out is defined by me. Do whatever you want inside.”
“So, so, so maybe, okay, maybe one, one sort of, I'll put this in my own words is agentic memory is sort of the stuff that you've done with the agent. And then the context graph is the stuff that you've done with everyone else. Because I think actually there's a lot of work By all the agent companies on querying their…”
“One observation I have, this, this kind of massive thesis I've been pursuing, which is what I've been calling an agent lab. Where you're sort of different than a model lab in the sense that you never train your own models, but you are the router evaluation layer, subject domain expert for choosing between models.”
“Basically, I think like the way that people ease into cloud co-work is like take a knowledge work task that you would normally be clicking around for, and then, uh, try to turn, turn that. And then you do the, okay, well, what if you went further? Okay. And then what if you went further, what if you sort of expand the…”
“all workloads are hybrid. Like, you know, uh, like you, you want the semantic, you want the text, you want the regex, you want SQL. I don't know. Um, but like, it's silly to like be all in on like one particular query pattern.”
“I think actually right now it's kind of a golden age of, um, pre LLM era infrastructure companies, because all your tooling is like inside the training data and like people who started it afterwards is not.”
“StackBlitz worked on WebAssembly containers for four years and got, was, I mean, they were making some progress, but like they, you know, nothing is compared to Bolt, which, which did the unlock. But like right now they can capitalize and just work on the agents because they already did the WebContainer stuff.”
“Everything can be broken down into some combo. Of compute, storage, and networking. If you fail to price one of them, they will, you will get abused because.”
“the slope of the Pareto Frontier is basically the state of the art of distillation, and the gentler the slope, the better distillation is.”
“your infrastructure shapes your apps. If your infrastructure doesn't make it easy for you, then you just won't do it.”
“It's basically an inversion of what Google Docs is, wants to do with Gemini. It's like Google Docs on the main screen and then Gemini on the side. And right, whatnot, what ChatGPT has done is Do the chat thing first, and then the docs on the side. But it's kind of like a reversal of, of what is the main thing.”
“And actually I found that that is a better way of doing this than when doing, you know, XML bracket, good example, close bracket, bad example, close bracket. Those, those examples tend to meet. There's an, there's an issue of prompts leaking, example leaking. Where there's a few short examples that you put in the…”
“Anyone who's doing, who's working with a ton of context has to put them at the end because they're swapping them out. That's just how it is. I don't see any way around it.”
“the labs that are not that frontier will, will keep measuring themselves on last year's benchmarks. And then the labs that are actually frontier will tell you about benchmarks you've never heard of.”
“The next frontier after RAG is planning and reasoning.”
“when you run a remote company, the general principle I want to tell people is that the thing, the money you would have spent on an office, you just have to spend on an offsite anyway.”