The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Copilot and Cursor Still Require Human Engineers to Drive Most Work
“Github Copilot or Cursor, which are incredible products. We use them internally. We're very happy with them. But again, this is, these are products where the engineer is driving most of the work. Like, even in agent mode, really the engineer is, like, driving …”
Needle-in-a-Haystack Tests Are Crude for Evaluating True Context Understanding
“So it's not just having long context, it's whether your model truly understands what's inside the context, and needle in the haystack tests are pretty crude and not very effective way of testing this kind of capability.”
Long-Context Attention Will Beat Agentic Localization for Codebase Indexing
“I bet would be on long context understanding and improving the attention mechanism over the long context.”
A 90% SWE-Bench Score Can Still Fall Flat in Customer Environments
“Autonomous coding benchmarks, let's say, like Sweetbench, are useful. I'm not going to discount them. They are useful. But let's say, you know, 90% on Sweetbench could still mean something that just falls over flat within a customer setting.”
Laskin: Other AI companies will converge on multi-agent retrieval architectures
“I'm sure that other companies will converge on it as well.”
Laskin: Pre-LLM AI breakthroughs spent more effort on environment design than model training
“And a lot of the project was in these big projects was not even on training the models. It was figuring out how the agents should actually interface with the environment that you're training it in.”
Laskin: Engineers spend 80% of their time comprehending complex systems
“When you look at what an engineer does in an organization, 80% of their time they're spending trying to comprehend complex systems and collaborating with teammates.”
Laskin: Evaluation Is the Most Important Differentiator for Frontier AI Labs
“This is I think the least spoken about part of what frontier labs do, but Possibly the most important, which is figuring out how they evaluate, like what makes Claude magically feel better at code than you know, another model out there. They did something righ…”
Laskin: Most Contributors on OpenAI's o1 Paper Worked on Evals
“When you look at the model card for, let's say, the O-one paper that came out, I think, last year. If you look at the distribution of what most people worked on, on that paper, it was evals.”
Laskin: Single-file coding questions do not require deep research agents
“If you're looking at a file, and there's like a specific thing in that file, and you're just trying to get a quick answer to it, you don't really need the hammer of like a deep research like experience. You don't need to wait, you know, like tens of seconds or…”
Laskin: Team-Wide AI Memory Governance Will Mirror Pull Request Workflows
“Where if you want to change the agents, the team wide memory, then it probably is going to look something like a pull request where the person who really understands that system Approves or, you know, edits it or something like this. I don't think it's going t…”
Laskin: Anthropic is generating massive revenue at an unprecedented growth rate
“When you look at how fast like Anthropics revenue is growing I think, right, they're kind of in this spot where it's like a massive revenue generating business that's growing at an unprecedented rate.”
Laskin: Current RL algorithms lack atomic credit assignment, causing meandering reasoning
“The RL methods we have today are quite bad, I would say, exploration and credit assignment. Like they, they're sort of just like the fundamental algorithms are take the things that work and make them happen more frequently, and the things that don't work and h…”
Laskin: AI models will beat humans at competitive coding within a year
“Code forces and other competitive coding environments. The models are almost best in the world, and within the year will probably be just the best in the world.”
Laskin: Sensory and vision-language model rewards are far more hackable than LLM rewards
“The challenge is that if we, if you think that language model rewards are hackable vision language model rewards or, you know, like other sensory signal rewards are infinitely more hackable.”
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”
AI Coding Agents Will Discover Unexpected 'Move 37' Breakthrough Solutions
“I think I think there are going to be a lot of move 37”
DeepSeek Revealed Attention Kernel Optimizations Kept Secret by Big Labs
“Deep seeks recent open sourcing of their various code components that they use to train that model, which I think outside of the big labs was not really well known to, right, it was not really well known how to write kind of a kernel that's optimized for this …”
Enterprise Engineers Spend Most of Their Time on Backlog Whack-a-Mole
“As a company gets larger, the, like, engineer goes from spending most of their time on, you know, working on the features that matter, and the kind of more, the more kind of value-driven work at a startup, To a very large company where you have giant code base…”
Laskin: Google Gemini's initial RLHF team was only 10 to 20 people
“I joined a small project at the time that you know, was tens of people. And that project became Gemini one and 1.5, and then obviously two and so forth. And I joined with my co-founder, my co-founder, Yannis was leading the reinforcement learning team, the RLE…”