Laskin: Organizational superintelligence will probably be an all-encompassing oracle
“We realized that probably the form factor of an organizational super intelligence, like the thing that's really gonna help organizations get a lot of stuff done, is probably gonna be something like an oracle. Like an oracle that understands the entire organiza…”
Laskin: AI models will interact with enterprise software primarily via APIs
“And so the way these language models are going to interact with any piece of software, not just Software engineering software, like Salesforce and other CRMs and creative tools and so forth. The majority of those interactions are going to be through function c…”
Laskin: Solving organizational code context yields all capabilities for superintelligence
“Like if you really solve this oracle for organizations just for coding, you've basically built all the capabilities you need to have super intelligence.”
Laskin: Self-improving code AI will mainly make algorithms more efficient
“I think what, I think coding intelligence that builds better coding intelligence will Make algorithms more efficient, basically.”
Laskin: The core ingredients to build AGI and ASI are now known
“But that was only meaningful when the ingredients for how to build artificial general intelligence, or ASI were not known. I think now they're known, and so.”
Laskin: Base language models before alignment are useless stochastic parrots
“If you play with one of these base models before they're aligned, they're really useless. They're, they are, they feel like stochastic parrots. They don't follow instructions.”
Laskin: Other AI companies will converge on multi-agent retrieval architectures
“I'm sure that other companies will converge on it as well.”
Laskin: Pre-LLM AI breakthroughs spent more effort on environment design than model training
“And a lot of the project was in these big projects was not even on training the models. It was figuring out how the agents should actually interface with the environment that you're training it in.”
Laskin: AI work awarded Physics Nobel Prize has had little impact on physics
“What's interesting is that the Physics Nobel Prize was given to something that has not really had that much impact in physics, but it is but I still buy it because it's kind of there's a physics smell to the breakthroughs that led to you know, these systems li…”
Laskin: Google Gemini's initial RLHF team was only 10 to 20 people
“I joined a small project at the time that you know, was tens of people. And that project became Gemini one and 1.5, and then obviously two and so forth. And I joined with my co-founder, my co-founder, Yannis was leading the reinforcement learning team, the RLE…”
Laskin: Coding agents are currently at an L4 junior engineer level
“There is in some sense, we are probably at a L four kind of junior, junior engineer level of autonomy, which is pretty incredible.”
Laskin: Principal-level AI engineers are a couple of years away
“And that the combination of this you know, L-Nine with Amnesia and the L-Nine's context core will together, you know, that will become the principal level engineer, the AI engineer. And so I actually think that that's not too far away. That's I would say in, y…”
Laskin: Sensory and vision-language model rewards are far more hackable than LLM rewards
“The challenge is that if we, if you think that language model rewards are hackable vision language model rewards or, you know, like other sensory signal rewards are infinitely more hackable.”
Laskin: Reinforcement learning is the only scalable path for synthetic data
“When we're generating synthetic data there is the only scalable path is really reinforcement learning.”
Laskin: Code reasoning models will generalize across other enterprise work
“The reason code is special is if you believe that the way a language model will interact with almost any piece of software is through function calls and therefore code, then if you build very capable reasoners coding reasoners that, you know, are sort of purpo…”
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
The Fundamental Language of AI Will Likely Remain Python
“The things that language models understand best tend to be the kind of piece of code that are represented on the internet. You know, maybe I don't think a lot of people would be happy with this, that maybe the kind of fundamental language of AI becomes Python …”
AI Coding Agents Will Discover Unexpected 'Move 37' Breakthrough Solutions
“I think I think there are going to be a lot of move 37”
DeepSeek Revealed Attention Kernel Optimizations Kept Secret by Big Labs
“Deep seeks recent open sourcing of their various code components that they use to train that model, which I think outside of the big labs was not really well known to, right, it was not really well known how to write kind of a kernel that's optimized for this …”
Frontier Labs May Hoard Superintelligent Models and Release Nerfed Versions
“You can imagine you know, the world converging on a few companies have really powerful coding models. They basically release a nerfed version of that to the public at large, and basically have a competitive advantage by having, you know, a super intelligent co…”
Needle-in-a-Haystack Tests Are Crude for Evaluating True Context Understanding
“So it's not just having long context, it's whether your model truly understands what's inside the context, and needle in the haystack tests are pretty crude and not very effective way of testing this kind of capability.”
Long-Context Attention Will Beat Agentic Localization for Codebase Indexing
“I bet would be on long context understanding and improving the attention mechanism over the long context.”
A 90% SWE-Bench Score Can Still Fall Flat in Customer Environments
“Autonomous coding benchmarks, let's say, like Sweetbench, are useful. I'm not going to discount them. They are useful. But let's say, you know, 90% on Sweetbench could still mean something that just falls over flat within a customer setting.”