Opinion
Kant: Solving knowledge work does not require 2-3 orders of magnitude scaling
“I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to You know, solve knowledge work the accounting, the legal, the code that we write.”
Opinion
Kant: Anthropic's MCP and explicit tool calling abstractions are stupid
“I think MCP and tools are stupid.”
Opinion
Kant: Restricting open-weight AI models at current capability levels will hurt innovation
“We are not at a level of capability right now. That we should start restricting, you know, open models in any way, shape or form. I think it will hurt innovation if we do so.”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Insight
Kant: Open weights alone do not allow developers to recreate AI models
“Weights are a binary. Let's call them what they are. Yes, we can modify it and we can change them, but like, Giving someone the weights does not allow them ultimately to recreate what you're doing.”
Opinion
Kant: Labs benefiting from Chinese open research have an obligation to give back
“The incredible Chinese lab has done an amazing job at sharing their research, and we have definitely take, like, been on the receiving end of taking advantage of that. So When you're on the receiving end of something coming to you, I think you also have kind o…”
Prediction Not checkable as stated
Kant: Catching up to frontier AI labs will soon become unfeasible
“We've got a small window before models are Really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible.”
Insight
Kant: Industry will squeeze far more capability from smaller models via behavior
“We are going to be able to squeeze so much more out of smaller models than I think we had imagined in the industry. Because yes, there's intelligence and larger models are more intelligent. Like no doubt about it. We should continue to scale up. But the behavi…”
Opinion
Kant: Focusing solely on small open-source models is a cop out
“I think we ultimately only succeed if we scale our models as large as our competition. I do not, like, I think we should not put our head in the sand and say we're going to be king of open source small models. I think that's Frankly, it's a cop out.”
Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Prediction Open · timeframe Jul 2027
Kant: Prompts stuffed with dozens of tools will vanish in 12 months
“I think we will in 12 months not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.”
Insight
Kant: Pre-training compute runs are not the expensive part of AI
“The training run is not the expensive part. The training run is a very anticlimactic event, right?”
Prediction Not checkable as stated
Kant: AI will become the world's most demanded commodity with commoditizing margins
“Intelligence is the most in life. You're going to be the world's most demanded commodity. It will more commoditize in margin and price.”
Insight
Kant: RL compute cannot scale like pre-training due to task batch constraints
“And RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the en…”
Disclosure
Kant: Poolside is releasing open weights to prevent five-company AI concentration
“I think it all just came down to one thing and I'll stop the monologue is the fact that I rather live in a world that has a hundred foundation model companies than a world that has five. Even if I was one of the five. And the smallest and most meaningful contr…”
Insight
Kant: Foundation model building is 90% engineering rather than research
“Model building is ultimately 90% engineering. And I think we all know it in the industry, because if you look at words, every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code.”
Assertion Not checkable as stated
Kant: Poolside trains and launches models in five to eight weeks
“Laguna access two that we launched. It was five weeks from the beginning of pre-training to launch. The model that we're going to talk about today was eight weeks from start to pre-training to launch.”
Insight
Kant: 95% of foundation model building is data and compute efficiency
“I actually think you can sum down. So I saw 90. Five percent of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency.”
Assertion Supported
Kant: Poolside's 8B-active Laguna S solved Erdős 397 on DGX Spark
“A 118,000,000,008 B active model, which is not that large. It fits on a DGX spark and still runs at, you know, 3040 tokens a second on a spark is able to solve. Erdos three 97 independently. It's able to do complex programming tasks.”
Insight
Kant: Mid-training stages exist due to organizational silos, not necessity
“Mid-training exists because there's a mid-training team now, right? There's people or like people decide to focus on like a mid-training effort. But what you really want is engineering. And skill of experiments that allows for a much more continuous spectrum t…”
Insight
Kant: Base model pre-training is required to unlock major capabilities
“You can't fine tune your way to success, right? Major capabilities emerge from training a base model made accurate and useful during fine tuning.”
Insight
Kant: Long-horizon coding tasks are the path to AGI
“We think focusing on coding and long horizon software tasks is a path towards AGI because it forces us to solve the hard problems.”
Opinion
Kant: Audio does not push AI models closer to AGI
“I don't think audio. Adds to that. I don't think it pushes us close to AGI. I think it is a necessary modality as you get close to AGI.”