base model

also referred to as: base models

8 statements across 8 episodes · 2 bullish · 3 bearish · 8 people on the record · first statement Apr 23, 2025 by Sid Bendre · across every show →

Everything said about base model, oldest first

Apr 23, 2025 positive
Insight
Prompt routing scales better than model routing as base models improve
“For us, we've actually like figured that the like unit, the like The return on it was not actually that beneficial. What makes sense for us instead is, you know, like we use base models for like the final, like response. What makes sense, what makes more sense…”
Sid Bendre Apr 23, 2025 ▶ 10:51 Tiny Teams: $6m ARR, 5m users with 4 employees — Sid Bendre, Oleve (Quizard AI/Unstuck AI)
May 23, 2025 neutral
Insight
Brown: Base LLMs will do anything up to their intelligence limit
“The base model in general of LLM is not artificially constrained in any way. Like, with the right prompt, it'll do whatever up to its intelligence limit.”
Will Brown May 23, 2025 ▶ 15:45 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Jul 29, 2025 negative
Insight
Fortuna: Base LLMs unsafely infer unconfirmed diagnoses from patient symptoms
“Patient, you know, they start to make these medical inferences. They're so smart, but they start to, like, infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say, like, I'm feeling sad and stressed out. Difficul…”
Brendan Fortuna Jul 29, 2025 ▶ 19:39 ⚡️Using RFT to Build Clinical Superintelligence
Sep 23, 2025 bullish
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00 ⚡️ Beyond Transformers with Power Retention
May 24, 2026 bearish
Insight
Updating mobile base models breaks per-app LoRAs, creating severe maintenance hurdles
“From a developer point of view, I think it will be very tricky because one, you don't want to have 20 different base models in the phone of the users. The battery will just die. You also don't want to have to update 20 LoRa every time you update the base model…”
Omar Sanseviero May 24, 2026 ▶ 15:56 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Jun 22, 2026 negative
Insight
Fredrikson: Prompt engineering cannot reliably enforce AI agent security policies
“Oftentimes people will try and prompt their way around it, like adjust the system prompt or like engineer the agent in a way where you're interjecting all the time and reminding it of what the original Goal and objective was, and that'll get you a little bit o…”
Matt Fredrikson Jun 22, 2026 ▶ 31:42 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jul 22, 2026
Insight
Kant: Base model pre-training is required to unlock major capabilities
“You can't fine tune your way to success, right? Major capabilities emerge from training a base model made accurate and useful during fine tuning.”
Eiso Kant Jul 22, 2026 ▶ 52:57 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Aug 3, 2026
Insight
Training speculative decoding models requires hidden states from the base model
“Now, with speculators today, you need to train the speculator using the base model itself, because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator.”
Philip Kiely Aug 3, 2026 ▶ 15:34 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.