Model Routing

topic on 4 shows · 8 statements across 7 episodes

Latent Space the a16z Podcast Big Technology 20VC

8 statements about Model Routing, every show

a16z Insight
Building evals is the hardest part of model routing
“Really the hardest part of routing is building the evals and trying to determine in what places a set of intelligences should be used for a particular application.”
Ryan Chi Sep 9, 2026 ▶ 16:04 Inside the Race to Measure Frontier Intelligence
20VC Assertion Supported
Greze: Anthropic spends significant effort keeping model personalities consistent
“One of the problems when you do model routing is there's some companies like Anthropix spends a lot of time. I know we make fun of them online, but they spend a lot of time actually making sure all of their model families roughly don't change too much in terms…”
Jean-Denis Greze Sep 7, 2026 ▶ 20:58 Town vs Instinct vs GrokBot | Why the AI Assistant Market Is Not a Bubble
20VC Insight
Greze: User-facing persona limits AI model routing unlike backend reasoning
“The way I think about our stack is there is the part of the stack that deals with the user interface, like the feel and the personality. And there it is harder for me to just route Wildly, because I need consistency of the experience that is sometimes hard to …”
Jean-Denis Greze Sep 7, 2026 ▶ 22:02 Town vs Instinct vs GrokBot | Why the AI Assistant Market Is Not a Bubble
BIG TECHNOLOGY Prediction Not checkable as stated
Kutylowski: Model routing will increasingly prioritize performance over cost
“I think right now, more and more of that in, in, in, in, in this recent trend, we're going to probably see also, also more of this coming up when it's not going to be coming down to cost only, but also to the performance of those models and making sure that a …”
Jarek Kutylowski Jul 7, 2026 ▶ 13:26 Why Specialized AI Models Are Challenging the Frontier Labs — With DeepL CEO Jarek Kutylowski
Roy: AI industry rapidly shifted from frontier models to model routing
“In fact, I think the Apple, so I can tell you model routing has become like so dominant in conversations in my world now. Like it went from, you know, like just there's, you have to have the latest frontier model and they're all encompassing. And within the sp…”
Ranjan Roy Jun 8, 2026 ▶ 50:12 Did Apple (Finally) Get AI Right At WWDC?, Anthropic’s Worry, Microsoft vs. OpenAI
Prompt routing scales better than model routing as base models improve
“For us, we've actually like figured that the like unit, the like The return on it was not actually that beneficial. What makes sense for us instead is, you know, like we use base models for like the final, like response. What makes sense, what makes more sense…”
Sid Bendre Apr 23, 2025 ▶ 10:51 Tiny Teams: $6m ARR, 5m users with 4 employees — Sid Bendre, Oleve (Quizard AI/Unstuck AI)
LATENT SPACE Prediction Open · timeframe Jan 2028
Swyx: Anthropic and OpenAI will launch automated API model routing
“And so I pick and OpenAI have both keys that model routing on APIs already. And so I think they'll launch them, especially at some point where you can sort of prioritize the three tradeoffs that are in model routing, cost, speed, intelligence.”
Shawn Wang Jan 17, 2025 ▶ 23:35 OpenAI o1 isn’t a chat model (and that’s the point)
Schulhoff: Researchers should pay for top models instead of engineering routing
“For the most part, designing these systems where you're kind of routing to different levels of intelligence is a really time-consuming and difficult task, and, like, it's probably worth it to just use the smart model And pay for it at this point if you're look…”
Sander Schulhoff Sep 20, 2024 ▶ 42:32 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.