Assertion Contradicted AI assessment confidence: 90% certainty 4/5 debate potential 1/5

Ubl: Cognition and Cursor shipped RL fine-tunes of open-source models

Malte Ubl · The Great Evals Debate — Ankur Goyal & Malte Ubl · Dec 7, 2025 · at 5:22

Vercel CTO Malte Ubl discusses recent reinforcement learning developments among AI coding agent startups with Swyx (Shawn Wang).

0:00 / 0:08exact quote · 8.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Just yesterday, I think we saw both Cognition, congrats, SWIX, and Cursor to ship RL fine tunes of unnamed open source models.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Malte Ubl

Opinion
Ubl: Anthropic's Boris Power Vibe-Codes From a Position of Privilege
“I think he comes from a particular position of extreme unusual privilege, which is that he works at an AI lab where like people in the office next door are like writing the evals and are like training the model like every day in exactly that way.”
Malte Ubl Dec 7, 2025 ▶ 5:44 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Malte Ubl Dec 7, 2025 ▶ 12:13 The Great Evals Debate — Ankur Goyal & Malte Ubl
Disclosure
Ubl: Vercel's Composite Models Are Faster Than Agentic Loops
“Basically what we do is we have this like composite model architecture. We run the frontier model and then we run the fine tune model after to fix its errors. That doesn't perform better than an agentic loop, but it's orders of magnitude faster, right?”
Malte Ubl Dec 7, 2025 ▶ 16:51 The Great Evals Debate — Ankur Goyal & Malte Ubl
Opinion
Ubl: Vercel leads the truly open-source deployment business model
“Vercel maybe has not invented this, but certainly kind of is the most successful at a model where you say, okay, I have this software library and it's truly open source. Everyone can run it. It comes with like adapters for every place on the planet, and that m…”
Malte Ubl Nov 1, 2025 ▶ 7:05 ⚡️ Ship AI recap: Agents, Workflows, and Python — w/ Vercel CTO Malte Ubl
Prediction Not checkable as stated
Future app platforms must extract auth and authorization completely
“Auth cannot be part of the app, because they're not going to get that right, right? So, Auth has to be extracted from the app. In fact, which data you can see, they also cannot be under control of the app, because again, you're going to get it wrong, right? So…”
Malte Ubl Nov 1, 2025 ▶ 40:25 ⚡️ Ship AI recap: Agents, Workflows, and Python — w/ Vercel CTO Malte Ubl
Disclosure
Ubl: Vercel publishes evals to influence OpenAI and Anthropic models
“I'm Vercel and I publish at Eval. That I want OpenAI and Anthropic to use to make sure when they ship the next model that they're better at the stuff that I care about.”
Malte Ubl Dec 7, 2025 ▶ 28:12 The Great Evals Debate — Ankur Goyal & Malte Ubl
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.