Qwen

includes Qwen 2.5, Qwen 3, Qwen 1.5, Qwen 72B, Qwen 7B, Qwen 2.5 Math, Qwen QwQ, Qwen 2, Qwen 2.5 Coder

10 statements across 8 episodes · 4 bullish · 3 bearish · 7 people on the record · first statement Dec 23, 2024 by Luca Soldani · said 91 times in 37 episodes since 2024 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (10), Elie Bakouch (9), Pratyush Maini (8), Vibhu Sapra (6), Ari Morcos (5), Mikhail Parakhin (4), Luca Soldani (4), Batuhan Taskaya (4)

tap a year for its mentions
0025105020202420252026episodesmentions
01020202420252026episodes it came up in
001.510320202420252026episodesmentions per episode
2026 22 mentions in 9 episodes 2 per episode
2025 46 mentions in 20 episodes 2 per episode
2024 23 mentions in 8 episodes 3 per episode

every mention, scene by scene, with the transcript →

Everything said about Qwen, oldest first

Dec 23, 2024 bullish
Assertion Supported
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
Luca Soldani Dec 23, 2024 ▶ 1:15 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Jan 2, 2025 negative
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Mar 23, 2025 positive
Insight
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
May 23, 2025 positive
Insight
Brown: Truncating reasoning model thinking mid-sentence still yields good outputs
“So it seems like artificially truncating the thought is actually like fine. Like the model can, even if like it got cut off mid-sentence with an injected like think token, these are smart enough models that they can kind of finish with the best that they got f…”
Will Brown May 23, 2025 ▶ 9:34 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
May 23, 2025 neutral
Opinion
Brown: Claude thinking and non-thinking modes likely use same underlying model
“I mean, I think these models should be the same model, and Anthropic knows what they're doing. Like, it's not that hard to, like, Quen did it in a very kind of, like, simple way, and they kind of talked about how they did it a little bit. But it's not, like, t…”
Will Brown May 23, 2025 ▶ 4:49 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Jul 31, 2025 neutral
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Oct 20, 2025 positive
Disclosure
Bakouch: Hugging Face plans to train an MoE model soon
“For example, we tried we are training MOE currently at TargetFace. I mean, we'll train soon. We start the training soon. And we tried with Megatron and we benchmarked, like, for example, the Mistral architecture with the Queen's three this one.”
Elie Bakouch Oct 20, 2025 ▶ 34:59 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Oct 20, 2025 negative
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Feb 5, 2026 neutral
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Feb 10, 2026 negative
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.