Qwen

product on 13 shows · 25 statements across 20 episodes · said 206 times in 104 episodes since 2024

Latent Space 77 the Startup Ideas Podcast 21 20VC 20 TBPN 15 Big Technology 13 All-In 13 the a16z Podcast 12 the MAD Podcast 11 No Priors 9 BG2 Pod 6 Top Founders 4 Lenny's Podcast 3 the Y Combinator Startup Podcast 2

Mentions by year, every show

tap a year for its mentions
00502510050202420252026episodesmentions
02550202420252026episodes it came up in
001.3252.550202420252026episodesmentions per episode

Latent Space 77the Startup Ideas Podcast 2120VC 20TBPN 15All-In 13Big Technology 13the a16z Podcast 12the MAD Podcast 115 more shows

2026 93 mentions in 49 episodes 2 per episode
2025 88 mentions in 45 episodes 2 per episode
2024 25 mentions in 10 episodes 3 per episode

every mention on every show, scene by scene, with the transcript →

25 statements about Qwen, every show

ALL-IN Insight
Baker: Divergent Chinese open-source model architectures favor Nvidia GPUs
“If you look at the underlying architectures of the three, what I call big Chinese open source models, and maybe even through a few, or four, You know, if we have Quinn, if we have Kimmy, if we have DeepSeq, and then we have GLM, they're actually evolving in ve…”
Gavin Baker Aug 13, 2026 ▶ 1:00:19 Anthropic's $2T IPO, Zuck's AI Manifesto, Nvidia's $500B AI Bet, Grok's Comeback
BIG TECHNOLOGY Prediction Not checkable as stated
Kedrosky: Minimal model differentiation will crush AI investment returns
“The convergence means that the model differences while there are so minimal as I can't tell the difference in the kind of Pepsi Coke phenomenon, which again, to cut to the investment chase suggests that the competition then becomes much more about marketing ex…”
Paul Kedrosky Aug 12, 2026 ▶ 31:54 Why The AI Bubble Will Burst: The Most Logical Case — With Paul Kedrosky
STARTUP IDEAS Disclosure
Finn Scans for Business Opportunities Every 20 Minutes Using Local Qwen
“Every 20 minutes, I have my agent going and using the Quen three seven model locally.”
Alex Finn Jun 6, 2026 ▶ 33:14 Hermes Agent Desktop: Full Setup + Real Use Cases
LATENT SPACE Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
LATENT SPACE Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
20VC Assertion Supported
Siddharth: Chinese open-source AI models like DeepSeek and Qwen are state-of-the-art
“I think it's very impressive, like the progress that they've made in open source with DeepSeek Kimi Ketu, Kuen. These models are state of the art.”
Jonathan Siddharth Dec 1, 2025 ▶ 1:09:21 Turing CEO Jonathan Siddharth: Who Wins in Data Labelling & Why 99% of Knowledge Work Will Disappear · 20VC with Harry Stebbings
MAD Assertion Not checkable as stated
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Luca Soldaini Nov 20, 2025 ▶ 10:52 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Nathan Lambert Nov 20, 2025 ▶ 17:01 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Partly supported
50% of Hugging Face model derivatives are now based on Qwen
“I think 50% of all model derivatives being downloaded from Hugging Face or Quen base now.”
Nathan Benaich Oct 30, 2025 ▶ 34:33 State of AI 2025 with Nathan Benaich: Power Deals, Reasoning Breakthroughs, Real Revenue
LATENT SPACE Disclosure
Bakouch: Hugging Face plans to train an MoE model soon
“For example, we tried we are training MOE currently at TargetFace. I mean, we'll train soon. We start the training soon. And we tried with Megatron and we benchmarked, like, for example, the Mistral architecture with the Queen's three this one.”
Elie Bakouch Oct 20, 2025 ▶ 34:59 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
LATENT SPACE Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
LATENT SPACE Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
BG2 Assertion Supported
Gerstner: Alibaba's Qwen open-source model surpassed 400M downloads
“So when the open source model out of Alibaba has passed, I think, four hundred million downloads.”
Brad Gerstner Jul 31, 2025 ▶ 5:10 China Open-Source, Compute Arms Race, Reordering Global Trade | BG2 w/ Bill Gurley and Brad Gerstner · Bg2 Pod
NO PRIORS Opinion
Krishnan: Chinese Models DeepSeek and Qwen Are the Best Open Source
“I think to even today, I would probably say the Chinese models, Deepsea, Quan are the best open source models”
Sriram Krishnan Jul 31, 2025 ▶ 6:04 No Priors Ep. 125 | With Senior White House Policy Advisor on AI Sriram Krishnan
NO PRIORS Opinion
Krishnan: Global Usage of DeepSeek and Qwen Is Geopolitical Soft Power
“When somebody is using DeepSeq or Quinn, that's an expression of soft power.”
Sriram Krishnan Jul 31, 2025 ▶ 18:19 No Priors Ep. 125 | With Senior White House Policy Advisor on AI Sriram Krishnan
NO PRIORS Assertion Not checkable as stated
Krishnan: Robotics Startups Are Heavily Using Distilled DeepSeek and Qwen Models
“When I was talking to a bunch of robotic startups, you're seeing a lot of distilled DeepSeq, a lot of distilled Quen out there.”
Sriram Krishnan Jul 31, 2025 ▶ 26:06 No Priors Ep. 125 | With Senior White House Policy Advisor on AI Sriram Krishnan
Brown: Claude thinking and non-thinking modes likely use same underlying model
“I mean, I think these models should be the same model, and Anthropic knows what they're doing. Like, it's not that hard to, like, Quen did it in a very kind of, like, simple way, and they kind of talked about how they did it a little bit. But it's not, like, t…”
Will Brown May 23, 2025 ▶ 4:49 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Brown: Truncating reasoning model thinking mid-sentence still yields good outputs
“So it seems like artificially truncating the thought is actually like fine. Like the model can, even if like it got cut off mid-sentence with an injected like think token, these are smart enough models that they can kind of finish with the best that they got f…”
Will Brown May 23, 2025 ▶ 9:34 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
BG2 Assertion Supported
Gurley: China has four deep-pocketed open-source AI models
“Deep Seek, Led to Quinn. Led to Xiaomi has a, has their own model as well. They've all gone open source. And so this will be the fourth deep pocket funded model in China that are all open source.”
Bill Gurley May 22, 2025 ▶ 1:07:21 AI, Middle East, China, Tariffs, Recon Bill, Invest America | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
TBPN Opinion
Brown: Alibaba Qwen makes the best model suites for research
“Like they make, I think, still the best model suites for like doing research.”
Will Brown Apr 26, 2025 ▶ 17:24 Why Humor Is the True Test of AI Intelligence | Will Brown on TBPN
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
a16z Assertion Supported
Mascorro: Distillations from DeepSeek-R1 Outperformed Direct RL on Smaller Models
“So it turns out in their experiments, they took Lama's EV and some of these are QN models, and they basically apply RL straight the same way they did it with R one on these base models. And it turns out that it improved in some fields, but it was not a signifi…”
Marco Mascorro Mar 5, 2025 ▶ 25:26 DeepSeek, Reasoning Models, and the Future of LLMs
LATENT SPACE Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
Luca Soldani Dec 23, 2024 ▶ 1:15 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
20VC Disclosure
UiPath uses Alibaba's open-source Qwen model for semi-structured documents
“We are using Gwen, which is a fantastic model built by Alibaba, which is totally open source. We are using it into understanding, like a lot of our semi-structured documents”
Daniel Dines Dec 18, 2024 ▶ 5:37 Daniel Dines, UiPath CEO & Founder: Why Agents Do Not Mean RPA is F*** | E1240 · 20VC with Harry Stebbings

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.