Baker: Divergent Chinese open-source model architectures favor Nvidia GPUs
“If you look at the underlying architectures of the three, what I call big Chinese open source models, and maybe even through a few, or four, You know, if we have Quinn, if we have Kimmy, if we have DeepSeq, and then we have GLM, they're actually evolving in ve…”
Kedrosky: Minimal model differentiation will crush AI investment returns
“The convergence means that the model differences while there are so minimal as I can't tell the difference in the kind of Pepsi Coke phenomenon, which again, to cut to the investment chase suggests that the competition then becomes much more about marketing ex…”
Finn Scans for Business Opportunities Every 20 Minutes Using Local Qwen
“Every 20 minutes, I have my agent going and using the Quen three seven model locally.”
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Siddharth: Chinese open-source AI models like DeepSeek and Qwen are state-of-the-art
“I think it's very impressive, like the progress that they've made in open source with DeepSeek Kimi Ketu, Kuen. These models are state of the art.”
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
50% of Hugging Face model derivatives are now based on Qwen
“I think 50% of all model derivatives being downloaded from Hugging Face or Quen base now.”
Bakouch: Hugging Face plans to train an MoE model soon
“For example, we tried we are training MOE currently at TargetFace. I mean, we'll train soon. We start the training soon. And we tried with Megatron and we benchmarked, like, for example, the Mistral architecture with the Queen's three this one.”
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Gerstner: Alibaba's Qwen open-source model surpassed 400M downloads
“So when the open source model out of Alibaba has passed, I think, four hundred million downloads.”
Krishnan: Chinese Models DeepSeek and Qwen Are the Best Open Source
“I think to even today, I would probably say the Chinese models, Deepsea, Quan are the best open source models”
Krishnan: Global Usage of DeepSeek and Qwen Is Geopolitical Soft Power
“When somebody is using DeepSeq or Quinn, that's an expression of soft power.”
Krishnan: Robotics Startups Are Heavily Using Distilled DeepSeek and Qwen Models
“When I was talking to a bunch of robotic startups, you're seeing a lot of distilled DeepSeq, a lot of distilled Quen out there.”
Brown: Claude thinking and non-thinking modes likely use same underlying model
“I mean, I think these models should be the same model, and Anthropic knows what they're doing. Like, it's not that hard to, like, Quen did it in a very kind of, like, simple way, and they kind of talked about how they did it a little bit. But it's not, like, t…”
Brown: Truncating reasoning model thinking mid-sentence still yields good outputs
“So it seems like artificially truncating the thought is actually like fine. Like the model can, even if like it got cut off mid-sentence with an injected like think token, these are smart enough models that they can kind of finish with the best that they got f…”
Gurley: China has four deep-pocketed open-source AI models
“Deep Seek, Led to Quinn. Led to Xiaomi has a, has their own model as well. They've all gone open source. And so this will be the fourth deep pocket funded model in China that are all open source.”
Brown: Alibaba Qwen makes the best model suites for research
“Like they make, I think, still the best model suites for like doing research.”
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Mascorro: Distillations from DeepSeek-R1 Outperformed Direct RL on Smaller Models
“So it turns out in their experiments, they took Lama's EV and some of these are QN models, and they basically apply RL straight the same way they did it with R one on these base models. And it turns out that it improved in some fields, but it was not a signifi…”
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
UiPath uses Alibaba's open-source Qwen model for semi-structured documents
“We are using Gwen, which is a fantastic model built by Alibaba, which is totally open source. We are using it into understanding, like a lot of our semi-structured documents”