Baker: Divergent Chinese open-source model architectures favor Nvidia GPUs
“If you look at the underlying architectures of the three, what I call big Chinese open source models, and maybe even through a few, or four, You know, if we have Quinn, if we have Kimmy, if we have DeepSeq, and then we have GLM, they're actually evolving in ve…”
Kedrosky: Minimal model differentiation will crush AI investment returns
“The convergence means that the model differences while there are so minimal as I can't tell the difference in the kind of Pepsi Coke phenomenon, which again, to cut to the investment chase suggests that the competition then becomes much more about marketing ex…”
Anthropic accuses major Chinese AI labs of distilling its Claude model
“Anthropic has accused DeepSeek, Moonshot, Minimax, ZAI, Alibaba, all of distilling Claude”
Calacanis: Major enterprise customers will abandon OpenAI and Anthropic for open-source
“And I'll make this prediction here, that you're going to see some of the major customers Of Anthropic and major customers of OpenAI. I'm talking about the eight and nine figure customers, people spending fifty million, a hundred million a year. They're leaving…”
Coogan: Proprietary AI lab revenues grew significantly despite the DeepSeek surge
“That's what we've seen throughout the whole deep seek saga when the lab revenues grew a ton, even though open source was big.”
Shao: Compute Limits Force Chinese AI Labs Into Distinct Specializations
“What you've seen is a lot of these tiny labs faced with compute constraint and in some ways capital constraint, they're forcing them to specialize rather than compete across every dimension. So, you know, for the sake of, you know, deep seek, it's really, real…”
Shao: DeepSeek and Moonshot Founders Focus on AGI Over Consumer Apps
“I think both, you know, DeepSeek and Kimi, if not seen as kind of the top two labs right now coming out of China, both founders have openly talked a lot about management of people you know, removing distractions, really focused on the pursuit of AGI, not kind …”
Shao: DeepSeek Lacks Capital Constraints and Pressure to Commercialize
“The fact that they don't actually have that much pressure to commercialize because they have enough money, he says, like, look, we have enough capital, we are not capital constraint compared to maybe other, other labs, and we are really, really focused on just…”
Shao: Quant Hedge Fund Backing Allows DeepSeek to Ignore Immediate Profits
“On the China side, deep seeks people have talked about how they are not revenue driven. So whatever money they, sorry, they're not profit driven. So whatever revenue they make, they want to put that money into R&D, given that they have the quant fund kind of l…”
Shao: Huawei and DeepSeek Are Collaborating Closely on Hardware Optimization
“Huawei is working very closely with deep seek on how to optimize hardware and software and try to figure out this.”
Liang: SambaNova serves 1.5T parameter models on one rack versus 10-20
“And so with sum it over, that minimum quantum is down to one rack. Right, where if you have other, other service providers, you just run, say, a DeepSeq model, which is now one and a half trillion parameters, just to run that, the minimum for some of the other…”
Evans: DeepSeek Proves Anyone With Two Billion Can Build Foundation Models
“Like anyone who's got a couple of billion dollars can make a foundation model which is what DeepSeq showed us as well.”
Banister: China will train AI on Western IP and sell output back
“And then what's really going to happen is China will sell our art back to us. So they're going to steal it all anyway, because they don't care. And they're just going to basically come back to us with these models, which we've seen with DeepSeek and everything…”
Ranjan Roy: DeepSeek and Chinese models are now in almost every agent conversation
“The conversation around moving towards deep seek and adding it into your agentic process or adding Chinese models did not exist in any conversation I was in 12 months ago and now is in almost every conversation, at least as an option, because cost has become s…”
Lemkin: DeepSeek's AI model is intentionally crippled inside China
“And DeepSeek is intentionally crippled in China. It's not as good as it is here. It can't search the web and it is trained on different data as near as I can tell.”
Roy: Due to cost, DeepSeek is discussed in every AI agent conversation
“The conversation around moving towards deep seek and adding it into your agentic process or adding Chinese models did not exist in any conversation I was in 12 months ago and now is in almost every conversation, at least as an option, because cost has become s…”
Siddharth: AI compute is shifting significantly toward post-training reinforcement learning
“In the past, it was a lot of the compute went into pre-training. Now a lot of compute goes into reinforcement learning in post-training as well. Especially after O-one came out and DeepSeek came out.”
DeepSeek Is One Of Ramp's Fastest Growing Vendors Despite Low Share
“Last month, DeepSeq was one of the fastest growing vendors on ramp. And yet, it's still a very small share of AI spend that is actually going through those rails.”
DeepSeek's Enterprise Growth Spike Unlikely To Withstand Rival Price Cuts
“Only about .4% of businesses are using it. Again, that's up from .1. So, Forex increase, but it's extremely small, and I think it's not going to be very durable, given that OpenAI and Anthropic are well positioned to respond to that with some price cuts.”
Chernin: Nebius had its best sales week during DeepSeek market selloff
“So I remember that Nebel stock went down 40% in one week or so in February, I think it was February or March, 20, 24, or 2025. And anecdotal story, the same exact week, we probably had the best week in sales.”
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
CommandCode offers 600 million DeepSeek tokens for $1 per month
“We launched a Go plan with just dollar one per month to, where you can do like, six hundred million tokens of DeepSQL for Pro in it, just to prove like, open models are actually really, really good, and they are catching up, right?”
Awais: Claude 3.7 Max is already CommandCode's second most used model
“But they will only be for deep seek when to 3.7 max is the second most used model on command code right now. It's just two or three days old.”
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Feldman says top Chinese open-source AI models lag closed-source models slightly
“This is made doubly worse by some of the best open source models were made by Chinese companies. And they are exceptionally good models. Kimi Ketu, Deep Seek. When the GLM, these are extraordinarily good models. They're not quite as good as the closed source m…”
Dubois: Models like Kimi and DeepSeek use ~1M RL data points
“Now when you look at reinforcement learning from models like Kimi or from DeepSeq models, it seems that they are closer to one million data points.”
DeepSeek Release Doubled Asian AWS GPU Rental Prices Due to Compute Demand
“And that was really strange because by Deep Seek Monday, it was super clear That this was going to be the most positive thing that had ever happened to compute demand. Prices in the AWS availability zones in Asia had already, like, doubled. You were seeing GPU…”
Rangwalla: AI Efficiency Breakthroughs Will Drive Adoption via Jevons Paradox
“If let's say that the deep seek moment that happened you know, last year, another moment like that happens where someone figures out a model that can do all the calculations with less power, less Semiconductors, less memory, all of that. If that happens, it is…”
Rao: Anthropic Reached Nearly $1B Run Rate by Late 2024 Series E
“At the end of 20, 24, we raised the Series E. You know, the business had scaled to, you know, close to a billion dollars of run rate revenue, but the day of our first close was the day of the DeepSeek news came out.”
Kantrowitz: DeepSeek blocked Tiananmen Square queries until user modifications removed censorship
“What happened when you tried to search for Tiananmen Square in DeepSeek? It was blocked. Now people were able to modify deep seek and, you know, strip that out, but it was only after modification.”
Paul Morris: DeepSeek Replicated OpenAI's AI for Just $26 Million
“Open AI took so many years and billions of dollars to create this and then deep seek, you know, yeah. Oh, we see how you did it, and you know, twenty six million dollars, six months later, you've got something not as good, but, you know, and so on and so on.”
Alex Kantrowitz: Governments cannot practically enforce bans on large language models
“They won't be banned. It's impossible. I mean, how can you ban them? Are you going to go and take your, is the government going to go grab your Mac mini out of your office where you've downloaded a version of deep seek and be like, all right, right to the poke…”
Moore: Russia Is DeepSeek's Second-Largest Market After China
“So Russia is the number two market for DeepSeek after, after China.”
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Hays: DeepSeek Is Hiring a 'PR Harmony Manager' Amid Distillation Scrutiny
“Deep seek is responding to distill gates and they are looking For a public relations harmony manager.”
Shaan Puri: DeepSeek outperformed OpenAI models without access to top-tier chips
“You ever heard of DeepSeek? They were in the news recently, created that big algorithm, didn't even have access to the best chips, and somehow outperformed all of the OpenAI models.”
Bishop: DeepSeek will launch its next model within eight or nine days
“And so now everyone's waiting for DeepSeek's next model, which is supposed to launch on or around Lunar New Year, which is next Wednesday the 17th. So we should have some sort of a DeepSeek model in the next eight or nine days.”
Gerstner: DeepSeek, Anthropic, and OpenAI will launch Blackwell-trained models in 4–8 weeks
“And just this year, like we're going to see the first models over the course of the next four to eight weeks out of deep seek, out of anthropic, out of open AI that are trained on blackwell servers, right? You're going to see a next generation of models Far mo…”
DeepSeek's architecture is still built on a GPT-2 scaffold
“And you can actually, in fact, Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest let's say deep seek version, 3.2 architecture. It's not like a big leap. It's still the same as scaffold.”
Dan Fu: DeepSeek-V3 was trained on ~2,000 H800s with 20% MFU
“If you look at the deep seek model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024. On last generation, kind of nerfed GPUs, H 800 instead of H 100, the 800 is nerfed by all sorts of ways fro…”
Mensch: DeepSeek built on Mistral's open-source MoE architecture
“Like we released the first
Sparse mixture of experts back at the beginning of 2024, and they built on top, and they released deep seek free, and deep seek was built on top of that. Well, it was, it's the same architecture, and we released, like, everything tha…”
Huang: DeepSeek paper was Silicon Valley's most important recent read
“This is, you know, let's face it, Deep Seek was probably the single most important paper that most Silicon Valley researchers read from in the last couple years.”
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Sacerdote: Inference-time reasoning will be adopted by all AI models
“And now deep seek is basically solidifying this inference time reasoning is going to be adopted by all the models.”
Altman: Google's Gemini 3 Exposed OpenAI Weaknesses But Failed to Hinder Growth
“Gemini three has not, or at least has not so far had the impact we were worried it might, but it did in the same way that deep seek did identify some weaknesses in our product offering and strategy.”
Siddharth: Chinese open-source AI models like DeepSeek and Qwen are state-of-the-art
“I think it's very impressive, like the progress that they've made in open source with DeepSeek Kimi Ketu, Kuen. These models are state of the art.”
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
AMD funded community DeepSeek kernel development through the GPU Mode Discord
“AMD's actually done this. Like there's some deep seek, like through GPU modes, discord, there have been some like deep seek kernels that they say you have, you know, guaranteed access to compute for, write them.”
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Wang: Chinese State Forbids DeepSeek From Communicating With Foreigners
“These companies generally don't like to talk to foreigners at the best of times. And then you have, An overbearing state that in some cases forbids deep seek from chatting with foreigners.”
Siddharth: Reasoning models ended demand for simple AI data factories
“This industry had a huge shift after the reasoning models came out late last year, O-one being the first. The shift is, before the reasoning models, this industry needed simple data. Gobs and gobs of simple data. You needed a data factory. Or, ah, this industr…”
Kokko: AI Will Autonomously Detect Market Shocks and Interview Experts
“When it doesn't find the information, it can recognize that autonomously and go and generate expert calls and have the AI do more expert calls, bring back the information, and now tell you that what does this new thing mean? When the next DeepSeek-like event h…”
Costica: Almost 10% of organizations adopted DeepSeek within one week
“Within a week, almost 10% of organizations were using DeepSeq.”
Costica: DeepSeek's database leak was a typical cloud exposure, not an exotic exploit
“It was a very, ah, what I would say, typical exposure that exposed a lot of sensitive data, but yet, ah, if we look historically, there were similar incidents, for instance, with Microsoft releasing a token that had access to sensitive data in, you know, a buc…”
Patel: NVIDIA GB200 yields 3x-4x performance per dollar on DeepSeek
“If you're running DeepSeq inference, the performance difference per GPU is like north of like six, seven X, and it continues to optimize you know, for DeepSeq inference. And so the, you know, then it's like, well, I'm only paying 60% more for six X, and it's l…”
Evans: Most users cannot distinguish Grok, Claude, Gemini, and DeepSeek in blind tests
“Like it seems to me right now you could do like a double blind test of the same prompt given to Grok, Claude, Gemini Mistral, Deep Seek. Do a double blind test. I bet most people wouldn't be able to tell which is which.”
Helberg: DeepSeek lied about compute capacity and distilled ChatGPT weights
“The main takeaways of how DeepSeq achieved the performance its performance was Incremental efficiency gains. They lied about their compute capacity because they have a billion dollar cluster, and they distilled ChatGPT's model weights.”
Weisbrot: Past experience living in China causes his distrust of DeepSeek
“I have not tried Deep Seek. I lived in China for long enough to know that I don't trust it.”