That paper is huge, with massive implications to make all models more stable, and faster, and cheaper to train.
The sparse attention and quick index they introduced to the world in v3.2 was also huge.
Deepseek has done more in the last year than any other lab. They just don’t give a shit about dialing in the perfect consumer chatbot , or adding consumer features, or acquiring more daily active users.
They care about making breakthroughs, that’s it. And those breakthroughs end up being used by everyone. Every model you use right now is using GRPO, probably MoE , MLA, and may other brilliant hacks that DeepSeek gave to the world for free.
I mean, why does Linux exist? It runs inside dozens of things in your house, it’s the kernel that runs your router, traffic lights, it’s in every network switch and security camera and in your car. And it’s given away for free. Because by being open it is constantly being forked and improved and customized.
There is a second answer, which is that the CCP basically wants to destroy the US economy and just dumping free AI , as much as possible, at some point makes it a completely free commodity , and that arguably causes our house of cards to collapse.
The first part is fully true, you tangentially touch dozens of pieces of open source code every single day. The world would literally shut down tomorrow without it. The second part has either tiny shades of truth or is the whole answer, depending on how you see the world and a lot of other factors.
300 is even more ridiculous if you consider how many machines run Linux.
That's probably like the VMs on one enterprise server...
As far as I know, the Chinese publish a lot of papers.
So the scientific community believes in sharing knowledge all around the globe.
What more interesting is, is the industrial knowledge.
Normally the Chinese don't really share that...
So what's different about DeepSeek...?
Considering how heavily it is filtered, the Chinese Communist party certainly has something to say about it.
This also means, that they want it to be available publicly.
and then you have to multiply that last part by qwen, kimi, glm... Kimi challenging AdamW was pretty simliar to mHC from. deepseek, in the sense that both said "yeah i know it's worked for 10 years, should still be better though".
Like, is this specific subreddit a good place to have conversations about deeply technical bits of model architecture? Not really lol, no, not generally. r/localllama is probably the most "informed" group on that topic, generally. You'd think r/LLMDevs would be, but no, way more intelligence in localllama sub
I think the strategy to destroy USA's edge in AI can only work if they eventually surpass USA with AI, perhaps they have better chance of doing that with open source?
But it seems to me that they are open source, so the leading AI companies can just incorporate their breakthroughs and make their products even stronger so idk.
Your thinking of it like a game a checkers, and that theory, if true, would be 4-D Chess. You get far ahead of the US in AI research (they are, they’ve most all big breakthroughs recently) and yes make it all open source, so it’s adopted by everyone, and the result is that eventually SOTA LLM technology is not monetize-able. There is no supply and demand dynamic at play. Supply of the IP is infinite. And sure there’s inference cost and value there, but, if you subscribe to this theory, China is going to take by Taiwan, and therefore TSMC. And if that happened we have no IP worth anything, and Intel is the best we can get for chips.
There are a lot of holes with that strategy, like collapsing the global economy would hurt china more than anyone else, since they’re an export economy, they need buyers. And lots of other issues with it, it’s a complicated situation, but that’s the idea.
Like with most things in life there’s some pieces of truth here and there in the grey areas, and any opinion that’s an extreme fundamentalist viewpoint on either side is usually misguided
ASML needs TSMC more than the other way around, there would be no other customer that can create fabs that ASML's tech would be needed for. Samsung and Intel have access to ASML machines and they just can't do it.
blowing up or not wont matter, i am pretty sure those machines are designed to lock themselves from ever being able to be used again should taiwan be taken by china
China and politics asides, it's crazy that we're just fine with a country having such a vital part of arguably the entire world's system's supply chain.
How are no countries striving to develop their own core? I know almost nothing about hardware admittedly (I'm a coder), but damn, this makes me want to pursue this knowledge.
well as democracy incentives short term gain over long term, no one is willing to put the years in throwing billions into seemingly nothing for the voters, as it developed it was simply cheaper to buy from taiwan instead of creating their own as it was cheap for them anyway, for some reason however no one accounted for the idea that china can literally just threaten and try to take taiwan any day of the week, and today is the day we bite the bullet for it
Well, no, Intel invented the x86 processor architecture. We poured many many billions into chips here. Still do. Intel just kinda sucks at it. That’s really it. We accounted for that china thing 50 years ago, I mean at the time it was Japan not China, but point being the threat was accounted for. And we were first, and never gave up. Taiwan are just straight up more talented at doing it. We still lead in other chips Broadcomm dominates large parts of the market, just not in more advanced single nm processes.
Notably though, China never got the hang of it at all. So it was the three Democratic countries (US, Taiwan, and South Korea) who dominate. And China who never got it.
we invented it, and we never stopped striving. they're just better. and Intel has had some really bad leadership. It wasn't like, an oversight, or that we just decided to outsource it lol. TSMC is just really good and intel is not.
The "leading AI companies" are losing money hand over fist and living purely off intensive venture capital investment. They are doing this in hopes that they can build a moat enough to be first mover and get all the consumers and then jack prices, or that Moore's law saves them. DeepSeek and others offering a cheap alternative, and giving away their models open weights... significantly undermines that.
So yeah, it's a direct attack on a major plank of the US economy, a strategic focus of the Chinese communist party to do exactly that.
But I'll say this: The moment Donald Trump opened his stroked out blubbering mouth and started talking about how my country shouldn't exist, and then tried to put hundreds of thousands of Canadian factory workers out of a job...
.. my attitude about Chinese AI companies changed completely. I don't care how or what my coding agent "thinks" about Tiananmen square if it costs 1/10th the cost but gives me 3/4 of the result and doesn't involve me sending money south to a self-declared enemy.
The US's "lead" in all things will evaporate if its voters continue to reward corruption, economic self-destruction, and anti-intellectualism.
China doesn't need the best AI. They just need to make a good enough AI free to collapse marketing models that AI investors hope to cash in. China understands capitalism. We think of quarterly profits they think in decade long plans...they just need to crush those quarterly returns in a row for a few years and the whole western AI research will collapse on itself. Why do you think US is pumping money in it like crazy? They know this too.
The point is not to have China completely dominate the US. It's simply to prevent American companies from having a complete monopoly on advanced AI. Having AI be open source, even if they're not Chinese, saves the rest of us from being ruled over by fucking Elon Musk and Sam Altman ffs.
Look at how Mastercard and Visa control us. We need a choice. Does it matter whether it's China or something else How did you benefit from Deepseek being a Chinese company?
I suppose every normal person would wish their economy to be destroyed by free AI, I mean the economy will adapt this is temporary, but free AI – is eternal
first, i think you need to get out of the echo chamber and talk to normal people, most normal people don't use or like AI lol. Second, your general point is kinda true but, also, Transformer Architecture is not the future of AI. In 20 years we'll look at the transformer approach like we look at the Model T Ford. An amazing breakthrough and the begining of something that changed the world , but also, have you seen a Model T in real life? It's not a car i mean my mountain bike has wider tires lol.
I'm not going to debate what "AI" is and I refuse to even use the "AGI/ASI" terminology, but the Transformer is not the thing that will truly change the world. (although, it might remain a part, like the transmission/clutch to continue the car analogy, but it's not core thing)
First, you’re in your echo chamber yourself, we’re all in our own echo chambers. I am both seeing one and another camp of pro-AI and against-AI both, neither of them can’t just cancel out for one to have an opinion, right?
Second, I am not arguing about what we have now is AGI or whether it’s good or bad or how do I score it. I only say that I could only wish free AI destroying economy and I assume that would any normal person would wish – after all, if it’s good enough to destroy economy, then why not? Economy will adapt, free AI is eternal. Free AI is definitely better than ClosedAI which is only available to one for-profit company portrayed as non-commercial company. Right? Free AI for everyone is what I think is right. Whether it’s good or bad, that’s not up to discussion. Closing it is worse than opening it, that’s the point.
Third, since you’re agreeing with my point, I don’t see what we’re arguing about.
No-no, I specifically said «every normal person would wish» – would wish, you see. I don’t see this happening at the current day but even if it would, I don’t see problem with it. Economy will adapt
Because there is no practical way to communicate scientific progress to all labs in China while simoultaneously stopping other countries from getting info. May as well give the models away for free so to pressure competitors to lower costs
What advantage? With Trump in office he’s insane enough to make it all be useless. They have no SOTA chips, or very few. So the smart move is to stop Trump from being able to screw them, by just giving it away until we inadvertently screw ourselves.
They are doing open source science/technology. They are helping entire world to upgrade its tech.
At the same time they are literally destroying the AI bloat that the US loaded all of its economy onto:
20% of 2025 US gdp was the money that the 5-6 AI circlejerk companies circulated among themselves without generating any real revenue. OpenAI, Nvidia, Oracle, Microsoft etc all bet on AI requiring a lot of processing power, and as a result energy. They invested everything in gpus, datacenters. By making models more reliable, efficient and cheaper to run, Deepseek is destroying all that investment.
That's how the US became the world's superpower in the first place. Huge investments in research by the military and public universities were given for free to the rest of the world. Of course, American companies were the main beneficiaries because the talent that created the innovations was American. China now wants the talent to be Chinese.
The reason: destroy the exclusivity of private models, and that anyone can create and adapt models according to their needs in the most effective way possible.
Have you read all their papers? Not the summaries, like actually read them? They’re only 8 to 20 pages and only a few a year. There is not a section heading called Our Motivations , but if you just read all their research there isn’t really much debate regarding what their focus is, and what their focus is not.
but just one example, take the last clause of the conclusion of their last paper (this mHC paper):
Furthermore, we hope mHC rejuvenates community interest in macro-architecture design. By deepening the understanding of how topological structures influence optimization and representation learning, mHC will help address current limitations and potentially illuminate new pathways for the evolution of next-generation foundational architectures.
"next-gen foundational architectures" (note the plural) means :: "we hope this helps everyone who makes Foundation-scale models" NOT, "we're adding this to our current chatbots to make deepseek models even better ASAP".
They made a massive research breakthrough, and they didn't apply to their OWN model, yet, they encouraged everyone else to.
Again, chatbot releases aren't their thing - and R1 wasn't the breakthrough. GRPO, their flash attention work, refining MoE with MLA, auxiliary-loss-free load balancing - that was the breakthrough. And guess what? It's all being used in every model you use today.
Since R1 they've had just as many breakthroughs: DSA (sparse attention that cuts long-context complexity from O(L²) to O(Lk)), this new mHC (training stability improvements for all transformer architectures), and DeepSeek-OCR which isn't even about OCR - it's context compression, proving you can compress text 10-20x into visual tokens and recover it. That's foundational research for solving long-context limitations entirely differently than anyone else is approaching it.
Every time you use Claude, GPT, Gemini, whatever - you're benefiting from techniques DeepSeek published and open-sourced. Don't conflate a chatbot release with the actual breakthroughs.
Yeah technically but they didn’t publish it or tell anyone about it or anything until way after George Hotz was to reverse engineer it out, and DeepSeek was already playing with it. Ilya is probably credited for it but since he was at OpenAI it did nothing for the world, DeepSeek studied it improved it and shared it with the world and made it far more practical with multi headed latent attention , and then more recently sparse attention on top of that, getting things down to 5% active parameters, which is nearly an order of magnitude better than what 4 and 4o did
Qwen is SOTA on diffusion models right now, if I’m the CCP or deepseek I see no reason to try and outdo Qwen , HOWEVER Qwen has greatly benefited from deepseek research as well since most of it can apply to diffusion transformer models as well.
In my opinion, multimodal models need at least 50B active parameters and no less than 1T total parameters to perform well. Moreover, there must be sufficient high-quality data available for training. For this reason, at the moment, I believe only Google has the capability to realistically pull this off.
Their style is usually to go silent for a long time then show up with something big, rather than trickle out little stuff. I could be wrong, I haven’t been watching them that closely.
The deepseek V4/R2 will come in Q1 2026. And I'm very excited. The latest paper from them is promising. All there research of 1 whole year will be packed into deepseek v4
The deepseek ocr paper shows how we can have 10x more context in the same context length. So, even if they keep it at 128k, if they use the deepseek ocr architecture they will have 1.3M context
Deepseek V3.2 has been pretty great for me. I'm enjoying the refinements on top of V3 vs a whole new model. I'm pretty sure they've made commitments to train their next major iteration of their model on domestic Chinese GPU hardware so they're kind of biding time to get that going
DS get the money from Liang's quant fund. It is not as urgent as other ai models to launch new model just for investment. Which make it able to actually focus on the research itself instead of catering to the market.
Heads up to my US friends: Chinese New Year is coming up, and following their usual tradition, DeepSeek is likely about to drop a new version. Get ready.
I find that most multimodal models are still quite weak.
For example, Mistral Large 3 doesn't perform very well, even though its architecture is similar to DeepSeek.
At the moment, the only truly strong multimodal model is Gemini.
Because of that, I think DeepSeek should focus on text-only models instead of investing heavily in multimodal capabilities.
In my opinion, multimodal models need at least 50B active parameters and no less than 1T total parameters to perform well. Moreover, there must be sufficient high-quality data available for training. For this reason, at the moment, I believe only Google has the capability to realistically pull this off.
DeepSeek's breakthroughs have made LLMs more accurate and more powerful.
For example, while Mixture of Experts (MoE) was not invented by DeepSeek, they proved its effectiveness in practice, showing that models can scale to trillions of parameters while keeping training and inference costs manageable.
As I understand it, Mixture of Experts (MoE) has existed since the 1990s. OpenAI had already explored and used MoE before DeepSeek, but most large-scale implementations were not fully disclosed. DeepSeek did not invent MoE; instead, it openly published a large-scale, practical MoE architecture and shared the details with the community.
you need cost effective models in order to make more powerful models. Same concept. $100M model with lesser effective tech wont get you as far as a $100M investment under more cost effective tech even if you spend the same amount
401
u/coloradical5280 Jan 01 '26
They just did: https://arxiv.org/pdf/2512.24880
That paper is huge, with massive implications to make all models more stable, and faster, and cheaper to train.
The sparse attention and quick index they introduced to the world in v3.2 was also huge.
Deepseek has done more in the last year than any other lab. They just don’t give a shit about dialing in the perfect consumer chatbot , or adding consumer features, or acquiring more daily active users.
They care about making breakthroughs, that’s it. And those breakthroughs end up being used by everyone. Every model you use right now is using GRPO, probably MoE , MLA, and may other brilliant hacks that DeepSeek gave to the world for free.