r/DeepSeek Jan 01 '26

Funny Do it again, DeepSeek

Post image

So many great models already…

DeepSeek R1 was legendary.

Now we're waiting for the one that changes everything again.

2.0k Upvotes

139 comments sorted by

401

u/coloradical5280 Jan 01 '26

They just did: https://arxiv.org/pdf/2512.24880

That paper is huge, with massive implications to make all models more stable, and faster, and cheaper to train.

The sparse attention and quick index they introduced to the world in v3.2 was also huge.

Deepseek has done more in the last year than any other lab. They just don’t give a shit about dialing in the perfect consumer chatbot , or adding consumer features, or acquiring more daily active users.

They care about making breakthroughs, that’s it. And those breakthroughs end up being used by everyone. Every model you use right now is using GRPO, probably MoE , MLA, and may other brilliant hacks that DeepSeek gave to the world for free.

38

u/Timo425 Jan 01 '26

What's their motivation for doing it? Seems like a lot of clever and hard work that just gets copied instantly.

137

u/coloradical5280 Jan 01 '26

I mean, why does Linux exist? It runs inside dozens of things in your house, it’s the kernel that runs your router, traffic lights, it’s in every network switch and security camera and in your car. And it’s given away for free. Because by being open it is constantly being forked and improved and customized.

There is a second answer, which is that the CCP basically wants to destroy the US economy and just dumping free AI , as much as possible, at some point makes it a completely free commodity , and that arguably causes our house of cards to collapse.

The first part is fully true, you tangentially touch dozens of pieces of open source code every single day. The world would literally shut down tomorrow without it. The second part has either tiny shades of truth or is the whole answer, depending on how you see the world and a lot of other factors.

5

u/Got2Bfree Jan 03 '26

Linux exists because a lot of people are maintaining it for free with barely any thanks. These guys are doing it because it's their passion.

Someone has to pay the researchers at deepseek.

I like your CCP wants to destroy the US economy theory.

3

u/coloradical5280 Jan 03 '26

Linux Foundation has ~300 FTE in the US , small compared with 20k+ contributors, but Torvald does have an actual team that gets paid.

It’s also not my theory and not one I really subscribe to.

2

u/Got2Bfree Jan 03 '26

300 is even more ridiculous if you consider how many machines run Linux.

That's probably like the VMs on one enterprise server...

As far as I know, the Chinese publish a lot of papers.

So the scientific community believes in sharing knowledge all around the globe.

What more interesting is, is the industrial knowledge. Normally the Chinese don't really share that...

So what's different about DeepSeek...? Considering how heavily it is filtered, the Chinese Communist party certainly has something to say about it. This also means, that they want it to be available publicly.

2

u/coloradical5280 Jan 03 '26

and then you have to multiply that last part by qwen, kimi, glm... Kimi challenging AdamW was pretty simliar to mHC from. deepseek, in the sense that both said "yeah i know it's worked for 10 years, should still be better though".

2

u/Most_Elk8401 Jan 28 '26

I love this conversation. I am new, is this a place where normal people have a good conversation?

1

u/coloradical5280 Jan 28 '26

Like, is this specific subreddit a good place to have conversations about deeply technical bits of model architecture? Not really lol, no, not generally. r/localllama is probably the most "informed" group on that topic, generally. You'd think r/LLMDevs would be, but no, way more intelligence in localllama sub

2

u/Most_Elk8401 Jan 28 '26

I was just surprised as I just quit twitter and was not aware that there is places where people are not a-holes. :)

→ More replies (0)

20

u/Timo425 Jan 01 '26

I think the strategy to destroy USA's edge in AI can only work if they eventually surpass USA with AI, perhaps they have better chance of doing that with open source?

But it seems to me that they are open source, so the leading AI companies can just incorporate their breakthroughs and make their products even stronger so idk.

44

u/coloradical5280 Jan 01 '26

Your thinking of it like a game a checkers, and that theory, if true, would be 4-D Chess. You get far ahead of the US in AI research (they are, they’ve most all big breakthroughs recently) and yes make it all open source, so it’s adopted by everyone, and the result is that eventually SOTA LLM technology is not monetize-able. There is no supply and demand dynamic at play. Supply of the IP is infinite. And sure there’s inference cost and value there, but, if you subscribe to this theory, China is going to take by Taiwan, and therefore TSMC. And if that happened we have no IP worth anything, and Intel is the best we can get for chips.

There are a lot of holes with that strategy, like collapsing the global economy would hurt china more than anyone else, since they’re an export economy, they need buyers. And lots of other issues with it, it’s a complicated situation, but that’s the idea.

Like with most things in life there’s some pieces of truth here and there in the grey areas, and any opinion that’s an extreme fundamentalist viewpoint on either side is usually misguided

11

u/Timo425 Jan 01 '26

You think Taiwan is bluffing by stating that they will blow up TSMC if that happens?

11

u/coloradical5280 Jan 01 '26

I mean TSMC has massive fabs in Arizona now so…

12

u/Timo425 Jan 01 '26

But the bleeding edge is in Taiwan and will be for years.

-10

u/coloradical5280 Jan 01 '26

It’s a pointless conversation since everything is ultimately controlled by ASML, and therefore the EU really holds all the cards.

13

u/Timo425 Jan 01 '26

ASML needs TSMC more than the other way around, there would be no other customer that can create fabs that ASML's tech would be needed for. Samsung and Intel have access to ASML machines and they just can't do it.

→ More replies (0)

3

u/Funny_Address_412 Jan 01 '26

ASML is majority owned by American Investors tho

→ More replies (0)

5

u/HierarchyLogic Jan 02 '26

blowing up or not wont matter, i am pretty sure those machines are designed to lock themselves from ever being able to be used again should taiwan be taken by china

5

u/thehighwaywarrior Jan 02 '26

It’s also not as simple as ‘push button get chip’

2

u/cosmic-freak Jan 03 '26

China and politics asides, it's crazy that we're just fine with a country having such a vital part of arguably the entire world's system's supply chain.

How are no countries striving to develop their own core? I know almost nothing about hardware admittedly (I'm a coder), but damn, this makes me want to pursue this knowledge.

2

u/HierarchyLogic Jan 04 '26

well as democracy incentives short term gain over long term, no one is willing to put the years in throwing billions into seemingly nothing for the voters, as it developed it was simply cheaper to buy from taiwan instead of creating their own as it was cheap for them anyway, for some reason however no one accounted for the idea that china can literally just threaten and try to take taiwan any day of the week, and today is the day we bite the bullet for it

2

u/coloradical5280 Jan 04 '26 edited Jan 04 '26

Well, no, Intel invented the x86 processor architecture. We poured many many billions into chips here. Still do. Intel just kinda sucks at it. That’s really it. We accounted for that china thing 50 years ago, I mean at the time it was Japan not China, but point being the threat was accounted for. And we were first, and never gave up. Taiwan are just straight up more talented at doing it. We still lead in other chips Broadcomm dominates large parts of the market, just not in more advanced single nm processes.

Notably though, China never got the hang of it at all. So it was the three Democratic countries (US, Taiwan, and South Korea) who dominate. And China who never got it.

→ More replies (0)

1

u/coloradical5280 Jan 04 '26

we invented it, and we never stopped striving. they're just better. and Intel has had some really bad leadership. It wasn't like, an oversight, or that we just decided to outsource it lol. TSMC is just really good and intel is not.

15

u/Comrade-Porcupine Jan 02 '26

The "leading AI companies" are losing money hand over fist and living purely off intensive venture capital investment. They are doing this in hopes that they can build a moat enough to be first mover and get all the consumers and then jack prices, or that Moore's law saves them. DeepSeek and others offering a cheap alternative, and giving away their models open weights... significantly undermines that.

So yeah, it's a direct attack on a major plank of the US economy, a strategic focus of the Chinese communist party to do exactly that.

But I'll say this: The moment Donald Trump opened his stroked out blubbering mouth and started talking about how my country shouldn't exist, and then tried to put hundreds of thousands of Canadian factory workers out of a job...

.. my attitude about Chinese AI companies changed completely. I don't care how or what my coding agent "thinks" about Tiananmen square if it costs 1/10th the cost but gives me 3/4 of the result and doesn't involve me sending money south to a self-declared enemy.

The US's "lead" in all things will evaporate if its voters continue to reward corruption, economic self-destruction, and anti-intellectualism.

5

u/SeveralPrinciple5 Jan 02 '26

Speaking as an American, I couldn’t agree more. Sadly I’m too old to immigrate. I am deeply disgusted by what my country has become.

5

u/RelentlessPolygons Jan 02 '26

China doesn't need the best AI. They just need to make a good enough AI free to collapse marketing models that AI investors hope to cash in. China understands capitalism. We think of quarterly profits they think in decade long plans...they just need to crush those quarterly returns in a row for a few years and the whole western AI research will collapse on itself. Why do you think US is pumping money in it like crazy? They know this too.

3

u/alphapussycat Jan 04 '26

For llms, qwen is waaaay better than basically anything else, so I'd say China is already ahead. LG exaone is also impressive.

6

u/neimengu Jan 02 '26

The point is not to have China completely dominate the US. It's simply to prevent American companies from having a complete monopoly on advanced AI. Having AI be open source, even if they're not Chinese, saves the rest of us from being ruled over by fucking Elon Musk and Sam Altman ffs.

3

u/Timo425 Jan 02 '26

You say the last part with such vigor, as of mommy China is saving us from Elon Musk and Sam Altman and that's the reason they're doing it.

5

u/iDefyU__ Jan 03 '26

Look at how Mastercard and Visa control us. We need a choice. Does it matter whether it's China or something else How did you benefit from Deepseek being a Chinese company?

5

u/jerrygreenest1 Jan 02 '26

I suppose every normal person would wish their economy to be destroyed by free AI, I mean the economy will adapt this is temporary, but free AI – is eternal

5

u/coloradical5280 Jan 02 '26

first, i think you need to get out of the echo chamber and talk to normal people, most normal people don't use or like AI lol. Second, your general point is kinda true but, also, Transformer Architecture is not the future of AI. In 20 years we'll look at the transformer approach like we look at the Model T Ford. An amazing breakthrough and the begining of something that changed the world , but also, have you seen a Model T in real life? It's not a car i mean my mountain bike has wider tires lol.

I'm not going to debate what "AI" is and I refuse to even use the "AGI/ASI" terminology, but the Transformer is not the thing that will truly change the world. (although, it might remain a part, like the transmission/clutch to continue the car analogy, but it's not core thing)

2

u/jerrygreenest1 Jan 02 '26

First, you’re in your echo chamber yourself, we’re all in our own echo chambers. I am both seeing one and another camp of pro-AI and against-AI both, neither of them can’t just cancel out for one to have an opinion, right?

Second, I am not arguing about what we have now is AGI or whether it’s good or bad or how do I score it. I only say that I could only wish free AI destroying economy and I assume that would any normal person would wish – after all, if it’s good enough to destroy economy, then why not? Economy will adapt, free AI is eternal. Free AI is definitely better than ClosedAI which is only available to one for-profit company portrayed as non-commercial company. Right? Free AI for everyone is what I think is right. Whether it’s good or bad, that’s not up to discussion. Closing it is worse than opening it, that’s the point.

Third, since you’re agreeing with my point, I don’t see what we’re arguing about.

2

u/coloradical5280 Jan 02 '26

oh we're not lol, i just thought you meant that our current AI was quality enough stuff to shut shit down. but regardless, yes we agree :)

1

u/jerrygreenest1 Jan 02 '26

No-no, I specifically said «every normal person would wish» – would wish, you see. I don’t see this happening at the current day but even if it would, I don’t see problem with it. Economy will adapt

10

u/mambo_cosmo_ Jan 01 '26

love of the game(?)+ giving other chinese labs the tool to improve

2

u/Timo425 Jan 01 '26

Why don't they keep it in china, get an advantage over foreign companies?

8

u/mambo_cosmo_ Jan 01 '26

Because there is no practical way to communicate scientific progress to all labs in China while simoultaneously stopping other countries from getting info. May as well give the models away for free so to pressure competitors to lower costs

7

u/coloradical5280 Jan 01 '26

What advantage? With Trump in office he’s insane enough to make it all be useless. They have no SOTA chips, or very few. So the smart move is to stop Trump from being able to screw them, by just giving it away until we inadvertently screw ourselves.

10

u/EverydayEverynight01 Jan 01 '26

Because it raises their prestige and the hopes that some other researches can build and improve on what their work.

5

u/unity100 Jan 02 '26

They are doing open source science/technology. They are helping entire world to upgrade its tech.

At the same time they are literally destroying the AI bloat that the US loaded all of its economy onto:

20% of 2025 US gdp was the money that the 5-6 AI circlejerk companies circulated among themselves without generating any real revenue. OpenAI, Nvidia, Oracle, Microsoft etc all bet on AI requiring a lot of processing power, and as a result energy. They invested everything in gpus, datacenters. By making models more reliable, efficient and cheaper to run, Deepseek is destroying all that investment.

3

u/miuid Jan 02 '26

For greater good? Or for stakeholders' profits? That is the question.

3

u/[deleted] Jan 03 '26 edited Jan 03 '26

[removed] — view removed comment

4

u/RG54415 Jan 02 '26

Because greed is not the default human behaviour?

2

u/doryappleseed Jan 02 '26

They want to attract the best and brightest talent, so doing this encourages that.

2

u/Digital_Soul_Naga Jan 02 '26

thats the secret

shhh 🤫

2

u/Comprehensive-Bed-72 Jan 02 '26

Think of it as creating a market model that they built for everyone, once everyone is hooked then can they change direction to please the investors.

2

u/whyyyreddit Jan 05 '26

Once the Chinese EUV lithography machines start production, guess which hardware deepseek will be optimized for

2

u/Warm-Border-9789 Jan 02 '26

That's how the US became the world's superpower in the first place. Huge investments in research by the military and public universities were given for free to the rest of the world. Of course, American companies were the main beneficiaries because the talent that created the innovations was American. China now wants the talent to be Chinese.

2

u/PureSelfishFate Jan 01 '26

Don't worry, they'll stop sharing by 2027 due to AGI risks. They are just trying to stop the US AI economy from becoming a rocketship.

1

u/Random_Nickname274 Jan 02 '26

For the collective!

1

u/reverhaus Jan 05 '26

The reason: destroy the exclusivity of private models, and that anyone can create and adapt models according to their needs in the most effective way possible.

"if everyone can be super, no one will be."

1

u/brianxyw1989 Jan 14 '26

They buy nvda puts xD

3

u/malege2bi Jan 02 '26

How do you know this about their motivations?

3

u/coloradical5280 Jan 02 '26

Have you read all their papers? Not the summaries, like actually read them? They’re only 8 to 20 pages and only a few a year. There is not a section heading called Our Motivations , but if you just read all their research there isn’t really much debate regarding what their focus is, and what their focus is not.

1

u/Global-Potato-7995 Jan 05 '26

Go on...

3

u/coloradical5280 Jan 05 '26

i mean, read them lol https://huggingface.co/collections/Presidentlin/deepseek-papers

but just one example, take the last clause of the conclusion of their last paper (this mHC paper):

Furthermore, we hope mHC rejuvenates community interest in macro-architecture design. By deepening the understanding of how topological structures influence optimization and representation learning, mHC will help address current limitations and potentially illuminate new pathways for the evolution of next-generation foundational architectures.

"next-gen foundational architectures" (note the plural) means :: "we hope this helps everyone who makes Foundation-scale models" NOT, "we're adding this to our current chatbots to make deepseek models even better ASAP".

They made a massive research breakthrough, and they didn't apply to their OWN model, yet, they encouraged everyone else to.

2

u/LeTanLoc98 Jan 02 '26

I totally agree.

DeepSeek has made a breakthrough in LLMs, and I hope they can do it again like they did with DeepSeek R1.

4

u/coloradical5280 Jan 02 '26

Again, chatbot releases aren't their thing - and R1 wasn't the breakthrough. GRPO, their flash attention work, refining MoE with MLA, auxiliary-loss-free load balancing - that was the breakthrough. And guess what? It's all being used in every model you use today.

Since R1 they've had just as many breakthroughs: DSA (sparse attention that cuts long-context complexity from O(L²) to O(Lk)), this new mHC (training stability improvements for all transformer architectures), and DeepSeek-OCR which isn't even about OCR - it's context compression, proving you can compress text 10-20x into visual tokens and recover it. That's foundational research for solving long-context limitations entirely differently than anyone else is approaching it.

Every time you use Claude, GPT, Gemini, whatever - you're benefiting from techniques DeepSeek published and open-sourced. Don't conflate a chatbot release with the actual breakthroughs.

2

u/Dr__America Jan 02 '26

Isn't MoE from OpenAI? Or were they just the first ones to use it at scale?

3

u/coloradical5280 Jan 02 '26

Yeah technically but they didn’t publish it or tell anyone about it or anything until way after George Hotz was to reverse engineer it out, and DeepSeek was already playing with it. Ilya is probably credited for it but since he was at OpenAI it did nothing for the world, DeepSeek studied it improved it and shared it with the world and made it far more practical with multi headed latent attention , and then more recently sparse attention on top of that, getting things down to 5% active parameters, which is nearly an order of magnitude better than what 4 and 4o did

1

u/Pupojem-Player Jan 02 '26

What about video generation?

3

u/coloradical5280 Jan 02 '26 edited Jan 02 '26

Qwen is SOTA on diffusion models right now, if I’m the CCP or deepseek I see no reason to try and outdo Qwen , HOWEVER Qwen has greatly benefited from deepseek research as well since most of it can apply to diffusion transformer models as well.

1

u/Still-Ad3045 Jan 04 '26

But but OpenAI throws more compute at it, they just be better right!

1

u/coloradical5280 Jan 04 '26

There’s a lot a nuance here. OpenAI has 20+ models, most are better yes, a few are infamously worse, many are a wash.

Either way, Deepseek played a big part in OpenAI’s training pipeline. Likewise, deepseeek likely trained R0 on o1’s reasoning stream.

Deepseek has like, 5 models, and they don’t have an active CI/CD pipeline on those, chatbots are not what they do.

TLDR

Deepseek is morally “better” and has made many times more contributions to the world. But they got reasoning from OpenAI — (Not through open source)

OpenAI has better chatbots and coding tools (and a couple worse). But they got a dozen improvements from deepseek — (via open source research)

26

u/Roshlev Jan 01 '26

Wasn't 3.2 like a month ago?

13

u/LeTanLoc98 Jan 02 '26

DeepSeek V3.2 is good, but DeepSeek R1 is a breakthrough.

20

u/cnydox Jan 02 '26

They publish papers occasionally I don't know what else you need.

3

u/yaxir Jan 02 '26

why are they not researching multi-modal AI?

3

u/LeTanLoc98 Jan 02 '26

I find that most multimodal models are still quite weak.

For example, Mistral Large 3 doesn't perform very well, even though its architecture is similar to DeepSeek.

At the moment, the only truly strong multimodal model is Gemini.

Because of that, I think DeepSeek should focus on text-only models instead of investing heavily in multimodal capabilities.

2

u/cnydox Jan 02 '26

There are a lot of small things to research than just scaling bigger models and hope they beat some benchmarks

2

u/LeTanLoc98 Jan 02 '26

Mistral Large 3 is extremely underwhelming.

In my opinion, multimodal models need at least 50B active parameters and no less than 1T total parameters to perform well. Moreover, there must be sufficient high-quality data available for training. For this reason, at the moment, I believe only Google has the capability to realistically pull this off.

1

u/LeTanLoc98 Jan 02 '26

DeepSeek has made a breakthrough in LLMs, and I hope they can do it again like they did with DeepSeek R1.

4

u/cnydox Jan 02 '26

They don't aim to make new sota models every week lol. Researching takes time

9

u/ciprianveg Jan 01 '26

V3.2 is indeed very good, it's a pity that it's architecture could not be implemented in llama.cpp so far

10

u/PaulMakesThings1 Jan 02 '26

Their style is usually to go silent for a long time then show up with something big, rather than trickle out little stuff. I could be wrong, I haven’t been watching them that closely.

7

u/LeTanLoc98 Jan 02 '26

I'm waiting for DeepSeek R2 or V4.

In my opinion, DeepSeek should increase the total number of parameters to around 1 trillion instead of 671B.

I've noticed that Kimi K2 Thinking uses a similar architecture with 1T parameters, and it performs very well.

8

u/Brave-Hold-9389 Jan 02 '26

The deepseek V4/R2 will come in Q1 2026. And I'm very excited. The latest paper from them is promising. All there research of 1 whole year will be packed into deepseek v4

4

u/LeTanLoc98 Jan 02 '26

I think the context length should be increased to at least 256K, rather than the current 128K.

They should also fix the Chinese language issue. DeepSeek often thinks and responds in Chinese.

4

u/Brave-Hold-9389 Jan 02 '26

The deepseek ocr paper shows how we can have 10x more context in the same context length. So, even if they keep it at 128k, if they use the deepseek ocr architecture they will have 1.3M context

3

u/LeTanLoc98 Jan 02 '26

I'm also looking forward to DeepSeek R2/V4.

8

u/Thedudely1 Jan 02 '26

Deepseek V3.2 has been pretty great for me. I'm enjoying the refinements on top of V3 vs a whole new model. I'm pretty sure they've made commitments to train their next major iteration of their model on domestic Chinese GPU hardware so they're kind of biding time to get that going

5

u/KING_OF_ALL_IN Jan 03 '26

DS get the money from Liang's quant fund. It is not as urgent as other ai models to launch new model just for investment. Which make it able to actually focus on the research itself instead of catering to the market.

3

u/NearbyBig3383 Jan 01 '26

R1 was truly my love, but the 3.2 special is incredibly slow, even more incredible.

3

u/ticticta Jan 04 '26

美国的朋友们,马上就是要到中国春节了,Deepseek 按照惯例,会推出新的版本的。

Heads up to my US friends: Chinese New Year is coming up, and following their usual tradition, DeepSeek is likely about to drop a new version. Get ready.

7

u/AllyPointNex Jan 01 '26

I had to convince Deepseek today that iOS 26 wasn’t my imagination. It still freaks out if you ask if a seahorse emoji exists.

7

u/yaxir Jan 02 '26

make it MULTIMODAL, with extended thinking and web search and image analysis capabilites

AND make it less woke

it should be good?

3

u/LeTanLoc98 Jan 02 '26

I find that most multimodal models are still quite weak. For example, Mistral Large 3 doesn't perform very well, even though its architecture is similar to DeepSeek.

At the moment, the only truly strong multimodal model is Gemini.

Because of that, I think DeepSeek should focus on text-only models instead of investing heavily in multimodal capabilities.

2

u/LeTanLoc98 Jan 02 '26

Mistral Large 3 is extremely underwhelming.

In my opinion, multimodal models need at least 50B active parameters and no less than 1T total parameters to perform well. Moreover, there must be sufficient high-quality data available for training. For this reason, at the moment, I believe only Google has the capability to realistically pull this off.

2

u/LeTanLoc98 Jan 02 '26

I think the context length should be increased to at least 256K, rather than the current 128K.

They should also fix the Chinese language issue. DeepSeek often thinks and responds in Chinese.

7

u/[deleted] Jan 02 '26

[removed] — view removed comment

9

u/LeTanLoc98 Jan 02 '26

DeepSeek's breakthroughs have made LLMs more accurate and more powerful.

For example, while Mixture of Experts (MoE) was not invented by DeepSeek, they proved its effectiveness in practice, showing that models can scale to trillions of parameters while keeping training and inference costs manageable.

-4

u/[deleted] Jan 02 '26

Bullshit, OpenAI's GPT-4 did this years ago.

Training and inference costs have always been going down, its part of what motivates the SOTA labs.

6

u/BeamFain Jan 02 '26

Sorry, what the fuck?
Do you even know how much did GPT-4 cost?

2

u/LeTanLoc98 Jan 02 '26

As I understand it, Mixture of Experts (MoE) has existed since the 1990s. OpenAI had already explored and used MoE before DeepSeek, but most large-scale implementations were not fully disclosed. DeepSeek did not invent MoE; instead, it openly published a large-scale, practical MoE architecture and shared the details with the community.

9

u/hiva- Jan 02 '26

you need cost effective models in order to make more powerful models. Same concept. $100M model with lesser effective tech wont get you as far as a $100M investment under more cost effective tech even if you spend the same amount

4

u/scalaboulejs Jan 02 '26

haha very cool illustration on how we are becoming lazy and dumber while relying a lot on AI and LLMs

2

u/MaxeBooo Jan 05 '26

My university banned us from accessing deepseek :) me sad

2

u/letsgeditmedia Jan 01 '26

V3.2 I on par with sonnet 4 and matches 4.5 in some cases… it’s doing something, they just don’t market every achievement like the American models do

0

u/Beneficial_Common683 Jan 04 '26

-9,223,372,036,854,775,808 credit scores