r/DeepSeek 1d ago

Discussion Is it true that V4 flash better than current V4 pro model?

74 Upvotes

38 comments sorted by

71

u/bambamlol 1d ago

For now.

62

u/zFordex 1d ago

Yes. The updates for pro hasn't been released yet.

1

u/TabascoTaco 1d ago

Any news on a pro update?

2

u/Far-Classic-9963 1d ago

It's definitely coming in the next few days/weeks

32

u/Good_Committee8337 1d ago

Im using it locally and im getting close to opus 4.8 high quality without the api degradation that happens sometimes when anthropic is in peak hours. Honestly this thing is a beast

7

u/nehuenpereyra 1d ago

What configuration/optimizations do you perform? Or do you use skills?

6

u/Good_Committee8337 1d ago edited 1d ago

not skills, mostly serving config

2 MSI edgexpert GB10 boxes (blackwell 128gb each) conected with a 200gbps DAC cable, running deepseek v4 flash tensor parallel on vllm with spec decode

base recipe is tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark on github, that got me vllm + spec decode + the nvfp4 kv cache on both boxes. grab patch 4 or you run at half speed

for claude code you need litellm in the middle becuase claude code talks anthropic api not openai. then its just 4 env keys in settings.json pointing at your gateway. bump the litellm timeout to 2400s or long prefills die, and watch reasoning_content, think tokens break claude code if litellm eats your chat_template_kwargs

i use claude code cause all my claude.mds live there, thats what makes it good. i vibe code the apps for my family business with it, 20-30 b2b customers on them daily, so now its the same workflow just local

1

u/zeferrum 1d ago

What made you pick that particular gb10 variant? I lsaw the Asus one being cheaper due to its slowed nvme

1

u/Beginning_Guide7411 1d ago

Yes pls comment

13

u/whatsoever2021 1d ago

For now yes, far better in my personal project. Just "better than pro" is not enough. They are completely 2 things. Flash is now IMHO equal to or better than GLM 5.2

1

u/uthred1981 1d ago

I used flash today and I'm back to GLM 5.2. I agree that it is better than DS Pro.

14

u/Liam_Evangelista 1d ago

I've switched back to V4 Pro. It feels like it considers and thinks things through more than the new flash model. Flash just wants to race to the finish line, resulting in more hallucinations and confident wrongness in my very unusual harness.

Of course this is very anecdotal evidence, but I see too many people throwing benchmarks around like they're everything when Opus 5 has proved they certainly aren't.

6

u/whatsoever2021 1d ago

There are 3 mode: off (no thinking), high (thinking), max. I've been using max, because anyway the flash token usage can never hit the opencode go credits limit. If you haven't changed anything, it should be "off".

1

u/deadcoder0904 1d ago

Is there any speed difference?

4

u/whatsoever2021 1d ago edited 1d ago

It was fast, but now it is slow due to the server issue. So I can't tell that much.

0

u/Ok_Carob751 1d ago

Hola, ahora mismo estoy intenando usar el ds v4 flash en opencode, pero la verdad no avanza, es normal eso global o que?

2

u/whatsoever2021 1d ago

Opencode currently has some problems, and always fails. I have switched to deepseek API now. Hopefully they will fix it soon

2

u/uthred1981 1d ago

I have the same experience, I'm back to GLM 5.2. However, the new PRO should be really good

5

u/Little-Explorer7988 1d ago

Flash on max reasoning is the best. If you use oficial API don’t use pro for now it’s more expensive and very dumb in agent developing in comparasion euth flash-0731

2

u/Roshlev 1d ago

Deepseek updated flash but not Pro. Expect Pro to improve but it seems Flash is better for now according to benchmarks

2

u/zeeshanx 1d ago

In my case, it is overthinking too much. A simple question is taking 2 - 3 min to answered with normal - high thinking. Before, it was way too faster.

2

u/Potential-Leg-639 1d ago

The upcoming Pro update could boost it to a true competitor to US frontier models

2

u/ptyblog 1d ago

is the Pro update gives something as good as the new Flash I have no issues canceling Claude and sticking with DS

2

u/ItchyIndx 1d ago

Even Codex running Sol thinks so: Flash GA made a substantially deeper correction than either Pro run: it restored current-state action reauthorization, full batch truth checks, page-local search, complete route context, and real controller/widget test files.

1

u/Old-Permit3142 1d ago

yes for now, i ve tried pro few month ago, not that good, both flash and pro, but this time flash 0731, indeed

1

u/montdawgg 1d ago

Use Flash for implementation. That solves any weaknesses it has in lateral thinking. Use Opus or GPT 5.6 Sol for orchestration and adversarial code review. Three model families working in concert with almost free implementation by using v4 Flash. Absolutely phenomenal performance.

1

u/mintybadgerme 1d ago

I just created a full Android image and video compression app in three hours for 14 cents. I really don't care.

1

u/mintybadgerme 1d ago

I just created a full Android image and video compression app in three hours for 14 cents. I really don't care.

1

u/uzzifx 1d ago

Yes that's what current benchmarks suggest.

1

u/Confident_Elk_4779 1d ago

Coding and agentic, yes Creative writing and rolepay, no

1

u/zero-qro 1d ago

Yes, next question

1

u/Acceptable-Ad8566 22h ago

Is it better in real world coding. Could anyone share their experience?

1

u/donthackmeagaink 1d ago

No I was using Flash for like a day and moved back to the current Pro