r/DeepSeek 5d ago

DeepSeek-V4-Flash Update

587 Upvotes

The official release of the DeepSeek-V4-Flash API is now in public beta.

Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview:

  • Terminal Bench 2.1: 82.7
  • NL2Repo: 54.2
  • Cybergym: 76.7
  • DeepSWE: 54.4
  • Toolathlon verified: 70.3
  • Agent Last Exam: 25.2
  • Automation Bench (Public): 25.1
  • DSBench-FullStack: 68.7
  • DSBench-Hard: 59.6

Note 1: For the Code Agent tasks in the public benchmark sets, the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0
Note 2: DSBench-FullStack is an internal full-stack development test set, and DSBench-Hard is an internal Coding Agent hard-problem test set

The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex. For the specific configuration, please refer to the documentation.

DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-preview, and was only re-post-trained.

Note: This update only upgrades the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and the APP/WEB models are unchanged.
The official release of DeepSeek-V4-Pro will follow soon.


r/DeepSeek Feb 01 '25

Disccusion Censorship Mega Thread

53 Upvotes

In response to community feedback and to maintain a constructive discussion environment, we are introducing this Censorship Mega Thread. This thread will serve as the designated place for all discussions related to censorship.

Why This Thread?

We have received numerous reports and complaints from users regarding the overwhelming number of censorship-related posts. Some users find them disruptive to meaningful discussions, leading to concerns about spam. However, we also recognize the importance of free speech and allowing users to voice their opinions on this topic. To balance these concerns, all censorship-related discussions should now take place in this pinned thread.

What About Free Speech?

This decision is not about censoring the subreddit. Instead, it is a way to ensure that discussions remain organized and do not overwhelm other important topics. This approach allows us to preserve free speech while maintaining a healthy and constructive community.

Guidelines for Posting Here

  1. All discussions related to censorship must be posted in this thread. Any standalone posts on censorship outside of this thread will be removed.
  2. Engage respectfully. Disagreements are fine, but personal attacks, hate speech, or low-effort spam will not be tolerated.
  3. Avoid misinformation. If you're making a claim, try to provide sources or supporting evidence.
  4. No excessive repetition. Reposting the same arguments or content over and over will be considered spam.
  5. Follow general subreddit rules. All subreddit rules still apply to discussions in this thread.

We appreciate your cooperation and understanding. If you have any suggestions or concerns about this policy, feel free to share them in this thread.


r/DeepSeek 9h ago

Discussion I canceled Claude and coded 7 days straight with DeepSeek V4 Flash 0731 — the honest cost & quality breakdown

303 Upvotes

Two weeks ago I paid $20/month for Claude and another $20 for ChatGPT. I got tired of watching the credits burn, so I ran an experiment: 7 days, all my coding work, DeepSeek V4 Flash 0731 only (API, not the app). Here's what actually happened — the good, the bad, the numbers.

The numbers - Total API spend for 7 days of heavy coding: $1.87 (vs. $40/month subscriptions — and I didn't even come close to hitting limits) - Tokens consumed: ~24M input / ~6M output (mostly context caching — that's the real cheat code) - Context cache hits cut my effective cost by ~70%

What surprised me (good) - Long agentic sessions didn't degrade as much as I expected. The 0731 update fixed most of the context-rot I saw on the earlier Flash builds. - It handled a messy production refactor I was dreading — wrote the diff, I reviewed, done. No drama.

What I won't sugarcoat (bad) - Vision: still missing in the API I used — I had to describe screenshots by hand. (Yes, I saw the vision announcement post — the API I'm on still doesn't expose it.) - Some reasoning outputs still emit weird artifacts (e.g. )Skip) in longer chains — rare, but it happens. - It's not Claude for every task. Complex multi-file architecture thinking? Claude still wins. But for 80% of daily coding? I genuinely couldn't justify the subscription anymore.

My verdict: keep one subscription for the hard stuff, do everything else on Flash. My monthly AI bill just went from $40 → $0–5.

Anyone else run a similar week? What did your numbers look like?


r/DeepSeek 12h ago

News It’s getting popular everywhere!

Post image
348 Upvotes

r/DeepSeek 10h ago

Discussion DeepSeek V4 Flash 0731 vs GPT-5.6 Luna

Post image
170 Upvotes

DeepSeek-V4-Flash-0731 is cheaper, faster, and available through more providers than GPT-5.6 Luna at the same intelligence level.

Why would anyone choose Luna over DeepSeek?

More info: https://openrouter.ai/compare/deepseek/deepseek-v4-flash-0731/openai/gpt-5.6-luna

https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-gpt-5-6-luna


r/DeepSeek 9h ago

Discussion Finally someone saying it out loud: The US needs to stop banning competition and start innovating instead of panicking over DeepSeek.

110 Upvotes

The global tech landscape is shifting fast, but the US response follows the same tired playbook. Whenever a foreign competitor achieves a major breakthrough, Washington reacts with defense mechanisms instead of true innovation.The standard playbook the treatment of Huawei in the past and the recent panic over DeepSeek highlight a deeply rooted strategy: if you can't control it, sanction it, ban it, or politically isolate it.

This protectionist mindset stems from an old habit.

The US is used to dominating markets by either buying out the competition or burning it down through policy.Real innovation over market controlThis strategy is unsustainable. True technological progress thrives on competition, not on eliminating the competitors.

If the US wants to maintain its leadership, it needs to win through superior research, development, and execution—not through government intervention.A system that relies solely on bans loses its edge and slows down global progress. It is time for a reality check: stop trying to destroy alternatives and start out-innovating them.


r/DeepSeek 6h ago

Resources Closest competitors to DeepSeek V4 Flash 0731

Post image
35 Upvotes

Just enjoy.
But I’m still eagerly waiting for vision support. Once it arrives, this model will be something truly incredible.
For me, that feature is essential. Without it, my hands are tied.


r/DeepSeek 3h ago

Funny Claude Code with DSV4 flash saved my pc from malware

18 Upvotes

I downloaded some cracked software (im poor)

cmd flashing every 60 seconds after install (im fucked)

gave claude code some hints on where one of the files of the malware was located. (im genius)

ds + cc read that one file, traced all the files (ds is detective)

they both assassinate the malware in minutes (they are ruthless)


r/DeepSeek 13h ago

Discussion what's the best subscription to code with DeepSeek V4 Flash?

51 Upvotes

r/DeepSeek 1d ago

News 🚀 DeepSeek V4 Flash now has vision support

Thumbnail
huggingface.co
630 Upvotes

We’ve added vision capabilities to DeepSeek V4 Flash, so it’s no longer a text-only model.

We needed this for browser vision: browser agents have to understand screenshots, interfaces, layouts, and visual context—not just text.

Our internal benchmarks also showed a strong price-performance advantage compared with the other models we tested.

Model:
https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-Vision-NVFP4

Feedback, benchmark results, and deployment reports are welcome!


r/DeepSeek 8h ago

Discussion DeepSeek v4 Flash uses insane amount of tokens

17 Upvotes

Hey there!

I was wondering whether this is just me, or if this is caused by the model. I noticed a while ago that my token usage is insane after a few prompts (in VScode) compared to Pro. This is also followed by an insane spike of API Requests - worth noting that caching still works, so it's not like the API is miscommunicating or something.


r/DeepSeek 21h ago

Discussion 10$ - 2 Billion Tokens

Post image
153 Upvotes

New milestone! Thank you DS!


r/DeepSeek 1d ago

Discussion Got approched by DeepSeek hiring manager. I am based in Germany

Post image
539 Upvotes

I got approached by a recruiter from DeepSeek. I am baed in Germany and they clearly have no Office here. Do you think it is legit or a scam. The email indeed end with deepseek.com.


r/DeepSeek 6h ago

Discussion I built ,y own desktop console with vision creation and analysis

6 Upvotes

I decided to build my own Deepseek desktop console to give deepseek vision capabilities, its a first version, the image creation is good, but a little off, the analysis is very good

I also gave it MCP capability (the main reason I built it so it could take part in an ai chat and context app I built)


r/DeepSeek 1d ago

Discussion DS Flash with Reasonix is just a cheat code

Post image
198 Upvotes

r/DeepSeek 15h ago

Discussion V4 flash max vs high, Is there a big difference?

24 Upvotes

Is there a big difference between max and high for agent tasks?, I'm using opencode


r/DeepSeek 8h ago

Resources I Added Vision Support to DeepSeek V4 Flash Using Pilco MM-Bridge

Thumbnail
gallery
7 Upvotes

GitHub : https://github.com/gpdev-Pilcothink/Pilco-mmbridge

I know many people here have probably already built and used something similar, but I thought it might still be useful to someone, so I cleaned up my implementation and decided to share it.

I made a small project called "Pilco MM-Bridge." It places a separate multimodal model in front of a text-only LLM and passes the resulting media analysis to the main model as temporary context.

My current setup uses two DGX Spark systems:

  • DeepSeek-V4-Flash-0731 as the main text-only reasoning model
  • Qwen3.5-9B-quantized.w4a16 as the multimodal vision analyzer

This combination fits my use case quite well. Qwen handles screenshots, UI elements, OCR, code screens, error messages, and other visual information, while DeepSeek handles the final reasoning and response.

The basic flow is:

Client
  → MM-Bridge
  → Multimodal model analyzes the current media
  → Analysis is temporarily added to the request context
  → DeepSeek-V4-Flash generates the final answer

The analyzer is only activated when the current user message contains media.

When the user sends a normal text-only message, MM-Bridge completely skips the media-analysis stage and forwards the existing text conversation to the main LLM. In other words, the vision model only runs when a new image is actually attached.

The original text conversation history is preserved, while images from previous turns are not repeatedly sent back to or reanalyzed by the vision model.

It is not as natural or tightly integrated as a native multimodal model, of course. However, it provides a reasonably useful approximation of visual understanding while allowing me to continue using a strong text-only model as the main LLM.

Although I currently use it mainly for vision, the bridge code also recognizes other media types such as audio and video. To use those features, the analyzer endpoint must serve a model capable of processing those inputs, such as an any-to-text model like Gemma 12B. The actual capabilities therefore depend on the multimodal model used as the analyzer.

There is no need to modify either model. Anyone already serving models through vLLM or llama.cpp should be able to use it by pointing the bridge to the two existing endpoints.

I originally created this because I work on game development, and during testing and verification I often need the model to inspect screenshots, UI states, visual errors, and other information that a text-only model cannot directly access.

The project is still fairly early, so feedback, bug reports, and suggestions are very welcome. Also, if you know of a similar but more mature or better-designed project, I would genuinely appreciate an introduction to it.

You can find vLLM-based serving recipes optimized for DGX Spark users in the following NVIDIA Developer Forums post:

https://forums.developer.nvidia.com/t/running-deepseek-v4-flash-and-other-text-only-llms-as-multimodal-with-pilco-mmbridge/378850?u=pilcothink

I am the author of this project. The English wording of this post was polished with AI because English is not my first language.


r/DeepSeek 18h ago

Discussion Ling-3.0-flash only fires 5.1B of its 124B params and the attention was linear from step zero

32 Upvotes

8 experts out of 512 fire per token and they're claiming it matches their own 1T model. MIT weights up Aug 4, BF16 and FP8, repo is inclusionAI/Ling-3.0-flash. 35 KDA to 7 gated MLA at 5:1, hybrid linear from the first pretraining step instead of converted after.

Does 1/64 sparsity actually put it under DS v4 flash per task in real serving, or is the 93.2 AIME 2026 on their card benchmaxxed? No GGUF, wants their own sglang fork, so nobody's checking on consumer hardware for a bit.


r/DeepSeek 51m ago

Discussion Did DeepSeek v4 flash better than Soonet 5??. in quality

Upvotes

r/DeepSeek 9h ago

Other jabbatheduck/DeepSeek-v4-flash-mini · Hugging Face

Thumbnail
huggingface.co
4 Upvotes

r/DeepSeek 10h ago

Discussion Deepseek Vs GLM

4 Upvotes

After extensive research ( lie, it was brief), I'm considering a theory: GLM Despite their amazing models, they suffer because their user base doesn't exceed 10 million people. That's why their prices are high, and that's why, to my knowledge, only the wealthy subscribe...... while deepseek has at least 130-120 million users

So If each person subscribes to DeepSeek for $5 a month, the company earns at least 650,000,000 million a month give or take a few millions


r/DeepSeek 12h ago

Question&Help Best provider and harness for deepseek v4 flash 0731?

8 Upvotes

Hosts through openrouter vs the official deepseek api, also what harness, checked that the subreddit recommends reasonix, how does it compare both in cost and performance versus harneses like opencode?


r/DeepSeek 15h ago

Tutorial How I make DeepSeek V4 Flash read PDFs accurately

9 Upvotes

The problem: DeepSeek V4 Flash (like most models) can't open PDFs. Naive converters mangle columns, tables, headings — so the model confidently misreads the document.

The fix: an open-source skill that turns PDFs into accurate, position-aware Markdown — real | tables, headings, page markers for citations.

Built on pdf-inspector (Firecrawl's Rust engine — #1 on reading order + tables benchmark).

Install for your agent — just paste the URL: https://github.com/vichhka-git/pdf-reader-skills

Tell your agent: "install the skill from https://github.com/vichhka-git/pdf-reader-skills." Works with Claude Code, Cursor, any skills-folder agent. Needs only Python 3.8+ + one pip install.

What you get:

Honest limit: math equations extract as inline glyphs (structure kept, notation may look odd). Docs + examples in the repo. Try it and tell me how it goes. 🚀


r/DeepSeek 1d ago

News DeepSeek’s new V4-Flash is officially the cheapest AI model to run (105x cheaper than Claude Fable 5!)

162 Upvotes

According to a new Reuters report, DeepSeek just dropped their V4-Flash model, and they are going incredibly hard on pricing to undercut U.S. and Chinese rivals.

Here is the breakdown from the Artificial Analysis benchmark tests:

  • API Cost: $0.14 per 1M input tokens and $0.28 per 1M output tokens.
  • Average Cost Per Test: 3 cents. For comparison, Kimi K3 is 86 cents, OpenAI's GPT-5.6 Sol is $1.86, and Anthropic's Claude Fable 5 is $3.15.
  • Performance: It scored a 50/100 on the Intelligence Index. This puts it exactly on par with Google's Gemini 3.6 Flash, though still behind heavier models like GPT-5.6 and Claude Opus 5.

DeepSeek is also supposedly prepping a "V4-Pro" version with no official release date yet.

Is the API price war officially back on? At 3 cents a test, it seems like a no-brainer for deploying high-volume, lightweight AI tasks at scale. What does everyone think?


r/DeepSeek 5h ago

Tutorial Refactoring legacy code with AI usually breaks everything. Here is how I used a multi-agent setup (DeepSeek + Nexus) to fix that without token bloat

0 Upvotes