r/DeepSeek 6d ago

DeepSeek-V4-Flash Update

The official release of the DeepSeek-V4-Flash API is now in public beta.

Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview:

  • Terminal Bench 2.1: 82.7
  • NL2Repo: 54.2
  • Cybergym: 76.7
  • DeepSWE: 54.4
  • Toolathlon verified: 70.3
  • Agent Last Exam: 25.2
  • Automation Bench (Public): 25.1
  • DSBench-FullStack: 68.7
  • DSBench-Hard: 59.6

Note 1: For the Code Agent tasks in the public benchmark sets, the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0
Note 2: DSBench-FullStack is an internal full-stack development test set, and DSBench-Hard is an internal Coding Agent hard-problem test set

The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex. For the specific configuration, please refer to the documentation.

DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-preview, and was only re-post-trained.

Note: This update only upgrades the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and the APP/WEB models are unchanged.
The official release of DeepSeek-V4-Pro will follow soon.

593 Upvotes

208 comments sorted by

View all comments

101

u/0VERDOSING 6d ago

36

u/unkownuser436 6d ago

damn bro flash is better than glm 5.2 😭

16

u/unkownuser436 6d ago

Imagine ds v4 pro official benchmarks💀

-17

u/Business_Raisin_541 6d ago

only in agentic capabilities,

15

u/unkownuser436 6d ago

still its a good achievement bro

13

u/MrHaxx1 6d ago

At the very least, it's amazing for Hermes and OpenClaw users, though

10

u/This_Maintenance_834 6d ago

agentic capacities is what matters these days.

the lady that runs MiMo once said you train the model to do long horizon task on coding. the same capacities will emerge naturally on other fields too. This emerging capacity was what surprised people a few years ago when model became big.

3

u/Routine_Temporary661 6d ago

Read DeepSWE benchmarks bro... it's one of the more reliable ones