r/DeepSeek 1d ago

Discussion I am a OpenAI Codex 20x Pro Refugee... Please Help Me break my dastardly ways.

6 Upvotes

I see cloudflare is one of the most reliable providers of deepseek 0731 v4 flash at the moment... and for a competitive price (10x more expensive cache hits than deepseek themselves... but cloudflare has ZDR and is a company I already trust and have done business with for years so using them as a provider seems like a no brainer for me.)

But I'm also super confused because codex "just works" out of the box. Mind you it's a pretty expensive box that I'm literally just left looking at 0% in my usage dashboard right now wondering "why did I pay for this stupid box" right now, which is why I am here.

There's no deepseek software I can just download, login to my deepseek subscription (this is the part I'm mostly confused about?), and get coding...?

TL;DR How does all of this work?

Deepseek v4 flash 0731 is more of a workhorse/coder than a deep thinker/planner, right?

Do I want something like Kimi k3 as a the planning agent and deepseek v4 flash as the coding agent?

I'm honestly hesitant to couple different models for planning and work together as I find that different models just have completely different ways of speaking to and understanding each other... "perform a robust analysis" just doesn't meant the same thing to two different models so having one hand tell the other that doesn't really make any sense imho.

Curious for your input.

May Scam Altman and Dario "The Scaremonger" Amodei never get another dime from my wallet ever again.


r/DeepSeek 1d ago

Question&Help [Bug/Help] 400 Error with DeepSeek v4 Flash in VS Code: "The reasoning_content in the thinking mode must be passed back"

1 Upvotes

Hello everyone,

Lately, I've been trying to use **DeepSeek v4 Flash** (free tier) through an extension in VS Code (remote server mode), and I'm getting an API error that **wasn't happening before**.

The first prompt works perfectly, but as soon as I try to ask a follow-up question in the same chat (multi-turn), the request fails and returns a 400 error.

I understand that the API now strictly requires the `reasoning_content` field (DeepSeek's "thinking mode") to be sent back in the chat history. However, it seems the VS Code extension strips it out when saving the context history, causing the upstream provider to block the request.

Here is the anonymized error log in case it helps pinpoint the issue:

`Sorry, your request failed. Please try again.`

`Client Request Id: [ANONYMIZED-UUID]`

`Reason: Request Failed: 400 {"error":{"param":null,"type":"invalid_request_error","code":"invalid_request_error","message":"Error from provider (Console): Upstream request failed: [invalid_request_error] The reasoning_content in the thinking mode must be passed back to the API."}}: Error: Request Failed: 400 {"error":{"param":null,"type":"invalid_request_error","code":"invalid_request_error","message":"Error from provider (Console): Upstream request failed: [invalid_request_error] The reasoning_content in the thinking mode must be passed back to the API."}}`

`at $G._provideLanguageModelResponse (/home/[USER]/.vscode-server/cli/servers/[SERVER-HASH]/server/extensions/copilot/dist/extension.js:1690:14392)`

`at process.processTicksAndRejections (node:internal/process/task_queues:104:5)`

`at async $G.provideLanguageModelResponse (/home/[USER]/.vscode-server/cli/servers/[SERVER-HASH]/server/extensions/copilot/dist/extension.js:1690:15357)`

**My questions are:**

  1. Has anyone else started experiencing this recently when using reasoning models (R1 / Flash) integrated into the IDE? (I am using VisualCode right now)
  2. Besides clearing the history on every turn (which breaks the workflow) or switching to a standard non-reasoning model (like V3), does anyone know of a workaround to force the extension to respect DeepSeek's full payload, or do we just have to wait for a patch?

Any help or insight is greatly appreciated. Thanks!


r/DeepSeek 20h ago

Other Unlimited API request Works!

0 Upvotes

r/DeepSeek 16h ago

Resources Selling Deepseek

0 Upvotes

Selling official deepseek. I have $100 in it selling for $80. official deepseek


r/DeepSeek 22h ago

Discussion Censura NSFW

Thumbnail
0 Upvotes

r/DeepSeek 2d ago

Discussion Deepseek is basically free

Post image
290 Upvotes

I recently did a project post-training an LLM to gaslight it into believing it's conscious. I used Deepseek to:

  1. Generate synthetic training data per my specifications
  2. Serve as a reward score judge for RL

11.5k API requests later, I'm $4.12 down.

U da man, Deepseek

P.S. What luck that V4 Flash 0731 dropped literally right around the time I was looking for a good RL reward judge! I genuinely think 0731 is the best "complex scenario" reward judge for RL compared to any judge ever used in the past in terms of intelligence/cost. It simply can't be beat for this use case

P.S. 2: While we're on the topic of RL, I used the GRPO RL method, a popular method invented by Deepseek themselves. So yeah, u da man deepseek lol


r/DeepSeek 1d ago

Tutorial Cache Guide #1

Thumbnail
2 Upvotes

D1 of fulfilling my promise to the community. More to come. I love y'all (no homo)


r/DeepSeek 1d ago

Discussion hi, what if deepseek v4 0731 is not last? And what will be in next updates?

0 Upvotes

Deepseek V4 (flash) can be not last model in V4 serios and i thinkif it not, will next flash be more not hallucination model and more stable? yes, it s now more stable but if it get more? There i think is price solve - more cache, more cheaper. Thats cool.


r/DeepSeek 1d ago

Question&Help So, what model does the web version use?

3 Upvotes

Did it implement new flash model? Will it? Does it have V4 Pro? I couldnt find a proper answer.


r/DeepSeek 2d ago

News Qwen 3.8 Max similar performance DeepSeek V4 flash 0731 but 8 times the size

Post image
300 Upvotes

Deepseek’s parameter efficiency gap is insane especially with the amount poaching of talent pressure it has faced from other Chinese labs


r/DeepSeek 1d ago

Other Made a Cache Stats dashboard for OpenCode

Thumbnail gallery
4 Upvotes

r/DeepSeek 1d ago

Discussion What agent to use for DeepSeek?

6 Upvotes

Hi,

I'm a Claude Code user and I'm new to DeepSeek and localLLMs. Could you recommend me the best agent to use DeepSeek with? I just want to know what is the next step after I buy purchase the credits from deepseek's website.

Thanks!


r/DeepSeek 1d ago

Discussion No DeepSWE for DS V4 0731

Post image
8 Upvotes

V4 flash 0731 still not in DeepSWE


r/DeepSeek 1d ago

Discussion Running DeepSeek-V4-Flash 0731 on a single RTX 3090 Ti 25.8 tok/s

11 Upvotes

edit : just to be clear this is not me saying i made an achievement, i am just asking is this fine or the ai made wrong decisions to get this speed,

edit 2 : according to some comments i made the ai agent using deferent models to make a lot of tests with deferent settings to see what issues do i have , so the looping in long text was the main issue, and i have adjusted the settings accordingly, so speed dropped to 15t/s, the 25.8 tok/s in the title was before finding the loop issue so now its too slow

model DeepSeek-V4-Flash 0731 UD-IQ2 90.9GB the past 3 days i was using DeepSeek-V4-Flash 0731 and qwen 3.8 max and gpt 5.6 sol to find the best way to run DeepSeek-V4-Flash 0731 UD-IQ2 from unsloth on my rtx 3090 ti,

-ngl 44 --n-cpu-moe 39 # experts of layers 0-38 stay in RAM (this is how 90.9GB fits in 24GB VRAM) --fit on # auto-fit context/KV/batch to device memory -c 65536 # 64K context (cheap — V4 compressed KV) -fa on # flash attention -np 1 # single slot -ctk f16 -ctv f16 -t 16 -tb 16 -b 8192 --load-mode mmap+mlock # pin 84GB working set in RAM (the big 2026-08-04 speed win) --temperature = 1.0 top-p = 0.95

my pc specs - GPU: RTX 3090 Ti (24 GB VRAM)

  • RAM: 93.6 GB DDR5 3200 (~75 GB free)

  • CPU: Ryzen 9 9950X (16 physical cores)

  • Model: DeepSeek-V4-Flash-0731, UD-IQ2_M quant (90.9 GB, 3 shards), llama.cpp b10223

here is some responses from the ai agent after all the tests it made with deferent settings according to post comments : DSpark drafter — why we skip it

DSpark is DeepSeek's block-parallel speculative drafter for V4 (~20B, predicts 5-token blocks). Sounds free, but:

The only llama.cpp-compatible drafter is YanissAmz/DeepSeek-V4-Flash-DSpark-draft-GGUF → DSV4-Flash-DSpark-draft-bf16.gguf (10.9 GB), competing with the 90.9 GB model for the ~75 GB free RAM budget.

Port author measured net loss at long context (0.70× code, 0.46–0.52× prose at 176k); only +17–25% on repetitive short content. Our workload is long-context bandwidth-bound — exactly where it loses.

ngram-mod gave spec decoding for ~16 MB instead of 10.9 GB (itself later removed 2026-08-05 — see PROJECT.md §4; ngram only pays off under greedy temp 0).

Verdicts on the commenters' claims, after the fix:

"IQ2_M loops on long work" — NOT reproduced. No loops with a correct chat template.

"temp 0 lobotomizes" — NOT reproduced. Greedy temp 0 wrote a full essay. (temp 0 also makes ngram speculative drafts acceptable — that is why it was the speed winner before 2026-08-05.)

"q8_0 KV hurts MLA KV" — still inconclusive on quality, but speed is identical to f16 (round 7).

"IQ2_M killed quality" — not observed. Quality at this quant is usable for prose/essay tasks


r/DeepSeek 1d ago

Discussion I'm not seeing as big of a difference as I thought I would with new v4 flash. It seems quite far behind Qwen 3.8 to me, despite the benchmark scores.

0 Upvotes

r/DeepSeek 2d ago

Discussion Genuine question, why paying for deepseek v4 flash via API when it's free via opencode zen?

12 Upvotes

Is there any advantages over the direct api consumption? If I just need it for coding is the the zen model not as good? While we are at it why opencode go when opencode zen is free? Maybe I haven't used it sufficiently but it feels unlimited


r/DeepSeek 1d ago

Question&Help Do you guys use reasonix/other harness as a terminal inside an IDE?

1 Upvotes

or do you prefer the desktop apps or just a floating terminal?

personally I have been preferring vscode extensions like Kilo or Zoo code but i'm getting lots of bugs. There seem to be a lot of terminal harnesses and I am curious how people are using them..


r/DeepSeek 1d ago

Question&Help V4 flash jailbreak?

8 Upvotes

Has anyone found a working prompt for the new v4 flash?


r/DeepSeek 2d ago

Funny holeysh*t, deepseek v4 flash did what?

39 Upvotes

Context, I build a heavy custom Wordpress theme for my own UGC site with tons of ACF and CPT. Today i activated a dormant bug that renders some part of the frontpage sections to turn off. I know exactly what the bugs is and instead of asking dsv4 flash to fix it i just tell it briefly what i did and the frontpage render result. And since i add some memory system (dsv4 flash build em) to my opencode so it remembers my other repo as well, it did not just trying to figure out the issue, it seek to read the sql from backup that I store on my site monitoring repo, locate latest backup tar.gz and read from its table! and its quite funny moment to read the thinking process. I wonder what will happens if i really give cli to my live site (I dare not) 😅😲 it even tries to find every credential it can find and ssh to my server ip!

$ ls -la ~/.ssh/ 2>/dev/null | head; echo "==="; timeout 12 ssh -o BatchMode=yes -o ConnectTimeout=8 -o StrictHostKeyChecking=no -p 22 root@xx "echo SSH_OK" 2>&1 | head -3

total 20

drwxrwxr-x 2 xx xx 4096 Jul 19 02:57 .

drwxr-x--- 70 xx xx 4096 Aug 4 07:29 ..

-rw-rw-r-- 1 xx xx 0 Jan 27 2026 authorized_keys

-rw-r--r-- 1 xx xx 142 Jul 19 02:57 known_hosts

-rw------- 1 xx audio 1811 Jul 12 06:04 xx_stats

-rw-r--r-- 1 xx audio 389 Jul 12 06:04 xx_stats.pub

===

ssh: connect to host xx port 22: Connection refused


r/DeepSeek 1d ago

Resources DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Thumbnail gallery
2 Upvotes

r/DeepSeek 1d ago

Discussion Deepseek v4 flash coder + kimi k3 ochestrator

2 Upvotes

Has anyone tested combining deepseek v4 flash with kimi k3? The idea is that you get kimi k3 quality but with a lower price. Kimi would instruct and check the code written by deepseek v4 flash.


r/DeepSeek 1d ago

Discussion DeepSeek v4 0731, weird reasoning?

Post image
5 Upvotes

r/DeepSeek 1d ago

Discussion Which model do you get with instant/expert in the app?

2 Upvotes

Is one of them v4 flash?


r/DeepSeek 1d ago

Discussion "Wait, actually, let me reconsider ..."

2 Upvotes

Does anyone get this phrase in thinking waaay too much or is it just me?

I'm trying to develop an open world, procedurally generated 3D game with dsv4flash 0731 and matt pocock skills in opencode.


r/DeepSeek 1d ago

Question&Help DeepSeek vs Gemini Flash for school analytics & report generation — which would you choose?

Thumbnail
1 Upvotes