The Call Center Doctors said it tested DeepSeek V4.1 Flash on a rented server with four Nvidia H200 GPUs on September 27, 2026. For a fully utilized day of its modeled coding workload, the consultancy estimated $184–$223 in equivalent DeepSeek API charges, compared with $440.88 for on-demand server rental.

The GPU test lasted about three hours. The September 1–27 dates refer to the consultancy’s broader Claude Code workload records, which it used to model the comparison.

The daily costs use different billing models

The on-demand rental cost $18.37 per hour, whether the server was busy or idle. A spot rental cost $9.19 per hour, or about $220 per day, but the provider could reclaim that capacity. The API figure is an estimate of usage charges for the modeled work, not a fixed server rental.

OptionCost per dayCost basisCondition
Four-H200 on-demand rental$440.88$18.37 per hourCharged whether the server was busy or idle
Four-H200 spot rentalAbout $220$9.19 per hourThe provider could reclaim capacity
DeepSeek API-equivalent workEstimated $184–$223API charges for the modeled workloadEstimate assumes a fully utilized server day

At the on-demand rate, the rental cost more than the estimated API usage for this workload. The spot rate was close to the API estimate, with the trade-off that the provider could reclaim the server. The consultancy put its roughly three-hour test at about $55 on-demand or $28 at the spot rate.

Repeated context drove much of the token volume

The September 1–27 records covered 388.5 billion tokens read and 393 million written. Of the tokens read, 374.2 billion were cached rereads—about 96.3%. That pattern matters because DeepSeek charges different rates for cached and uncached input.

DeepSeek’s listed off-peak rates for DeepSeek-V4.1-Flash are $0.003 per million cached input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens. Weekday peak rates are double. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. For more on how cache hits affect the API bill, see NeoTeo’s earlier explanation of DeepSeek V4.1 Flash API pricing.

The consultancy’s estimated API cost for that September workload ranged from $3,500 to $7,000, depending on peak pricing, with an estimated average of about $4,200. Its Claude Code subscription bill for the same period was about $5,500. Those amounts use different billing models: the API figure is estimated token usage, while the Claude Code figure is subscription spending.

DeepSeek reviewed code but did not write it

The consultancy said sandbox-escape concerns kept its DeepSeek builder agents offline. It used 48–64 read-only reviewer agents instead; they read 2,377 folders and filed 32 bug reports. DeepSeek’s builder agents shipped no code in the test, so the trial did not show it replacing the consultancy’s Claude Code workflow for writing and merging changes.

The reported concern included a possible route involving a settings file in a shared temporary folder that could allow agent-written code to run with administrator privileges. The consultancy’s account describes a potential escape path, not a successful exploit.

What the “80x cheaper” framing measures

The “80x cheaper” framing refers to a comparison of per-token list prices. It does not describe an equivalent subscription saving or the cost of producing the same completed coding work. The consultancy’s own comparison puts the September API estimate between $3,500 and $7,000, against about $5,500 in Claude Code subscription spending.