A four-system NVIDIA DGX Spark setup running DeepSeek V4.1-Flash was reported to deliver a peak 494 tokens per second on code with 32 concurrent requests. That is a workload-level result; the reported single-request code rate was about 96 tokens per second.
What the reported rates measure
Concurrent requests are separate prompts being handled at the same time. The 494 tokens/s figure belongs to a 32-request code workload, while the single-request figure describes a different condition. They answer different performance questions: aggregate output under concurrent work versus the rate reported for one request.
| Task | Reported output rate | Stated request condition |
| Code | 494 tokens/s | 32 concurrent requests |
| Prose | 280 tokens/s | 32 concurrent requests |
| Code | Approximately 96 tokens/s | Single request |
| Prose | Approximately 58 tokens/s | Single request |
The figures are specific to their listed tasks and request conditions. In particular, the 494 tokens/s peak should not be read as the rate for a single code request.
What DeepSeek says about V4.1-Flash
In its September 10, 2026 announcement, DeepSeek describes V4.1-Flash as a 552-billion-parameter Mixture-of-Experts (MoE) model. DeepSeek says 8 billion parameters are active during input processing and 16 billion during output generation. “Active parameters” refers to the portion of the model engaged for that stage; it is distinct from the model’s total parameter count.
Inside the four-system setup
The reported configuration used four NVIDIA DGX Spark systems, each described as having 128 GB of unified memory. Together, they were reported to provide approximately 512 GB of pooled unified memory. The model’s reported memory footprint was 476 GB—a separate figure from the setup’s pooled capacity. The systems were described as connected in a switchless QSFP ring.
For broader background, NeoTeo has an overview of NVIDIA DGX Spark as a local AI system. Running a model on local hardware is a different access route from using an API; NeoTeo has also explained DeepSeek V4.1-Flash API pricing and cache.