Run DeepSeek-V4-Flash on a DGX Spark: measured tok/s
DeepSeek-V4-Flash, 304B total parameters, does not fit on one 128 GB Spark at any standard quantization. It has been squeezed onto a single GB10 with 2-bit expert planes and a patched vLLM fork, reported at 14 tok/s on prose and 21 tok/s on code, single-stream, with roughly 753 tok/s of prefill. GPUwerk has not run this configuration itself and treats it as a research curiosity, not something to build a service on.
Reported, not GPUwerk-measured
| Engine | Quantization | Decode, single user | Prefill |
|---|---|---|---|
| patched vLLM fork | 2-bit experts | 14 to 21 tok/s | ~753 tok/s |
From the model's own Hugging Face discussion thread, quoted on GPUwerk's benchmarks page, updated August 26, 2026. The thread does not state an exact run date or hardware other than a single GB10, and no concurrency sweep is published. The fork is not stock vLLM, so nothing here is reproducible with the standard vllm/vllm-openai container GPUwerk uses elsewhere on this site.
Memory footprint
Not measured in gigabytes on GPUwerk's pages. What GPUwerk's models page does establish: DeepSeek-V4-Flash is 304B total parameters, placed in its "what does not fit" section alongside multi-node-only models like DeepSeek-V4-Pro. Fitting it on one 128 GB unit at all required dropping to 2-bit expert planes, well below the 4-bit floor that GPUwerk's own quantization guidance calls the point where quality starts to visibly degrade on most models. No total weight figure for the 2-bit build is given, and none is estimated here rather than guessed.
How to run it
GPUwerk does not publish a run command for this configuration, because it depends on a patched, non-stock vLLM fork rather than the vllm/vllm-openai container used throughout the vLLM doc. Reproducing the reported numbers means locating that fork from the model's discussion thread and building it yourself; there is no docker run line to copy here in good conscience, because none has been verified by GPUwerk.
If you want a supported path to a large DeepSeek model, that means more than one Spark, which is outside what a single-node setup in the vLLM doc covers.
When this model is the wrong choice
Almost always, on this hardware. 2-bit quantization on a model this size is a research configuration meant to prove the model can be squeezed onto one GB10, not a deployment a customer should be served from. A patched fork with no stated maintenance status is also an operational liability: no upgrade path, no guarantee the next vLLM release keeps compatibility, no vendor to file a bug against.
If the workload genuinely needs 300B-class quality, the honest answer on a single Spark is that it does not fit. Either scale to more than one machine, or pick from the models on GPUwerk's models page that fit at 4-bit or better, such as Nemotron-3-Super-120B at NVFP4, which has a GPUwerk-adjacent measured decode number and no fork required.
Try it
A Spark is $0.79/hour, billed per minute against a prepaid balance, on the pricing page, cheap enough to spend an hour confirming this model is not what you want before reaching for a properly supported one instead. If the workload is real and needs this scale, GPUwerk's first engagement is the place to talk through a multi-node setup rather than a 2-bit workaround.
FAQ
Does DeepSeek-V4-Flash fit on one DGX Spark?
Only with a patched vLLM fork and 2-bit expert planes, not with stock vLLM or a standard quantization format. At 304B total parameters it is well beyond what a single 128 GB unit runs at any quality-preserving quantization.
How fast is DeepSeek-V4-Flash on a Spark?
14 tok/s on prose and 21 tok/s on code, single-stream, with roughly 753 tok/s of prefill, reported in the model's own Hugging Face discussion thread using the patched fork. GPUwerk has not reproduced these numbers on its own fleet.
Should I run DeepSeek-V4-Flash in production on a single Spark?
No. It is a research configuration built to prove the model can be squeezed onto one GB10 at 2-bit, not a supported deployment path. For a 300B-class model in production, use more than one machine.