Docs/Models/DeepSeek-V4-Flash
Models

Run DeepSeek-V4-Flash on a DGX Spark: measured tok/s

Samuel Seidel · Published September 7, 2026

DeepSeek-V4-Flash, 304B total parameters, does not fit on one 128 GB Spark at any standard quantization. It has been squeezed onto a single GB10 with 2-bit expert planes and a patched vLLM fork, reported at 14 tok/s on prose and 21 tok/s on code, single-stream, with roughly 753 tok/s of prefill. GPUwerk has not run this configuration itself and treats it as a research curiosity, not something to build a service on.

Reported, not GPUwerk-measured

Engine Quantization Decode, single user Prefill
patched vLLM fork 2-bit experts 14 to 21 tok/s ~753 tok/s

From the model's own Hugging Face discussion thread, quoted on GPUwerk's benchmarks page, updated August 26, 2026. The thread does not state an exact run date or hardware other than a single GB10, and no concurrency sweep is published. The fork is not stock vLLM, so nothing here is reproducible with the standard vllm/vllm-openai container GPUwerk uses elsewhere on this site.

Memory footprint

Not measured in gigabytes on GPUwerk's pages. What GPUwerk's models page does establish: DeepSeek-V4-Flash is 304B total parameters, placed in its "what does not fit" section alongside multi-node-only models like DeepSeek-V4-Pro. Fitting it on one 128 GB unit at all required dropping to 2-bit expert planes, well below the 4-bit floor that GPUwerk's own quantization guidance calls the point where quality starts to visibly degrade on most models. No total weight figure for the 2-bit build is given, and none is estimated here rather than guessed.

How to run it

GPUwerk does not publish a run command for this configuration, because it depends on a patched, non-stock vLLM fork rather than the vllm/vllm-openai container used throughout the vLLM doc. Reproducing the reported numbers means locating that fork from the model's discussion thread and building it yourself; there is no docker run line to copy here in good conscience, because none has been verified by GPUwerk.

If you want a supported path to a large DeepSeek model, that means more than one Spark, which is outside what a single-node setup in the vLLM doc covers.

When this model is the wrong choice

Almost always, on this hardware. 2-bit quantization on a model this size is a research configuration meant to prove the model can be squeezed onto one GB10, not a deployment a customer should be served from. A patched fork with no stated maintenance status is also an operational liability: no upgrade path, no guarantee the next vLLM release keeps compatibility, no vendor to file a bug against.

If the workload genuinely needs 300B-class quality, the honest answer on a single Spark is that it does not fit. Either scale to more than one machine, or pick from the models on GPUwerk's models page that fit at 4-bit or better, such as Nemotron-3-Super-120B at NVFP4, which has a GPUwerk-adjacent measured decode number and no fork required.

Try it

A Spark is $0.79/hour, billed per minute against a prepaid balance, on the pricing page, cheap enough to spend an hour confirming this model is not what you want before reaching for a properly supported one instead. If the workload is real and needs this scale, GPUwerk's first engagement is the place to talk through a multi-node setup rather than a 2-bit workaround.

FAQ

Does DeepSeek-V4-Flash fit on one DGX Spark?

Only with a patched vLLM fork and 2-bit expert planes, not with stock vLLM or a standard quantization format. At 304B total parameters it is well beyond what a single 128 GB unit runs at any quality-preserving quantization.

How fast is DeepSeek-V4-Flash on a Spark?

14 tok/s on prose and 21 tok/s on code, single-stream, with roughly 753 tok/s of prefill, reported in the model's own Hugging Face discussion thread using the patched fork. GPUwerk has not reproduced these numbers on its own fleet.

Should I run DeepSeek-V4-Flash in production on a single Spark?

No. It is a research configuration built to prove the model can be squeezed onto one GB10 at 2-bit, not a supported deployment path. For a 300B-class model in production, use more than one machine.

See alsoFull benchmark numbers across engines See alsoWhich models fit in 128 GB

Run it yourself, not our numbers.

A dedicated Spark, deployed in minutes, with $20 in credit for your first $10 top-up.

Deploy a Spark