Docs/Models/GLM-5.3-Flash
Models

Run GLM-5.3-Flash on a DGX Spark: quantization, context and memory

Samuel Seidel · Published September 14, 2026

GLM-5.3-Flash is a 320B-A18B mixture-of-experts model that GPUwerk serves as the idle model filling spare capacity on its own Sparks. This is a configuration note about the quantization and context choice behind that deployment, not a benchmark. GPUwerk has not published a tokens-per-second figure for this model, and none appears below.

On this page
  1. How GPUwerk runs it
  2. The UD-Q2_K_XL to UD-IQ1_S switch
  3. What a Dynamic GGUF quant is
  4. Observed memory headroom
  5. Why there is no tok/s number here
  6. FAQ

How GPUwerk runs it

GLM-5.3-Flash fills the idle and rental-spare role across GPUwerk's Sparks: a model kept resident to use capacity that would otherwise sit empty between customer workloads, served through llama.cpp. As of 13 September 2026 it runs the UD-IQ1_S quantization at -c 262144 -np 1, the model's full native context length at one concurrent slot, with --no-mmap set, which the community reports as effectively required for stable behavior on the Spark's unified memory, per a llama.cpp GitHub discussion. The previous UD-Q2_K_XL preset stays staged on the fleet as a fast rollback rather than removed, since a sub-2-bit quant carries a real accuracy cost worth being able to reverse quickly.

The UD-Q2_K_XL to UD-IQ1_S switch

Before 13 September 2026, the idle preset ran UD-Q2_K_XL, loading at roughly 107 GB. On a Spark's unified memory, weights, KV cache and prompt-processing buffers all draw from the same pool the host uses, and GPUwerk's Sparks have run without swap since a memory-hardening pass on 9 September 2026. Under that constraint, UD-Q2_K_XL's footprint left too little headroom to fit the model's native 262,144-token context at all: a 131,072-token profile had already failed cold initialization on one node on 10 September 2026, and the steady-state configuration was capped at 65,536 tokens.

UD-IQ1_S loads at roughly 93 GB, about 15 GB smaller, and that freed headroom is what makes -c 262144 -np 1 fit. The tradeoff is real: a sub-2-bit quant carries a known coherence cost, especially reported for code on large MoE models, which is why the larger Q2_K_XL build is kept staged rather than deleted.

What a Dynamic GGUF quant is

UD-Q2_K_XL and UD-IQ1_S are both Unsloth Dynamic GGUF quantizations (accessed 14 September 2026). Rather than compressing every layer of a model to the same bit width, a dynamic quant assigns a different precision per layer, holding layers judged more sensitive to quantization error at higher precision while compressing the rest more aggressively; Unsloth's own documentation describes the method as dynamically adjusting "the quantization type of every possible layer," with the resulting mix differing by model.

In general terms, the tradeoff runs one direction: a smaller quant frees memory, which on a fixed-memory machine like a Spark can buy back context length or concurrency, at the cost of accuracy on the model's own outputs. Unsloth's documentation is explicit that its smallest, roughly 1-bit tier should not be used for agentic workloads because tool-calling and reasoning quality degrade at that compression level. GPUwerk has not run an accuracy evaluation of its own IQ1_S deployment against that caution; it is stated here as the general risk the memory-for-context tradeoff carries, not as a GPUwerk finding about this specific configuration.

Observed memory headroom

At a 131,072-token context with the host's llama-server prompt cache disabled, one Spark showed roughly 0.7 to 1.4 GiB of MemAvailable while Docker itself accounted for only about 4.6 GiB, with zero memory pressure recorded at the moment of inspection and KV-cache capacity retries appearing in the logs. That is a tight operating envelope, observed under those specific conditions, not a load certification: a long-prompt or high-concurrency soak test at that configuration has not been run. The switch to UD-IQ1_S and -np 1 at the full 262,144-token context is a different, larger memory footprint on the KV cache side than that observation, and has not been soak-tested to the same standard either.

Why there is no tok/s number here

GPUwerk's own measured throughput figures exist for exactly two models, published with raw logs on the benchmarks page: Qwen3-Coder 30B-A3B AWQ and Llama 3.3 70B AWQ, both under vLLM. GLM-5.3-Flash has not been through that same measurement pass, and this page does not estimate a number for it. If a figure matters for your decision, run it yourself on a rented Spark rather than relying on the memory observations above, which describe headroom, not speed.

Try it

A Spark rents at $0.79/hour, billed per minute against a prepaid balance, on the pricing page. To serve your own GGUF quant with llama.cpp, follow the build and serve pattern in the llama.cpp doc; for background on what a quantization format trades off in general, see the quantization doc.

FAQ

What tok/s does GLM-5.3-Flash get on a DGX Spark?

GPUwerk has not published a throughput figure for GLM-5.3-Flash. It runs as the idle/spare model on GPUwerk's own Sparks, but this page is a configuration note about quantization and context, not a benchmark.

Why did GPUwerk switch GLM-5.3-Flash from UD-Q2_K_XL to UD-IQ1_S?

UD-Q2_K_XL's roughly 107 GB loaded footprint left too little unified-memory headroom, after GPUwerk's OOM hardening, to fit the model's native 262,144-token context at all; 131,072 context had already failed cold initialization on one node. UD-IQ1_S loads at roughly 93 GB, and the freed headroom is what makes -c 262144 -np 1 fit on 13 September 2026.

What is a Q2_K_XL or IQ1_S quant?

Both are Unsloth Dynamic GGUF quantizations, which assign a different bit-width per layer instead of one uniform precision across the whole model, aiming to keep the more sensitive layers at higher precision while compressing the rest harder. IQ1_S compresses further than Q2_K_XL, trading more memory for more accuracy loss.

See alsoQuantization formats explained See alsoServing with llama.cpp

Test your own quant, not ours.

A dedicated Spark, deployed in minutes, with $20 in credit for your first $10 top-up.

Deploy a Spark