Running vision and multimodal models on a DGX Spark
Everything GPUwerk has published about the Spark so far is text-only: chat, coding, batch generation. This page is upfront about that gap and stays qualitative, because there's no measured vision-model number on this site to cite yet.
The capacity argument still applies
The reason GPUwerk leads with 128 GB of unified memory for text models, that it fits 70B-class models a typical consumer GPU's 24 to 32 GB of VRAM can't hold, per the hardware page, applies the same way to vision-language models. A VLM combines a language model backbone with a vision encoder, and the combined weight footprint of a large open VLM is usually well inside what 128 GB accommodates, with room left for a KV cache and a reasonable batch of images. That's a capacity statement, and it's the honest limit of what can be said without a measured number: it tells you the model will load, not how fast it will run.
What decode speed probably looks like, and what's unverified
GPUwerk's benchmarks page documents a formula for text decode speed: tokens per second is roughly memory bandwidth (273 GB/s on a Spark) divided by active parameter bytes per token. A VLM's text-generation phase, once an image has been encoded into tokens, is still decode in the same sense, so the same bandwidth ceiling should govern it the same way it governs the text-only models on that page. What that formula doesn't cover is the image encoding step itself: turning pixels into tokens the language model can consume is a different, largely compute-bound operation, closer in character to prefill than to decode, and GPUwerk has no measured figure for how long that takes on a Spark for any specific vision encoder. Until there's a run to point at, any number here would be a guess dressed up as a fact, and this page isn't going to do that.
What's plausible to build
Qualitatively, and consistent with what unified memory capacity and the bandwidth ceiling above imply: a single-user image-understanding assistant, describing a photo, reading a chart, answering a question about a screenshot, is the kind of workload the Spark's memory headroom is built for, the same "fits when a consumer GPU doesn't" argument the site already makes for large text chat models. High-throughput video or batch image-captioning at scale is a different question entirely, one this page can't answer, because it depends on encoder speed and batch behavior that haven't been measured here. If that's the target workload, benchmark it yourself before committing to it.
Where to actually get numbers
The vision-language models worth testing (and their exact memory footprints, quantized and unquantized) are documented on their own model or Hugging Face pages, not on GPUwerk's site. Pull the specific model's published weight size, check it against the 128 GB ceiling, and run vllm bench serve or the model's own inference harness on a rented node before assuming a speed number. GPUwerk's benchmark methodology page covers how to run that kind of test and log it usefully, which is exactly the exercise this page is deferring to you rather than guessing at.
A first engagement can help scope and benchmark a specific vision or multimodal workload if you'd rather not run that exercise solo.