Cost
DGX Spark/Cold start vs always-on: choosing an instance uptime pattern

Cold start vs always-on: choosing an instance uptime pattern

By Samuel Seidel · Published September 9, 2026

A rented Spark has three states, not two: running, stopped, and terminated, and each carries a different cost and a different guarantee about what's still there when you come back. Which pattern to use, stopping between sessions or leaving it running around the clock, comes down to how predictable and how latency-sensitive the actual usage is, not a general rule that applies to every deployment.

What each state actually costs and keeps

Per the pricing page, a running Spark bills the full $0.79/hour on-demand rate. Stopping an instance keeps it reserved, the same machine, same SSH port and hostname, workspace intact, at 75% of the running rate: $0.59/hour. Terminating ends billing entirely, but it also deletes the workspace from the node for good, with no copy kept, which is why the console asks you to type the instance name to confirm.

That middle state, stopped, is the one that makes a cold-start pattern viable at all. It's not free, a stopped Spark still holds its allocation and still costs something, but it's meaningfully cheaper than running while preserving everything on disk. If the pattern is closer to "start it, terminate it, start a fresh one next time," there's no state to preserve between sessions and every restart means re-fetching models and re-doing setup from scratch.

The case for stopping between uses

Stopping fits usage that's genuinely intermittent: a research or evaluation workload run for a few hours a day, a batch job that only needs the GPU during a nightly window, or a team still in early testing where the Spark sits idle most of the time. At 75% of the running rate while stopped, the savings scale directly with how much of the day the instance would otherwise sit idle but running. A Spark used eight hours a day and stopped the other sixteen costs noticeably less over a month than one left running continuously, even accounting for the stopped-hour charge.

The tradeoff is restart time. Model presets may need several minutes to warm up after a restart, per the pricing page's own FAQ, and that's after the instance itself comes back online. For an interactive workload where someone is waiting on the first response of the day, that warm-up is a real, if bounded, delay. For a batch or scheduled job, it's mostly irrelevant since nothing is waiting on it synchronously.

The case for always-on

Always-on fits the opposite profile: an endpoint other systems depend on, a chat interface teammates use throughout the day without warning, or any workload where a multi-minute cold start would be visible to a user at an unpredictable moment. Running continuously also avoids the small operational overhead of a stop/start routine, whether that's a person remembering to do it or a script that has to be maintained and monitored.

The cost tradeoff runs the other way here: 24 hours running costs $18.96 against $14.16 for 24 hours stopped, per the pricing page's worked example, a difference that compounds daily if the instance genuinely doesn't need to be up around the clock. Always-on is the more expensive default, and it's worth confirming that the workload actually needs it rather than defaulting to it out of convenience.

A rough way to decide

The deciding question is usually: does anything depend on this endpoint responding within seconds, at any hour, without advance notice? If yes, always-on is close to a requirement, not a preference. If the workload has a predictable active window, business hours, a nightly batch, a specific team's working hours in one timezone, stopping outside that window is close to free savings, since the warm-up delay lands at a time nobody's waiting synchronously anyway.

For a workload that's genuinely unpredictable, bursty usage with no clear pattern, the honest answer is that no uptime policy solves it cleanly on a single Spark. A scheduled stop/start based on observed usage patterns, revisited periodically, tends to work better than either extreme applied blindly.

Whichever pattern is chosen, terminating (not stopping) is the only state that fully ends billing, and it permanently deletes the workspace. Keep a backup of anything that matters before terminating; GPUwerk keeps no copy once an instance is gone.

For working through the actual monthly cost of either pattern against a specific usage estimate, see the cost forecasting guide. Full pricing detail, including the stopped-rate math, is on the pricing page. A first engagement can help model the right uptime pattern against a specific workload before committing to one.

First top-up: pay $10, get $20 in credit

Stop, don't terminate, and pick up right where you left off.

Deploy a Spark and choose the uptime pattern that matches actual usage, not a fixed monthly plan.

Deploy a Spark See pricing