Docs/First evaluation
Getting started

Your first model evaluation on a rented DGX Spark

Updated September 4, 2026 · 5 min read

Start with one model and a concrete question: does it produce useful answers fast enough for your application? This guide uses the preconfigured Qwen3 Coder 30B image, so you can test the hardware without installing a serving stack.

Sign in and configure See current pricing

1. Choose a small, measurable test

Prepare ten public or synthetic prompts representative of your workload. Write down what a correct answer looks like and your acceptable response time. Start with one request at a time; a coding-model test does not establish suitability for every document or chat workload.

2. Plan your spend before adding credit

The current single-Spark rate is $0.79/hour before applicable tax, billed per minute. One hour of running time costs $0.79; eight hours costs $6.32. Loading and testing take time, so keep enough credit for the whole session. A top-up is prepaid credit, not a booking or a reservation of capacity.

A stopped Spark stays reserved and continues billing at 75% of its running rate, rounded to $0.59/hour for this tier. Termination ends compute billing and deletes the workspace. Zero credit or a reached budget stops the instance and releases the machine; the workspace is saved for 7 days, then deleted. Export your results while you still have access.

3. Configure one Spark and your SSH key

  1. Open the console and sign in. Add prepaid credit when you are ready; your balance does not reserve capacity.
  2. Add your public SSH key. Keep the private key on your computer. See the SSH setup guide.
  3. Select one Spark and Qwen3 Coder 30B (vLLM). Review the displayed billing conditions and live availability before deploying.
  4. Deploy, wait for readiness, and copy the SSH connection command from the instance details.

The base OS images do not start this model automatically. For another model or runtime, use the vLLM guide or llama.cpp guide.

4. Run a first request inside the SSH session

After connecting, run these commands on the rented instance. The preconfigured image serves on port 8888. First check the model list:

curl --fail --max-time 30 http://127.0.0.1:8888/v1/models

For the Qwen3 Coder image, the configured model name is qwen3-coder-30b. Then save a small response:

curl --fail --max-time 180 http://127.0.0.1:8888/v1/chat/completions   -H 'Content-Type: application/json'   -d '{"model":"qwen3-coder-30b","messages":[{"role":"user","content":"Write a Python function that returns the unique items in a list, preserving order."}],"max_tokens":256,"temperature":0}'   -o /workspace/first-response.json
cat /workspace/first-response.json

If the connection is refused, the model may still be loading. Check instance status and retry after it is ready. A timeout is not a reason to launch another paid instance. If the model list differs, use its returned model ID.

The image also exposes a model endpoint through its instance hostname. Calling localhost over SSH does not disable that public endpoint. Use public or synthetic inputs for this first test; review endpoint access considerations before using confidential material.

5. Compare useful answers, not just token speed

Run the same prompts against your current model or provider. Record correctness, response time, prompt/output lengths, model configuration and total compute cost. Keep concurrency consistent. Our benchmark results are reference workloads, not a guarantee for your prompts.

6. Export, then terminate

Copy /workspace/first-response.json to your computer before terminating. With an SSH alias configured as described in the SSH guide, run this on your own computer, replacing my-spark with your alias:

scp my-spark:/workspace/first-response.json ./first-response.json

Open the local file to verify the copy. Then terminate in the console and confirm the instance is terminated. Stopping is a paid reservation, so it does not finish the billing step. GPUwerk does not keep a backup of your workspace.

Need help choosing the test?

Tell us the model, approximate context length, expected concurrent users and what you need to establish. Do not send confidential datasets in the first email.

Discuss your workload