Private LLM hosting vs AWS Bedrock
We rent DGX Sparks, so weigh that against everything below. If you need Claude, Amazon Nova, or another closed model in Bedrock's catalogue by name, this page cannot get you there, a Spark only runs open-weight models. If an open model already covers your task and the reason you're evaluating Bedrock is data location and who can read your prompts, a dedicated Spark in EU-Central gives you a shorter answer to both questions, because the machine is yours alone and nothing about it is shared with AWS, a model provider, or another tenant.
Side by side
| AWS Bedrock | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| Where data is processed | Whichever AWS region you pick for your Bedrock endpoint, including eu-central-1 (Frankfurt) and eu-west-1 (Ireland) for EU-only inference. Cross-region inference profiles, if you enable them, can route requests to a wider set of regions for latency or capacity reasons. | EU-Central (Prague), on one dedicated machine. No other region, no cross-region routing to configure because there's only the one machine. |
| Who can see your prompts | Not other customers, and AWS states inputs and outputs are not used to train the underlying foundation models or shared with model providers for that purpose. Requests still pass through AWS's managed service and its automated content filters, and AWS operational staff can access infrastructure under AWS's standard support and security procedures, the same as for any managed AWS service. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." There is no abuse-monitoring layer at all, because GPUwerk operates infrastructure, not a model API. |
| Model choice | Closed frontier models (Anthropic Claude, Amazon Nova) and open-weight models (Meta Llama, Mistral, and others) side by side through one API and one bill. | Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB. Claude and Nova are not available on a Spark, full stop, that is not a current limitation, it is the shape of the product. |
| Speed and concurrency | AWS does not publish a tokens/second figure for Bedrock models on its pricing or product pages; on-demand throughput is instead governed by per-account, per-model request and token quotas, and provisioned throughput is sold in fixed capacity units rather than a stated tok/s number. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). On our own fleet, Qwen3-Coder 30B-A3B AWQ decodes at 80.9 tok/s single-stream, and Llama 3.3 70B AWQ decodes at 6.0 tok/s single-stream but reaches 485 tok/s aggregate at 128 concurrent before its performance collapses past that point. Raw logs and methodology on our benchmarks page. |
| Pricing model | On-demand pricing is per token, billed separately for input and output and varying by model. Provisioned throughput is billed hourly for a fixed capacity commitment regardless of how much of it you use, similar in shape to a rented machine but sized in model-specific throughput units rather than GPU-hours. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model or how many requests you send. USD, tax extra, per pricing. |
| Worked cost example | See the arithmetic below the table. | |
| Contracts and DPA | Governed by the AWS GDPR Data Processing Addendum, incorporated into the AWS Customer Agreement. Enterprise Support and negotiated terms are typical at volume. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| Lock-in and portability | Bedrock's primary interface is its own Converse API and per-model invoke calls, not the OpenAI-compatible chat completions shape a Spark serves. Some Bedrock model endpoints also accept an OpenAI-compatible request shape for a subset of models, but Knowledge Bases, Guardrails, and Agents are Bedrock-only platform features with no equivalent to carry over. | Open weights mean you can move the exact model file to another Spark, another GPU box, or your own hardware; the model itself is never AWS's or GPUwerk's to withhold. You do lose Bedrock's managed platform layer (Guardrails, Knowledge Bases, Agents, provisioned scaling) and have to build any of that yourself. |
| What you operate yourself | Nothing at the infrastructure layer; AWS manages capacity, scaling, and model serving. You manage prompts, application code, and your own use of quota or provisioned throughput. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, content filtering, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
AWS Bedrock. On-demand pricing is per token, billed separately for input and output and varying by model, and provisioned throughput is billed hourly for a fixed capacity commitment sized in model-specific units. GPUwerk could not independently verify a current per-token or per-hour price for a comparable open-weight model (Llama or Mistral) on Bedrock from a source already cited in this repository, and did not fetch AWS's pricing page to produce a new number for this page. Run your own model and volume through AWS's own Bedrock pricing page for a number you can trust before comparing it against the Spark figure below.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back, with no idle time and enough real traffic to keep the queue full. Real traffic arrives in bursts, not a constant stream at exactly your saturation concurrency, so a Spark billed by the hour that sits half-idle waiting for requests can lose ground fast, you're paying $0.79 for every hour on the clock whether the GPU is generating tokens or not. Bedrock's on-demand per-token price does not care about idle time because there is no idle time to bill; provisioned throughput, on the other hand, is a fixed hourly commitment much like a Spark and carries the same underutilisation risk. Put your own number from AWS's Bedrock pricing page next to $127.19 for your volume, and the crossover point is a utilisation question specific to your traffic pattern, not a fixed answer either provider can give you.
Migration path
The two do not speak the same wire protocol by default. Bedrock's primary interface is its Converse API and per-model invoke calls, shaped around AWS's own request and response format, while a Spark serving a model through vLLM exposes an OpenAI-compatible chat completions API. Client code written directly against the Bedrock SDK needs a real integration change to move, not a base URL swap. Put LiteLLM in front of both and you get a single OpenAI-compatible gateway that can route to either a Spark or Bedrock's OpenAI-compatible model endpoints (where a given model supports that shape), which is the practical way to keep an application portable between the two without rewriting it per provider. Bedrock-only platform features, Knowledge Bases, Guardrails, and Agents, have no Spark equivalent because there is no managed platform layer underneath it, you would build retrieval, safety filtering, and orchestration yourself.
When AWS Bedrock is the right choice
- Your evaluation requires a specific closed model by name, Claude or Amazon Nova, and no open-weight model has matched it on your task yet.
- Your token volume is spiky or low, so on-demand per-token billing beats paying for a machine that sits idle between requests.
- You want AWS managing capacity, scaling, and platform features like Guardrails, Knowledge Bases, or Agents, rather than building and operating that layer yourself.
- You already run on AWS and want one bill, one IAM boundary, and one compliance story spanning Bedrock and the rest of your infrastructure.
When a dedicated Spark is the right choice
- An open-weight model, gpt-oss-120b, Llama 3.3, Qwen3-Coder, or similar, already meets your quality bar, so Bedrock's closed-model catalogue doesn't add anything for your task.
- Your traffic is steady enough to keep the machine busy for most of the hours you're paying for, batch document processing, an internal assistant used all day, or an agent fleet running continuously.
- You want the shortest possible answer to "who can access our data": one dedicated machine, root access under your own SSH keys, outside AWS entirely, and an operator who states in writing that it does not read your instance's content.
FAQ
Can I run Claude or Amazon Nova on a DGX Spark?
No. Claude, Amazon Nova, and any other closed model in Bedrock's catalogue run only on AWS's infrastructure, not on hardware you rent or own. A Spark runs open-weight models you deploy yourself: Llama 3.3, Mistral, gpt-oss-120b, Qwen3-Coder, DeepSeek, and similar. Bedrock also offers open-weight models like Llama and Mistral through the same managed API, so a Spark's model catalogue is a subset of Bedrock's, not a different one, only closed models are off the table entirely.
Is AWS Bedrock GDPR compliant if I use the Frankfurt region?
Choosing eu-central-1 (Frankfurt) as your Bedrock region keeps your inference requests within that AWS region, which is a data-residency control, not an automatic GDPR compliance statement. AWS's GDPR Data Processing Addendum is the governing document, and as with a GPUwerk Spark, you remain the controller responsible for the lawfulness of your own processing. Read AWS's own compliance documentation for your specific setup rather than taking any third-party summary, including this one, as legal advice.
Who can see my prompts on AWS Bedrock?
AWS states that Bedrock does not use customer inputs or outputs to train the underlying foundation models and does not share them with the model providers for their own training. AWS does apply automated content and abuse filtering to requests, and, as with any managed cloud service, AWS operational staff can access infrastructure under AWS's own support and security procedures. On a Spark, GPUwerk's published terms state GPUwerk hosts the machine but does not access, read, copy, index, or analyse the content of your instance, and there is no automated content scanning at all.
What happens to my data on a Spark when I stop or terminate the instance?
Terminating deletes your container and its /workspace volume from the node; a node is sanitised before it is offered to another customer. GPUwerk keeps only a periodic recovery copy of /workspace for hardware failures, not a live backup: while running or stopped, the workspace itself is on the node. Stopping keeps the workspace in place at 75% of the running hourly rate. Keep your own backup regardless.
Is a DGX Spark cheaper than AWS Bedrock?
It depends on the model, the token volume, and how continuously you can keep a Spark busy. Bedrock bills either per token on demand or by provisioning fixed throughput capacity by the hour, depending on the model and plan; AWS's own Bedrock pricing page has the current figures for your model and region, and this page does not repeat a specific number because GPUwerk could not independently verify one for a comparable open-weight model at the time of writing. A Spark is a flat $0.79/hour regardless of volume, so it favours steady, high-volume traffic; a lightly loaded Spark can lose to per-token billing that charges nothing when idle.
Can I switch between AWS Bedrock and a Spark without rewriting my application?
Partly. Bedrock has its own API shape (the Converse API and model-specific invoke calls), not the OpenAI-compatible chat completions shape a Spark serves through vLLM. Client code written against Bedrock's SDK needs a real integration change, not just a base URL swap, unless you already route through a gateway like LiteLLM that normalises both behind one interface. What never moves is anything tuned to a specific model's behaviour, and Bedrock-only features like Knowledge Bases, Guardrails, or Agents have no Spark equivalent because there is no managed platform layer, you would build any of that yourself.
GPUwerk did not fetch AWS's Bedrock documentation or pricing pages while writing this page; the statements about Bedrock's data handling, regions, and pricing model above reflect AWS's generally published product structure, not a page fetched and quoted on a specific date, and are marked accordingly where a specific number would otherwise be expected. If you need a citable, dated quote for Bedrock's own wording the way the Azure OpenAI comparison on this site has for Microsoft, treat this page as a starting point and verify against AWS's current documentation.