Private LLM hosting vs Amazon SageMaker
We rent DGX Sparks, so weigh that against everything below. SageMaker isn't Bedrock, our other AWS comparison covers Bedrock's managed model API, this page is about SageMaker's self-managed hosting: you deploy a model to an AWS-managed instance and operate the serving stack yourself, closer in shape to renting a Spark than to calling a hosted API. The difference is who runs the machine, AWS with instance-hour billing and its own platform layer on top, or GPUwerk with one dedicated box at a flat hourly rate, and how much orchestration you get for it.
Side by side
| Amazon SageMaker | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What it is | A managed ML platform: you deploy a model artifact or container to a SageMaker endpoint running on an AWS-selected instance type, with training, autoscaling, and monitoring tools built around it. | One dedicated physical machine, rented by the hour, with root access. No orchestration platform underneath. |
| Where data is processed | Whichever AWS region you deploy the endpoint in. AWS operates EU regions including Frankfurt and Stockholm; AWS states it does not move customer content out of the selected region except to serve the request or comply with law. | EU-Central (Prague), on one dedicated machine, no other region to configure. |
| Who can see your data | Governed by the AWS Customer Agreement and the AWS GDPR Data Processing Addendum, which AWS states applies automatically to in-scope services. AWS states it does not access customer content except as necessary to provide the service or as legally required. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | Open-weight models through the SageMaker JumpStart catalogue, or any model you package in a custom container, per AWS documentation. No restriction on model family beyond what fits the instance you choose. | Open-weight models only: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else you can fit in 128 GB. Same open-model category as JumpStart, on one fixed piece of hardware instead of a choice of instance types. |
| Speed and concurrency | Depends entirely on the GPU instance type you provision (AWS offers several, from single- to multi-GPU) and how you configure autoscaling; AWS does not publish a single tokens/second figure because it varies by instance and model. | A single Spark, single-stream: gpt-oss-120b decodes at 33.5 tok/s and reaches 862.8 tok/s aggregate at 256 concurrent requests (Dendro Logic's concurrency benchmark, run on their own Spark). Raw logs and methodology on our benchmarks page. |
| Pricing model | Per instance-hour for the GPU instance type you provision, plus SageMaker's platform fee on top of the raw EC2 rate, published on AWS's own pricing page and varying by region and instance family. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop that holds your reservation. One rate regardless of model, per pricing. |
| Worked cost example | See the full arithmetic below the table. | |
| Contracts and DPA | Governed by the AWS Customer Agreement and the AWS GDPR Data Processing Addendum. | A standard GDPR Article 28 DPA published free at /legal/dpa, no negotiation required. Sub-processor list at /legal/sub-processors states none are engaged for instance workloads. |
| What you operate yourself | The model container, serving code, and monitoring configuration; AWS manages the underlying instance provisioning, patching, and autoscaling infrastructure. | Everything, root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
The worked cost example
Take a workload that generates 500 million output tokens a month, roughly what a busy internal support bot or a document-processing pipeline running continuously might produce. Two ways to serve it:
Amazon SageMaker. Real-time inference endpoints bill per instance-hour for the GPU instance type you provision, and SageMaker adds a platform fee on top of the raw EC2 rate for that instance. AWS publishes current instance-hour prices by region on its own SageMaker pricing page, but the exact figure for a given instance family and EU region needs to come from there, GPUwerk did not attempt to reconstruct a single blended number that would hold across instance choices.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per Dendro Logic's concurrency benchmark on a single Spark (cited on our benchmarks page). That is 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency the whole time. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That Spark figure only holds under the assumption stated: the node is busy at 256 concurrent requests for the full 161 hours, back to back. SageMaker's per-instance-hour price on a comparable GPU is a number you'd need to pull from AWS's pricing page for your chosen instance and region, then run through the same idle-time logic, a SageMaker endpoint left running with light traffic bills for every hour on the clock the same way a Spark does. Where SageMaker earns its platform fee is autoscaling: an endpoint configured to scale to zero, or to add instances under load, changes this arithmetic in ways a single fixed Spark cannot match.
Migration path
A model served through vLLM on a Spark and a model deployed to a SageMaker real-time endpoint both typically expose an HTTP inference endpoint, though the request and response shape differs unless you specifically configure SageMaker's endpoint to match an OpenAI-compatible format. Put LiteLLM in front of either and you get a consistent OpenAI-compatible gateway with request logging and key management. What doesn't move is SageMaker's autoscaling configuration, multi-model endpoint setup, or its integration with SageMaker Pipelines and Model Monitor, none of that has an equivalent on a single Spark, you'd build your own scaling and monitoring layer.
When SageMaker is the right choice
- You need autoscaling that adds or removes GPU instances under load, or scales an endpoint to zero when idle, rather than paying a fixed hourly rate for one machine.
- Your ML workflow already spans SageMaker Pipelines, Feature Store, or Model Monitor, and keeping inference in the same ecosystem is worth more than the simplicity of a separate box.
- You want a choice of GPU instance types and can size the instance to the model, rather than working within one fixed hardware spec.
When a dedicated Spark is the right choice
- Your traffic is steady enough that a fixed always-on machine is simpler to reason about than autoscaling configuration and its edge cases.
- You want a single, flat hourly rate you can put in a spreadsheet without pulling instance-hour prices from a region-specific pricing page.
- You want the shortest possible answer to "who can access our data": one dedicated machine, root access under your own SSH keys, and an operator who states in writing that it does not read your instance's content.
FAQ
How is SageMaker different from Amazon Bedrock?
Bedrock is a managed API to closed and open foundation models, you send prompts and get completions, Amazon runs the model. SageMaker is a broader ML platform for training and hosting models yourself, you pick an instance type, deploy a model artifact or container to a SageMaker endpoint, and Amazon manages the underlying infrastructure while you manage the model and its serving stack. SageMaker is closer in shape to a dedicated Spark than Bedrock is, both hand you an instance to run a model on, but SageMaker's instance is AWS-managed and billed by instance-hour plus platform fees, while a Spark is a single dedicated machine billed at a flat hourly rate.
Can I run open-weight models on a SageMaker endpoint?
Yes. SageMaker supports deploying open-weight models through its JumpStart catalogue or as a custom container on a SageMaker real-time or serverless inference endpoint, per AWS documentation. That's the same category of model a Spark runs, gpt-oss-120b, Llama 3.3, Qwen3-Coder, and similar, the difference is who operates the machine underneath and how it's billed.
Where does SageMaker process my data?
Whichever AWS region you deploy the endpoint in. AWS operates EU regions including Frankfurt and Stockholm, and AWS's own documentation states that AWS does not move customer content outside the region a customer selects, except as needed to serve requests or comply with law. A Spark has no region to configure, it runs in EU-Central (Prague) because that is the only place it runs.
Is a DGX Spark cheaper than SageMaker?
It depends on the instance type and utilization. In the worked example on this page, 500 million generated tokens a month costs about $127 in Spark compute at $0.79/hour, assuming the node sustains 256 concurrent requests the entire time. SageMaker's real-time inference endpoints bill per instance-hour for GPU instance types AWS publishes on its own pricing page, plus SageMaker's platform fee on top of the raw EC2 rate; check that page for the instance type and region you'd actually use before comparing.
What does SageMaker manage that a Spark doesn't?
Endpoint autoscaling, multi-model endpoints, built-in A/B testing between model versions, managed training jobs, and integration with the rest of the AWS ML stack (Feature Store, Pipelines, Model Monitor). A Spark is one machine with root access and none of that orchestration layer; you would build any of it yourself on top.
AWS pricing and region statements referenced above come from AWS's own SageMaker pricing page and its published data residency and GDPR documentation; GPUwerk did not find a single blended instance-hour figure that holds across instance types and regions, so the pricing row reflects that variability rather than a guess. Spark throughput figures come from Dendro Logic's concurrency benchmark, detailed on our benchmarks page.