Private LLM hosting vs Together AI
We rent DGX Sparks, so weigh that against everything below. Together AI has done real work making open models easy to call: a large catalog, an OpenAI-compatible API, fine-tuning and batch tooling, and infrastructure that scales without you touching a GPU. That's a genuinely useful product for teams that want an open model without operating hardware. The trade is tenancy: Together's standard serverless API runs your requests on shared GPU capacity that Together operates and allocates across customers. A dedicated Spark is the opposite shape, one machine, one tenant, at the cost of doing the setup yourself.
Side by side
| Together AI | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| Tenancy | Shared, multi-tenant GPU pool for the standard serverless API. Together also sells dedicated endpoints as a separate, typically pricier tier, check Together's docs for current terms. | Single-tenant by default: one physical Spark, one SSH key set, nobody else's workload runs on it. |
| Where data is processed | Together's own data centers; GPUwerk did not fetch Together's current documentation on which specific regions back a given request. | EU-Central (Prague), on one dedicated machine, no region choice needed because there's only the one machine. |
| Who can see your prompts | Together, as operator of the API, on the shared tier. GPUwerk did not fetch Together's current data-usage and training policy for this page; check Together's own privacy policy and API terms for what is retained and by whom. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | A large, actively maintained catalog of open-weight models available immediately via API, no download or setup required. | Any open-weight model you choose that fits in 128GB unified memory: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and others, but you download, configure, and serve it yourself. |
| Speed and concurrency | Together publishes its own throughput figures per model; GPUwerk did not fetch current numbers for this page. | Measured, not estimated: gpt-oss-120b decodes at 33.5 tok/s single-stream and reaches 862.8 tok/s aggregate at 256 concurrent requests; Qwen3-Coder 30B-A3B AWQ decodes at 80.9 tok/s single-stream. Full methodology on our benchmarks page. |
| Pricing model | Per token on the serverless tier, priced by model size and type. Check Together's own pricing page for current rates, GPUwerk did not fetch a current figure to avoid quoting a stale price. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop. One rate regardless of model or request volume, per pricing. |
| What you operate yourself | Nothing at the infrastructure layer for the serverless tier; Together manages capacity and model serving. You manage prompts and application code. | Root SSH access to your own container: the model server, any RAG or agent layer, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
Where the honest difference actually is
Together's catalog is the real draw: dozens of open models ready to call the moment you have an API key, with no GPU procurement, no driver setup, no capacity planning. For a team that wants to try five different open models in an afternoon, that's hard to beat, and a Spark cannot match it without real setup time per model.
The trade is in what "shared" means in practice. On Together's standard serverless tier, your requests run on GPU capacity Together allocates across its customer base, the same shape as most hosted inference APIs. Together does sell dedicated endpoints for customers who want isolated hardware, and if that's the tier you're comparing, the tenancy gap narrows considerably, though you'd still be running on infrastructure Together operates rather than a machine under your own keys. A Spark's difference is that the isolation isn't a paid upgrade, it's the only mode: one machine, your SSH keys, and GPUwerk's stated policy of not accessing instance content.
Whether that matters depends on what you're running. A prototype or an internal tool with no sensitive data probably doesn't need it. A workload touching client data, regulated records, or anything under a contract that specifies where processing happens is where the difference stops being abstract.
The worked cost example
Take a workload generating 500 million output tokens a month.
Together AI. GPUwerk did not fetch current per-token pricing for a specific Together-hosted model at the time of writing, and rates vary meaningfully by model size. Run your own volume and model choice through Together's pricing page for a number you can trust.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per the concurrency benchmark on our benchmarks page. That's 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency throughout. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax.
That figure assumes constant saturation for the full 161 hours. Together's per-token pricing scales cleanly with actual usage and needs no minimum commitment, so at low or bursty volume it can beat an hourly-billed machine even before considering setup time; the comparison flips as volume climbs and utilization on the Spark stays high.
Migration path
Together's API and a Spark serving a model through vLLM both expose an OpenAI-compatible chat completions shape, so client code that only calls that endpoint moves with a base URL and API key change. What doesn't move automatically is any model-specific behavior tuning done against a particular Together-hosted checkpoint; if you switch quantization or serving stack on the Spark side, re-validate output quality before cutting over.
When Together AI is the right choice
- You want to try or run many different open models without setting up infrastructure for each one.
- Your volume is variable enough that per-token pricing beats paying for a machine around the clock.
- You don't have a hard requirement for single-tenant hardware, or you're willing to pay for Together's dedicated tier if you do.
When a dedicated Spark is the right choice
- Single-tenant hardware is a requirement, not a nice-to-have, and you want it as the default, not a paid upgrade.
- You've already settled on one or two models and don't need a large ready-made catalog.
- Your traffic is steady enough to keep the machine busy for most of the hours you're paying for.
FAQ
Is Together AI multi-tenant?
Together AI's serverless inference API serves requests from a shared pool of GPU capacity across customers, the standard shape for a hosted inference API. Together also offers dedicated endpoints and reserved capacity as a separate, typically more expensive tier, check Together's own documentation for what that includes.
Does Together AI train on my API data?
GPUwerk did not fetch Together AI's current data-usage and training policy for this page. Check Together's own privacy policy and API terms for what is retained and whether it differs by plan.
Is a Spark more private than Together AI's shared API?
For Together's standard serverless tier, yes on the tenancy question: a Spark is one dedicated machine under your own SSH keys, and GPUwerk's terms state GPUwerk does not access, read, copy, index, or analyse instance content. Together's dedicated-endpoint tier narrows that gap since it also isolates hardware per customer, compare the two directly if you're considering that tier.
Does Together AI have more models than GPUwerk?
Together AI hosts a large, actively updated catalog of open-weight models ready to call via API with no setup. On a Spark you can run any open-weight model that fits in 128GB unified memory, but you have to download, configure, and serve it yourself.
GPUwerk did not fetch Together AI's current pricing, data-usage policy, or infrastructure documentation for this page; every Together-specific claim above is either general public knowledge about the product's shape or explicitly hedged. Check together.ai directly for current terms. Spark figures come from our benchmarks page and our pricing page.