Private LLM hosting vs the Perplexity API
We rent DGX Sparks, so weigh that against everything below. This comparison is a bit unusual because Perplexity's API and a dedicated Spark aren't really competing for the same job. Perplexity's product is retrieval: it searches the live web, ranks sources, and has a model synthesize an answer with citations attached. A Spark gives you a model and nothing else, no search, no citations, unless you build that layer yourself. If what you actually need is grounded, cited answers about current events or recent documents, Perplexity does something a bare model on a Spark cannot do out of the box.
Side by side
| Perplexity API | Dedicated DGX Spark (GPUwerk) | |
|---|---|---|
| What the product does | Retrieval-augmented answers: live web search plus a model that synthesizes a response with source citations, via the Sonar model family. | Raw model inference on hardware you control. No built-in search; whatever context reaches the model is what you send it. |
| Where queries are processed | Perplexity's infrastructure, plus outbound calls to the search index and the sites it retrieves from. GPUwerk did not fetch Perplexity's current documentation on specific hosting regions. | EU-Central (Prague), on one dedicated machine, no outbound search calls unless you add them yourself. |
| Who can see your queries | Perplexity, as operator of the API, and necessarily the search layer it queries on your behalf. GPUwerk did not fetch Perplexity's current data-usage and training policy for this page; check Perplexity's own terms for retention and staff access. | Nobody at GPUwerk. Our DPA states GPUwerk "hosts the machine but does not access, read, copy, index, or analyse the content of the controller's instance." |
| Model choice | Perplexity's Sonar models, tuned specifically for search and synthesis, plus access to some third-party frontier models through the same API depending on plan. | Open-weight models you choose and run yourself: gpt-oss-120b, Llama 3.3 70B, Qwen3-Coder 30B, DeepSeek, and anything else that fits in 128 GB unified memory. |
| Freshness of information | Current by design: answers reflect what the live web says today, with citations you can check. | Frozen at the model's training cutoff unless you feed it fresh context yourself, whether that's a document store, a scraped feed, or a search API you wire up separately. |
| Pricing model | Per request and per token, with search-specific line items. Check Perplexity's own pricing page for current rates, GPUwerk did not fetch a current figure to avoid quoting a stale price. | Per hour, billed per minute: $0.79/hour on-demand, $0.59/hour for a customer-requested stop. One rate regardless of model or request volume, per pricing. |
| What you operate yourself | Nothing at the infrastructure or search layer; Perplexity manages the index and capacity. You manage how you call the API and what you do with citations. | Root SSH access to your own container: the model server, any retrieval layer you choose to add, monitoring, and backups. GPUwerk keeps only a recovery copy of /workspace for hardware failures, refreshed roughly every six hours; it is not a backup service, so that responsibility is entirely yours. |
Where the honest difference actually is
Perplexity built something genuinely hard to replicate with a bare model: a search index, a ranking layer, and a model tuned to cite its sources instead of confabulating them. If your workload is "answer questions about things that happened this week" or "summarize what's currently on the web about a topic," that's the product, and a Spark running an open-weight model with no search access simply cannot do the same job without you building a retrieval pipeline on top of it.
What a Spark changes is who touches the query on the way to an answer. Perplexity's API necessarily involves at least two hops outside your control: the model call, and the search calls it makes on your behalf to third-party sites and its own index. A dedicated Spark keeps everything on one machine you control, with GPUwerk's stated policy of not accessing instance content, but it only does that for the model-inference half of the job. If you need live retrieval, you either accept a third party in that loop, whether it's Perplexity or a search API you wire up yourself, or you give up freshness.
Some teams run both: a Spark for anything touching regulated or proprietary data that doesn't need live web context, and Perplexity's API for genuinely open, current-events questions where citations matter more than data control.
The worked cost example
Take a workload generating 500 million output tokens a month from a model, ignoring the search-specific costs Perplexity would add on top since those depend heavily on request pattern.
Perplexity API. GPUwerk did not fetch current per-token or per-request pricing for a specific Sonar model at the time of writing. Run your own request volume through Perplexity's pricing page for a number you can trust, keeping in mind that search-grounded requests typically carry a per-call fee on top of tokens.
Dedicated Spark. gpt-oss-120b reaches 862.8 tok/s aggregate at 256 concurrent requests, per the concurrency benchmark on our benchmarks page. That's 862.8 × 3,600 = 3,106,080 tokens per hour if the node stays saturated at that concurrency throughout. 500,000,000 ÷ 3,106,080 ≈ 161 hours. At $0.79/hour, 161 × $0.79 = $127.19, before tax, and before whatever you spend building and running your own retrieval layer.
That $127.19 figure only covers plain inference. It is not a fair like-for-like comparison to Perplexity's price, because Perplexity is selling search-and-synthesis as one product and a bare Spark is not. Add your own retrieval cost, whether that's a search API's fees or the engineering time to build one, before comparing totals.
Migration path
There isn't a clean drop-in migration here, because the products aren't the same shape. If you're calling Perplexity purely for its underlying model and not using search grounding, moving that call to a Spark running an OpenAI-compatible endpoint via vLLM is straightforward: change the base URL and API key. If you rely on citations and live retrieval, you'd need to build or buy a separate search layer to feed a Spark-hosted model equivalent context, which is a real engineering project, not a config change.
When the Perplexity API is the right choice
- Your workload needs answers grounded in the current web, with citations users can verify.
- You don't want to build and maintain your own search or retrieval infrastructure.
- Query volume is variable enough that per-request pricing beats paying for a machine around the clock.
When a dedicated Spark is the right choice
- Your workload doesn't need live web access, and a frozen-at-training-time model plus your own documents is enough.
- You want no third party touching the query at all, and can live without built-in citations.
- You already have or are willing to build your own retrieval layer, and want full control over the model underneath it.
FAQ
Is the Perplexity API a substitute for a private LLM?
Not directly. Perplexity's API is built around retrieval, returning answers grounded in live web search with citations. A private LLM on your own hardware answers from the model's weights and whatever context you give it, with no built-in web access. They solve different problems, and some teams need both.
Does Perplexity's API see my queries?
GPUwerk did not fetch Perplexity's current data-retention or training policy for this page. Check Perplexity's own privacy policy and API terms for what is logged, for how long, and whether staff or sub-processors can access query content.
Can I run web search on a dedicated Spark instead?
You can wire a Spark-hosted model to your own search API or a self-hosted retrieval layer, but that is infrastructure you build and operate yourself. Perplexity ships search and synthesis as one product; a Spark ships compute for a model you choose.
Is a Spark more private than the Perplexity API?
On the question of who else can read a given prompt, yes, for the model-inference part: a Spark is one dedicated machine under your own SSH keys, and GPUwerk's terms state GPUwerk does not access, read, copy, index, or analyse instance content. Perplexity's API additionally routes queries out to live web search, which is a different data flow entirely and one a dedicated Spark alone does not replicate.
GPUwerk did not fetch Perplexity's current pricing, data-usage policy, or infrastructure documentation for this page; every Perplexity-specific claim above is either general public knowledge about the product's shape or explicitly hedged. Check perplexity.ai directly for current terms. Spark figures come from our benchmarks page and our pricing page.