Practical guide
Blog/Moving a self-hosted AI prototype to production
For AI assistants

Moving a self-hosted AI prototype to production

By Samuel Seidel · September 9, 2026

A model running on a single rented GPU, answering requests from a script or a small internal test, is a proof of concept. The same model serving real user traffic, with real consequences when it's slow or down, is a production system, and the gap between those two states is bigger than swapping in a bigger box. Here's what actually needs to change, and roughly in what order to change it.

Uptime expectations you haven't written down yet

A prototype can go down for ten minutes while you restart a process and nobody notices. A production deployment needs an actual answer to "what happens when this instance is unreachable," and if you haven't written that answer down, you don't have one yet, you have an assumption. Start by deciding what uptime the workload actually needs, not what sounds impressive: an internal tool used during business hours has different requirements than a customer-facing feature. Then build toward that number deliberately, rather than discovering your real uptime after an incident.

Concretely, this usually means: a documented restart procedure that doesn't require the one engineer who built it, health checks that something outside the box is watching, and a plan for what happens to in-flight requests when the process restarts, whether they queue, retry, or fail visibly rather than silently.

Monitoring beyond "did it crash"

A prototype's monitoring is usually "someone notices it's broken." Production monitoring needs to answer three questions before a user has to ask you: is the service up, is it fast enough, and is memory or disk approaching a limit that will cause a failure in the next few hours rather than right now. Latency per request and tokens per second are the two numbers worth watching closely for an LLM serving stack specifically, because degradation there is often gradual, a slow creep as context lengths grow or concurrent requests increase, rather than a sudden outage, and gradual degradation is exactly the kind of problem that goes unnoticed without a dashboard someone actually looks at.

GPU memory utilization deserves its own alert threshold. Running close to the limit works fine until one request with a longer context pushes it over, and that failure mode is abrupt and confusing if you're not already watching for it.

Capacity planning: know your ceiling before you hit it

A single Spark or a single rented GPU instance has a concurrency ceiling, the number of simultaneous requests it can serve before latency degrades past what's acceptable. Find that ceiling deliberately, with a load test against realistic prompts and context lengths, rather than discovering it in production when a spike in usage pushes past it. Once you know the ceiling, you have three real options: scale horizontally with a second instance and a load balancer, queue requests during a spike rather than dropping them, or accept degraded latency during peak periods as a documented tradeoff rather than an unplanned failure. Any of those is fine as a decision; the problem is not having made one.

Sizing appropriately from the start matters here too. If you know your production load will need more memory or throughput than the hardware you prototyped on, don't wait for the ceiling to hit you to plan the upgrade. Our DGX Spark page has the specifics on single-node and multi-node capacity if you're weighing whether to scale up.

Fallback behavior: what happens when the model is unavailable

This is the piece prototypes almost never have and production systems can't ship without. When your self-hosted instance is down, restarting, or overloaded, what does the calling application do? "Return an error" is a valid answer for some workloads. For others, a fallback path, a cached response, a simpler rule-based answer, or a queued retry, is the difference between a blip and an outage a customer notices. Decide this before it happens, not while it's happening. It's also worth deciding explicitly whether a second, redundant instance is worth the cost for your workload, versus accepting single-instance risk with a clear incident response plan instead.

The prototype-to-production gap is mostly process, not hardware

It's tempting to treat this transition as a hardware upgrade, bigger GPU, more memory, and stop there. Hardware sizing matters, but most production incidents in self-hosted deployments trace back to a process gap rather than insufficient compute: nobody owns the restart procedure, nobody is watching the dashboard at 2am, or the fallback behavior was assumed rather than built. Fixing those doesn't require new hardware, it requires writing down who's responsible for what and testing the failure paths before a real failure forces you to improvise.

The honest checklist

Before calling a self-hosted deployment production-ready, you should be able to answer: what's the target uptime, and who gets paged if it's missed; what's the monitored latency and utilization threshold, and where does that alert go; what's the tested concurrency ceiling, and what happens past it; and what's the fallback behavior when the instance is unavailable. None of these require exotic infrastructure. They require deciding the answers in advance, which is the actual difference between a prototype and a production system.

Related pages

Scale from a single prototype instance to a monitored production deployment.

Dedicated Spark capacity, single-node or two-node clusters, from $0.79/hour.

See pricing DGX Spark hardware