What are guardrails in LLM apps?
Guardrails are the filtering and validation layers wrapped around a language model in production: checks that run on a request before it reaches the model, and checks that run on the response before it reaches the user. They catch things the model's own training doesn't reliably catch on its own, and they can be changed without retraining or redeploying the model itself.
Where guardrails sit
An LLM app is usually more than the model. A request comes in, passes through some input handling, reaches the model, and the model's output passes through some output handling before it's returned. Guardrails live in those handling stages, not inside the model's weights. That separation matters: a guardrail is a piece of application logic you can update, swap, or turn off in an afternoon, while retraining a model to change its behavior is a much larger undertaking. It also means guardrails work the same way regardless of which model sits behind them, which is useful if you route requests across models by cost or capability.
What input-side guardrails do
On the way in, guardrails typically check for jailbreak attempts (prompts trying to override the system prompt or extract hidden instructions), prompt injection embedded in retrieved documents or tool outputs, and requests that fall outside the app's intended scope. Some also validate structure: rejecting a request that's missing a required field before it burns a model call at all.
What output-side guardrails do
On the way out, common checks include content moderation (blocking hate speech, self-harm content, or other categories the app shouldn't produce), PII redaction (scrubbing names, emails, or account numbers a model might have echoed from context), and schema validation (rejecting or repairing a response that was supposed to be valid JSON and isn't). Some setups also run a second, cheaper model as a judge over the first model's output, catching factual or tonal problems a rules-based filter would miss.
Why this is separate from model training
A model's safety tuning shapes what it's inclined to say, but it isn't a hard boundary. Models can be prompted around their training, they can make mistakes under ambiguous instructions, and a fine-tune done for one purpose can loosen behavior meant to hold for another. Guardrails give an operator a deterministic, auditable layer that doesn't depend on the model behaving as trained every time, which matters more, not less, when the app is self-hosted and the operator owns the failure mode directly rather than pointing at a vendor's usage policy.
Cost of getting this wrong
Guardrails that are too loose let bad output through; guardrails that are too strict block legitimate requests and frustrate users, and every check adds some latency, particularly on the input side, where a blocked request never reaches the model at all. That latency budget is the same one covered by time to first token: a heavyweight input classifier ahead of the model adds directly to the wait before generation even starts, so guardrail design on a self-hosted stack is a real engineering tradeoff, not a checkbox, and it's worth benchmarking with your own request patterns rather than assuming defaults are correct. The benchmarks page covers the underlying latency arithmetic on DGX Spark hardware that guardrail overhead adds on top of.