Serving
DGX Spark/Structured output and function calling on a self-hosted LLM

Structured output and function calling on a self-hosted LLM

By Samuel Seidel · Published September 9, 2026

Getting a self-hosted model to return valid JSON, or to call a tool with correctly typed arguments, is a different problem from getting a hosted API to do it. The mechanics live in your serving engine rather than behind someone else's endpoint, which means you're responsible for choosing a model that was actually trained for tool use, wiring the right flags, and handling the failure mode where the model almost gets the format right.

Two different guarantees

It's worth separating what "structured output" actually means before configuring anything, because vLLM and llama.cpp offer two distinct guarantees that get conflated in casual use.

Grammar-constrained decoding forces the output to match a JSON schema at the token level: the engine masks out any token that would make the output invalid against your schema, so malformed JSON is structurally impossible. This is what you want for anything downstream that parses the response with json.loads and breaks on a stray comment or an unescaped quote.

Prompted JSON mode, asking the model nicely to return JSON via the system prompt, with no grammar enforcement, is weaker. Better-instruction-tuned models comply most of the time, but "most of the time" means your parser needs a retry loop regardless. Use grammar-constrained decoding wherever your engine supports it; treat prompted JSON as a fallback for cases the constrained path doesn't cover.

Setting it up in vLLM

vLLM's OpenAI-compatible server accepts a response_format with json_schema in the chat completions request, backed by a constrained decoding backend (xgrammar by default in recent versions). Start the server normally, per the vLLM install and flags guide, then pass the schema per request:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-model",
    "messages": [{"role": "user", "content": "Extract the invoice fields."}],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "invoice",
        "schema": {
          "type": "object",
          "properties": {
            "vendor": {"type": "string"},
            "total": {"type": "number"},
            "due_date": {"type": "string"}
          },
          "required": ["vendor", "total", "due_date"]
        }
      }
    }
  }'

For tool calling, vLLM supports the OpenAI tools and tool_choice parameters, but only for models with a matching chat template and parser; you need to launch the server with --enable-auto-tool-choice and a --tool-call-parser matching your model family (for example hermes, llama3_json, or mistral, depending on how that model was trained to emit calls). Check your model's card for which parser it expects; passing tools to a model or template that doesn't support them produces output that looks like a tool call but isn't parsed as one.

Setting it up in llama.cpp

llama.cpp's server exposes a json_schema or raw grammar field (GBNF format) on completion requests, enforced via grammar-constrained sampling at generation time, the same category of guarantee as vLLM's constrained decoding, different implementation. For OpenAI-style tool calling, recent llama.cpp server builds support --jinja to use the model's own chat template plus tools in the request body, but coverage of tool-calling templates varies more by model here than in vLLM; test with your specific GGUF before relying on it in production, per the notes in the llama.cpp guide.

Model choice matters more than engine choice

Constrained decoding guarantees the shape of the output is valid; it says nothing about whether the content is correct, a model that wasn't trained on tool-calling data will produce syntactically valid JSON with the wrong tool picked, or arguments that don't match the actual task, filled in to satisfy the schema rather than to answer the question. Check your model's documentation for explicit tool-calling or function-calling support before building a pipeline around it. Among the models in the GPUwerk model catalog, check each model's own docs page for tool-calling support before committing a pipeline to it, this varies by release and isn't safe to assume from parameter count alone.

Handling the failure mode

Even with grammar constraints, two things still go wrong in practice. First, the model can satisfy the schema with a technically valid but semantically empty response, an empty string where a real value was expected, because the grammar only enforces type, not content. Second, very long or deeply nested schemas slow decoding, since the constraint check runs on every token; keep schemas as flat and specific as the task allows, and avoid schemas with many optional fields the model has to needlessly reason about. Log the raw output alongside the parsed result during development so a wrong-but-valid response is visible before it reaches production silently.

Practical checklist

Testing before you trust it

Don't take a model's tool-calling documentation as a guarantee that it behaves correctly on your specific schemas. Build a small test set of representative requests, including a few adversarial ones (ambiguous instructions, a request that should trigger no tool call at all, arguments near a type boundary like a negative number where the schema expects a positive one), and run it against your actual server configuration before wiring the pipeline into anything that acts on the output automatically. This catches the gap between "supports tool calling" in general and "calls the right tool with the right arguments for your prompts", which is where most production issues in this area actually originate.

See the vLLM guide and the llama.cpp guide for the underlying install and flag references, or the model catalog for which models are documented as supporting tool calling.

First top-up: pay $10, get $20 in credit

Wire up tool calling on your own box.

Deploy a dedicated Spark and configure structured output with vLLM or llama.cpp from the console.

Deploy a Spark Read the vLLM guide