n8n with a self-hosted LLM on a DGX Spark
n8n's AI nodes talk to language models through the same OpenAI-compatible shape most self-hosted engines speak. That means a workflow that summarizes support tickets, drafts replies, or extracts fields from an invoice can call your own Spark instead of a public API, with the same node configuration either way.
Why route n8n's AI nodes yourself
A workflow built around an AI node usually carries whatever triggered it: a support email, a form submission, a row from a customer database. Every one of those calls out to whichever model the node is configured to use, and by default that means a public API, running on infrastructure you don't control, under a provider's own retention policy. n8n itself doesn't require that. Its AI nodes accept a custom base URL, so the same trigger, the same prompt template, and the same downstream steps can run against a model on your own Spark, and the request never leaves your network.
Start with a running endpoint
Get an OpenAI-compatible endpoint running first, either vLLM or LiteLLM in front of it if more than one workflow or team will share the model. Confirm it answers before touching n8n:
curl http://<spark-host>:8000/v1/models
If n8n runs on a different machine or in its own container, make sure that host can actually reach the Spark. On Docker, a Spark endpoint reachable from the n8n host's network is enough; there is no special n8n-side networking beyond that.
Add the credential in n8n
n8n's built-in OpenAi credential type takes a base URL, not just an API key, which is what makes pointing it at a self-hosted model possible without a community node. In n8n, go to Credentials, New, OpenAi, and set:
API Key: not-needed-unless-you-put-a-gateway-in-front Base URL: http://<spark-host>:8000/v1
vLLM does not check the API key field unless you started it with --api-key, but n8n's credential form requires something non-empty, so any placeholder string works. If you've put LiteLLM in front of the engine for per-workflow keys and budgets, use the gateway's URL and a real virtual key here instead, and the rest of the setup is unchanged.
Which nodes to use
n8n's AI functionality splits across a few nodes that all accept this same credential:
- OpenAI Chat Model is the sub-node that plugs into an AI Agent or a chain, and it's where the credential above gets attached.
- AI Agent is the higher-level node for tool-using workflows, ticket triage that can look up an order, a chatbot that can call other n8n workflows as tools. It uses whichever chat model sub-node you connect to it.
- Basic LLM Chain is the simpler option when there's no tool use, just a prompt in and a completion out, useful for straightforward summarization or classification steps.
All three route through the same credential, so switching a workflow from a public API to your Spark is a matter of changing which OpenAi credential the chat model node uses, not rebuilding the workflow.
A minimal workflow
A common shape: a webhook or email trigger feeds a Basic LLM Chain node that summarizes the incoming text, then a downstream node writes the result somewhere, a database row, a Slack message, a ticket field. The chain node's prompt is just a text field:
Summarize the following support ticket in two sentences,
and flag the priority as low, medium, or high:
{{ $json.body }}
Set the chain's model to whatever --served-model-name your Spark reports at /v1/models, or the alias you gave it in LiteLLM's config. n8n sends the request to your Base URL exactly the way it would to a public API; nothing about the prompt template or the downstream nodes changes.
What actually breaks
- Network reachability, not n8n config. If n8n itself runs in Docker, "localhost" inside that container is not your Spark. Use the Spark's actual hostname or IP, or an internal DNS name if both are on the same private network.
- The model name has to match exactly. A mismatch between the model string in the credential's chat node and what the server has loaded returns an error that looks like a broken credential, not a wrong model name.
- Tool-calling support varies by model. The AI Agent node relies on function/tool calling to decide when to invoke a connected tool. Not every model serves this reliably; check the model's documentation or test a simple tool-using workflow before building something that depends on it.
- Long-running workflows and short context windows. An agent workflow that accumulates conversation history or tool output across several steps can exceed
--max-model-lenfaster than a single-shot chat. Size the context window for the workflow, not just for one prompt.