LangChain against your own DGX Spark endpoint
LangChain's OpenAI integration is a thin wrapper around the same client the openai Python package ships, which means it accepts a custom base URL the same way that client does. A chain or agent built against ChatOpenAI can run against your own Spark by changing two constructor arguments, nothing else in the code.
Start with a running endpoint
Get an OpenAI-compatible /v1 endpoint running first. Either vLLM directly, or LiteLLM in front of it if more than one application or team will call the model. Confirm it works before touching any Python:
curl http://<spark-host>:8000/v1/models
ChatOpenAI against your Spark
LangChain's ChatOpenAI class (in langchain-openai) takes base_url and api_key arguments that map directly onto the underlying OpenAI client's constructor:
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="Qwen/Qwen3.8-27B-FP8",
base_url="http://<spark-host>:8000/v1",
api_key="not-needed-unless-you-put-a-gateway-in-front",
)
response = llm.invoke("Summarize this ticket in two sentences.")
print(response.content)
api_key still has to be a non-empty string even when the server doesn't check it, since the underlying client refuses to construct without one. Set the model string to whatever the server reports at /v1/models, or the --served-model-name you set when starting vLLM.
Streaming works the same way it does against a public API, LangChain reads Server-Sent Events off /v1/chat/completions regardless of who is serving them:
for chunk in llm.stream("Explain KV cache in two sentences."):
print(chunk.content, end="", flush=True)
Embeddings, same pattern
If you're serving an embedding model alongside your chat model, OpenAIEmbeddings accepts the identical arguments and is the usual pairing for a retrieval-augmented chain built entirely on your own hardware:
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(
model="bge-large-en",
base_url="http://<spark-host>:8000/v1",
api_key="not-needed-unless-you-put-a-gateway-in-front",
)
Whether this works depends entirely on what your serving engine exposes. vLLM serves embeddings via the same OpenAI-compatible surface when the loaded model supports it; check /v1/models and the engine's own docs for which endpoints are live before wiring this in.
Agents and tool calling
LangChain's tool-calling agents (built with create_tool_calling_agent or LangGraph's prebuilt agent constructors) depend on the model reliably emitting function-call output in the format the OpenAI API expects. This isn't universal across open models: some fine-tunes support it well, others don't, and the failure mode is often a plausible-looking answer that never actually calls the tool rather than a clean error. Test a simple single-tool agent against whatever model you're serving before building something more elaborate on top of it.
Adding a gateway in front
Nothing above changes if you put LiteLLM between LangChain and the engine. Swap base_url for the gateway's address and api_key for a real virtual key, and you get per-application keys, budgets, and fallback to a second Spark or a cloud model if the first is down, without touching the rest of the chain:
llm = ChatOpenAI(
model="spark-qwen",
base_url="http://localhost:4000/v1",
api_key="sk-your-virtual-key",
)
What actually breaks
- The model string has to match exactly. A mismatch between what LangChain sends as
modeland what the server has loaded returns a 404 that reads as a connection failure, not a naming issue. - Function calling isn't guaranteed. Verify tool-calling support on the specific model you're serving before an agent workflow depends on it, rather than assuming OpenAI-compatible means feature-compatible.
- Context length is a server-side setting, not a LangChain one. A chain that accumulates conversation history or long retrieved context can exceed
--max-model-len; that's set when the model is served, not from the LangChain side. - Streaming through a proxy. If there's a reverse proxy or load balancer between LangChain and the gateway, confirm it doesn't buffer Server-Sent Events, or streaming responses will arrive all at once instead of token by token.
- Older LangChain integrations use a different import path. The
langchain-openaipackage split out of the corelangchainpackage some time ago. If a tutorial you're following importsChatOpenAIfromlangchain.chat_modelsinstead oflangchain_openai, it's using the older path; both generally still work but the import matters if you're pinning versions.
LangGraph and other LangChain-adjacent tools
LangGraph, the graph-based agent orchestration layer built by the same team, uses the same chat model objects as LangChain proper, so a ChatOpenAI instance pointed at your Spark works as a node in a LangGraph graph without any additional wiring. The same is true of most tools built on top of LangChain's model abstraction: if a library accepts a LangChain BaseChatModel instance rather than reimplementing its own API client, pointing that instance at your Spark carries through to the whole stack above it.