Cursor and Continue.dev against your own DGX Spark endpoint
The coding agents guide covers Continue, Cline and aider as editor and terminal agents pointed at a self-hosted endpoint. This page is narrower: the IDE-integrated assistants built into an editor's own UI, specifically Cursor, plus Continue.dev's own IDE panel, and what each one needs to route through your own Spark instead of its default cloud backend.
Start with a running endpoint
Both assistants below need the same thing: an OpenAI-compatible /v1 endpoint reachable from your laptop. Get that running first with vLLM, and confirm it before touching either editor:
curl http://<spark-host>:8000/v1/models
If the Spark isn't reachable from your laptop directly, tunnel it over SSH per the SSH keys guide, or bind it to a public port on an instance you control.
Cursor
Cursor supports custom OpenAI-compatible providers through its model settings. Open Settings, Models, add a custom OpenAI API endpoint, and set the base URL to your Spark:
Base URL: http://<spark-host>:8000/v1 API Key: not-needed-unless-you-put-a-gateway-in-front Model: Qwen/Qwen3.8-27B-FP8
Cursor's tab-completion and chat features are separate model slots in the settings panel; a custom endpoint typically covers chat and edit requests, while Cursor's own fast-completion model is a separate, more tightly integrated feature that may not be redirectable. Check Cursor's current model settings screen for what's configurable, since this is one of the areas that changes between releases. The model name field has to match what your server reports at /v1/models exactly, or requests fail with an error that looks like an auth problem rather than a naming mismatch.
Continue.dev
The coding agents guide already covers Continue's config.yaml setup in detail, same file, same openai provider entry, whether you're using Continue as a terminal agent or through its IDE sidebar in VS Code or JetBrains. There's nothing IDE-specific to add here: point apiBase at your Spark or your LiteLLM gateway, and the sidebar chat, inline edit, and autocomplete features all use whatever model that config entry names.
Other IDE assistants
Several other editor-integrated assistants accept a custom OpenAI-compatible base URL in their settings, though exact support and menu locations vary by release and are worth checking against the tool's current documentation rather than assuming parity with Cursor or Continue. If an assistant's settings panel has a field for "API base URL" or "custom provider" alongside its API key field, the same pattern applies: point it at your Spark's /v1 path, or at your LiteLLM gateway if you want per-user keys and usage tracking across the team.
JetBrains' own AI Assistant plugin and several VS Code extensions in this category follow the same shape too. The detail worth checking case by case is scope: some plugins route only their chat panel through a custom endpoint while keeping an unrelated feature, spell-check style suggestions, commit message generation, on their own built-in model regardless of what you configure. Read the plugin's own settings documentation for what a custom endpoint actually covers before assuming full coverage.
Rolling this out across a team
A single developer pointing Cursor or Continue at a raw Spark URL is fine for one person. Once several people share the box, put LiteLLM in front rather than handing out the same shared base URL: each developer gets their own virtual key, you can rate-limit a runaway agent loop before it starves everyone else's completions, and revoking access for one person doesn't mean rotating a URL everyone else has hardcoded into their editor settings. The base URL in each editor's config just changes from the Spark's address to the gateway's, and the model alias comes from whatever you named it in config.yaml.
What actually breaks
- Fast-completion features may stay cloud-only. An IDE assistant's chat panel and its inline autocomplete are often separate systems under the hood; redirecting one doesn't necessarily redirect the other. Check what a given custom-endpoint setting actually covers before assuming code stopped leaving your machine.
- The model name has to match exactly. Same failure mode as everywhere else on this site: a mismatch reads as an auth or network error, not a naming one.
- Context gets eaten fast. An assistant that pulls in open files, recent edits, and a system prompt before it even sees your message can burn a large share of the context window before the actual question. Set
--max-model-lengenerously on the server. - Editor updates change settings locations. Both tools ship frequently; a menu path in this guide can drift. If a setting described here isn't where expected, check the tool's own current documentation.