Coding agents against your own endpoint
Every popular editor agent, Continue, Cline, aider, was built against OpenAI's API shape, which means every one of them can be pointed at a self-hosted endpoint by changing a base URL. The code you paste into an agent is some of the most sensitive text you'll ever send anywhere; it doesn't have to leave your own machine.
Start with a running endpoint
Every agent below just needs an OpenAI-compatible /v1 endpoint and a model name. Get that running first, either engine works:
- vLLM if more than one person on the team will share it.
- llama.cpp if it's just you, it's noticeably faster for a single stream.
If the endpoint isn't reachable from your laptop directly, either bind it to a public port on the instance (fine on a dedicated box you control) or tunnel it over SSH per the SSH keys guide. Confirm it works before touching any editor config:
curl http://<spark-host>:8000/v1/models
Continue (VS Code / JetBrains)
Continue reads its model list from config.yaml (or the legacy config.json). Add an openai-compatible provider entry pointed at your instance:
models:
- name: My Spark
provider: openai
model: Qwen/Qwen3.8-27B-FP8
apiBase: http://<spark-host>:8000/v1
apiKey: not-needed-unless-you-put-a-gateway-in-front
If you put LiteLLM in front of the engine for per-user keys and budgets, point apiBase at the gateway instead and set a real apiKey. The model name has to match whatever the server actually reports at /v1/models, or whatever --served-model-name you set.
Cline (VS Code)
Cline's settings panel has an "OpenAI Compatible" provider option directly. Set the base URL to your endpoint, paste any string as the API key if nothing is enforcing one, and enter the model ID exactly as the server reports it. Cline is more agentic than Continue by default, longer tool-use chains, more file edits per turn, so it leans harder on context length. Make sure whatever you're serving isn't capped at a short --max-model-len.
aider (terminal)
aider talks to any OpenAI-compatible endpoint via environment variables, no config file needed for the simple case:
export OPENAI_API_BASE=http://<spark-host>:8000/v1 export OPENAI_API_KEY=not-needed-unless-you-put-a-gateway-in-front aider --model openai/Qwen/Qwen3.8-27B-FP8
The openai/ prefix on --model tells aider's underlying LiteLLM client which provider format to speak, it isn't optional even though there's no OpenAI involved.
Which model to run
Coding-specific fine-tunes earn their keep here over general chat models. Qwen3-Coder 30B-A3B is the one staged on every GPUwerk node for this reason: strong on real repositories, small enough to leave most of a Spark's 128 GB for context and concurrent editor sessions. See which models fit in 128 GB for the full picture, including 70B-class options if 30B isn't enough for your codebase.
What actually breaks
- Context gets eaten fast. An agent that reads several files before it edits one can burn tens of thousands of tokens per turn. Set
--max-model-lengenerously on the server, not at the editor's default. - Streaming has to actually work. All three agents expect Server-Sent Events streaming from
/v1/chat/completions; both vLLM and llama.cpp support it by default, but a gateway or proxy in between can silently buffer it and make every response feel like it hangs. - The model name has to match exactly. A mismatch between what the agent sends and what the server has loaded returns a 404 that looks like a network problem, not a config problem, and costs more debugging time than anything else on this page.