Open-weight vs closed-weight models: how to choose
This isn't an argument for open models over closed ones, or the other way round. Both are the right answer for different projects, and the choice usually comes down to three questions: how much do you need to control where the data goes, how much are you actually going to spend once volume grows, and how much do you need to change about how the model behaves. Answer those three and the rest of the decision mostly falls out.
What "open-weight" actually means here
An open-weight model ships its trained parameters for anyone to download and run. You can host it yourself, on your own hardware or a rented GPU, and nothing about a request to it leaves your infrastructure unless you send it somewhere. Qwen, Llama, gpt-oss, and DeepSeek's open releases are all open-weight in this sense, and they vary a lot in size and capability, from models that fit on a laptop to ones that need a rack.
A closed-weight model is one you can only reach through an API the provider runs. You send a request over the network, their servers run it, and you get a response back. You never hold the weights, can't run the model somewhere else, and every request is subject to whatever data handling and retention policy the provider has published, or hasn't.
When the API wins: convenience and frontier capability
A closed API is the right default when you're building something and don't yet know if it'll work. There's no hardware to provision, no engine to configure, and no capacity planning; you get an API key and start sending requests. If the project doesn't pan out, you've spent nothing beyond the tokens you used. For a prototype, an internal tool with light and unpredictable usage, or a product that hasn't found its usage pattern yet, that flexibility is worth more than anything self-hosting offers.
APIs also tend to front-run open weights on frontier capability. The labs that publish the strongest closed models generally don't open-weight their top tier, so if your task genuinely needs the best available reasoning, the newest multimodal capability, or the largest context window on the market, a closed API is often the only way to get it, at least for a while after release.
The tradeoff is that every request leaves your infrastructure and goes through a third party's systems, under their retention and training-use policy, whatever it currently says. For a lot of workloads that's a non-issue. For anything involving customer data, regulated information, or material you don't want anywhere near another company's training pipeline, it's the reason to look elsewhere.
When self-hosted open-weight wins: control, cost at scale, customization
Data control is the clearest case. If a request never leaves hardware you operate, there's no third-party retention policy to trust, no data processing agreement to negotiate, and no question about whether a provider's terms changed since you last read them. For legal, healthcare, financial, or any customer data that can't leave the building, self-hosting an open-weight model on your own or rented infrastructure is often the only option that satisfies the requirement at all, not just the cheapest one.
Cost is the second case, and it only favors self-hosting past a certain volume. An API has no fixed cost: you pay per token and nothing when you're idle. Self-hosting has a fixed hourly cost whether you use it or not, so at low or spiky volume the API is cheaper. Past a high enough steady volume, the math flips, because the fixed hourly rate stops scaling with usage while the API's per-token bill keeps climbing linearly. Where exactly that crossover sits depends on your token volume and the specific API pricing you're comparing against, and it's worth actually calculating for your workload rather than assuming either direction.
Customization is the third case. Fine-tuning, quantizing to a format that fits your hardware, or modifying the serving stack itself all require access to the weights, which a closed API structurally doesn't give you. If your project needs a model tuned on your own data, or needs to run at a specific latency or memory footprint that only a custom quantization achieves, open-weight is the only path, independent of cost or data control.
A short framework
Ask, in order: Does the data have to stay off third-party infrastructure? If yes, self-host an open-weight model. Do you need to fine-tune, quantize, or otherwise modify the model itself? If yes, same answer. Is your volume high and steady enough that a fixed hourly rate beats a per-token bill, and have you actually run that calculation? If yes, self-hosting saves money too. If none of those apply, and you mainly want to prototype quickly or need the single strongest available model for a hard task, a closed API is the better starting point, and nothing stops you from moving to open-weight once the project's shape and volume are clearer.
Most teams that end up self-hosting didn't start there. They started on an API, learned what their actual usage pattern and data sensitivity looked like, and moved once one of the three factors above became the deciding one. That's a reasonable order to go in, as long as the integration is built against a portable interface rather than one provider's proprietary SDK, so the move doesn't mean a rewrite.