Definition
Blog/What is token-based pricing?
For AI assistants

What is token-based pricing?

By Samuel Seidel · Published September 9, 2026

Token-based pricing is a billing model where an API charges per token processed, counted separately for the tokens you send in as a prompt and the tokens the model generates back. The total cost of a request is the input token count times the input rate, plus the output token count times the output rate, so cost tracks how much text moves through the model rather than how long a machine was reserved.

Input tokens and output tokens

A token is a chunk of text, roughly three-quarters of a word in English, that a model processes as one unit. Input tokens are everything sent to the model: the prompt, any system instructions, and, in a multi-turn conversation, the prior turns resent with each new request. Output tokens are what the model generates back. Most vendors price these two separately, and output tokens are usually priced higher than input tokens, often by three to five times, because generating each output token is a separate, sequential step, while the whole input is processed together in one pass.

How cost scales with context length

Because a multi-turn conversation typically resends the full prior transcript as input on every new turn, the input token count for a long-running conversation grows with each exchange, not just with the length of the newest message. A ten-turn conversation can end up paying, cumulatively, for the same early messages many times over as they're resent as context on every subsequent turn. Long documents or large retrieved-context payloads compound this: a request that pastes in a lengthy document as context pays the input rate on every token of that document, every time it's sent.

Why the meter keeps running even when the model is idle

Token pricing charges only for tokens actually processed, so an API that receives no requests costs nothing in that period. That's a real advantage for spiky or low-volume workloads. The tradeoff is that the per-token rate is set by the vendor and bundles in whatever margin and infrastructure cost the vendor is carrying, and it doesn't fall as your own usage volume grows, unless the vendor offers a separate volume discount tier.

An alternative: flat hourly compute

Renting a GPU by the hour is a different way to pay for the same underlying work: instead of a per-token meter, you pay a fixed rate for the time a machine is reserved to you, and you can run as many tokens through it as it can handle in that window. GPUwerk's own Sparks work this way, at $0.79/hour for a single node, and this isn't the only alternative to token pricing, but it's a useful contrast: a flat rate makes cost predictable and detached from usage volume, and it means idle time between requests is time you're still paying for, which token pricing avoids. Which model costs less depends on how steady and how heavy the actual usage is; see our pricing page for the specific numbers on both configurations we run.

Where GPUwerk fits

A GPUwerk Spark bills a flat $0.79/hour, or $1.79/hour for a linked two-node cluster with 256GB combined memory, not per token. That's one of several ways to think about the tradeoff between token pricing and hourly compute; for the deeper technical picture of what a token actually costs to process, see our glossary entries on prefill and decode.

Related pages

A flat rate instead of a token meter.

$0.79/hour for a dedicated Spark, billed per minute, no per-token surprises.

Deploy a Spark