What is a token (in LLMs)?
A token is the basic unit of text a language model reads and generates. Before a model sees any text, a tokenizer splits it into a sequence of tokens, each mapped to a number, and the model works entirely in that numeric sequence rather than on raw characters or words.
Tokens are usually pieces of words
Most modern tokenizers use a method called byte-pair encoding or a close variant, which builds a vocabulary of common substrings from a large body of text. Frequent short words end up as single tokens; longer or rarer words get split into two or more pieces. The word "token" itself is one token in many vocabularies; "tokenization" might split into "token" and "ization". Spaces, punctuation, and even parts of numbers count as tokens too. A rough rule of thumb for English is that one token is about four characters, or about three-quarters of a word, though the exact ratio varies by tokenizer and language.
Why models work on tokens instead of raw text
A fixed, finite vocabulary of tokens (typically tens of thousands of entries) lets a model represent any input text as a bounded sequence of numbers, which is what the underlying neural network actually operates on. Working at the character level would make sequences much longer and harder to model efficiently; working at the whole-word level would require an enormous vocabulary to cover every word form and would still fail on words never seen during training. Sub-word tokenization is the compromise that's become standard.
Why tokens matter for cost and context
Two things about running an LLM scale directly with token count. First, cost: nearly every inference API and GPU rental bills by tokens processed, since token count tracks the actual compute work involved. Second, context window: a model can only attend to a fixed number of tokens at once (the "context length" advertised for a given model), and that budget covers the prompt, any documents fed in, and the reply combined. Long documents or long conversation history all draw from the same token budget, which is why techniques like retrieval-augmented generation exist to fetch only the relevant text instead of stuffing everything into context.
Where token processing actually happens
Every token in a request has to pass through the model's layers on GPU memory, and every token generated in a reply is produced one at a time, each depending on all the tokens before it. That's why GPU choice matters for LLM workloads: more memory means a longer context fits, and more memory bandwidth means tokens generate faster. On a DGX Spark, 128GB of unified memory is enough to hold a large model's weights plus a substantial amount of token cache, which is the practical limit on how long a conversation or document context a self-hosted model can handle.
Where GPUwerk fits
GPUwerk rents dedicated Spark hardware for running your own models rather than metering by API token, so cost is a flat $0.79/hour for a single Spark or $1.79/hour for a two-node cluster, regardless of how many tokens you push through it. For workloads with heavy, sustained token volume, that flat-rate model can work out cheaper than per-token API pricing; see private LLM hosting for how that comparison plays out in practice.