What is fine-tuning?
Fine-tuning is a further round of training applied to a model that's already been pretrained, using a smaller and more specific dataset to adjust its weights toward a particular task, style, or domain. The model comes out of the process able to do something the base model couldn't reliably do, because its internal parameters have actually changed, not because it's being fed extra instructions at request time.
What actually happens during fine-tuning
A pretrained model already has weights learned from a large, general training run. Fine-tuning takes that starting point and runs additional gradient updates against a new, narrower dataset, an example set of the input-output pairs you want the model to get better at. Those updates move the weights, sometimes across the whole model, sometimes only through a small set of added parameters (a common efficient approach called LoRA). Either way, the result is a model whose parameters differ from the base checkpoint you started with.
How it differs from prompting and retrieval
Fine-tuning is often confused with, or compared against, retrieval-augmented generation, but the mechanisms are different. RAG leaves the model's weights untouched and instead retrieves relevant documents at query time, inserting them into the context the model reads before answering. Fine-tuning changes what the model itself has learned; RAG changes what it's shown. The two solve different problems and are frequently combined rather than chosen between. We cover that comparison, and when each is the better fit, in fine-tuning vs RAG.
What fine-tuning is good at
Fine-tuning tends to help with consistent style or format, domain-specific terminology, and tasks where the desired behavior is hard to fully specify in a prompt but easy to demonstrate with examples. A support team that wants every response to follow a specific tone and structure, or a codebase with unusual internal conventions, are typical cases. It doesn't add new factual knowledge as reliably as it changes behavior; a model fine-tuned on last year's documentation doesn't automatically know this year's, in the way a retrieval step reading the current docs would.
What it costs, roughly
Fine-tuning a smaller open-weight model with a parameter-efficient method is within reach of a single GPU for hours rather than days, though the exact time depends heavily on model size, dataset size, and method. Full fine-tuning of a large model's entire parameter set is far more compute-intensive and usually isn't practical on a single consumer or workstation-class GPU. Check the specific method and model size before assuming either extreme.
Full fine-tuning vs parameter-efficient methods
Full fine-tuning updates every weight in the model, which gives the most flexibility but needs enough memory to hold the full optimizer state alongside the weights, often several times the size of the model itself. Parameter-efficient methods, LoRA being the most common, freeze the original weights and train a much smaller set of added parameters instead, cutting memory requirements sharply at some cost to how much the model's behavior can shift. Most fine-tuning done on a single machine today uses a parameter-efficient method for exactly that reason, since it makes adapting a mid-sized model practical on one GPU rather than a cluster.
Where GPUwerk fits
A GPUwerk Spark is sized for running and fine-tuning models in the smaller-to-mid range that fit within 128GB of unified memory, at $0.79/hour from EU-Central. It's not built for full-parameter fine-tuning of frontier-scale models; for that scale of training job, a multi-GPU cluster is the right tool. See inference vs training for where that line sits.