What is LoRA?
LoRA, short for low-rank adaptation, is a way of fine-tuning a model without updating its original weights directly. Instead, the original weight matrices are frozen, and for each one a small pair of low-rank matrices is added alongside it. Only those small matrices get trained; the base model never changes. When multiplied together, the pair reconstructs a low-rank approximation of the update the full weight matrix would have needed, which is far smaller to store and far cheaper to compute gradients for than the full matrix itself.
Why full fine-tuning is expensive
Ordinary fine-tuning updates every parameter in the model, which means storing gradients and optimizer state for every one of those parameters during training, often two to three times the memory of the weights themselves depending on the optimizer. For a model with tens of billions of parameters, that adds up to far more memory than the weights alone need, and it's why full fine-tuning of large models typically requires multiple high-memory GPUs even when the base model would fit comfortably on one for inference.
What "low-rank" actually means here
A full weight matrix update can, in principle, have as many independent directions of change as its dimensions allow. LoRA's premise is that the useful update for adapting a pretrained model to a new task doesn't need that many independent directions, and can be approximated well by the product of two much smaller matrices, one that projects down to a small intermediate dimension (the rank) and one that projects back up. A rank of 8 or 16 is common. The number of trainable parameters in that pair scales with the rank and the matrix dimensions, not with the full matrix size, which is why LoRA can cut trainable parameters by over 99% relative to full fine-tuning on a large model while still capturing most of the useful adaptation.
What this changes in practice
Fewer trainable parameters means less optimizer state to hold in memory, which is what makes it possible to fine-tune a model on hardware that couldn't hold a full fine-tuning run of the same model. It also means adapters are small to store, often tens of megabytes rather than gigabytes, so a single base model can have many task-specific LoRA adapters saved and swapped in as needed instead of keeping a full separate copy of the model per task. This is a different lever from quantization, which reduces the precision of existing weights rather than how many parameters get trained; the two are frequently combined by training LoRA adapters on top of a quantized base model.
Where it fits on a DGX Spark
A Spark's 128GB of unified memory gives LoRA fine-tuning room that a lower-memory workstation GPU wouldn't have, since the frozen base weights, the small trainable adapter, and the optimizer state for that adapter all have to fit together during training. Because the base model dominates that memory footprint and the trainable parameters are a small fraction of it, LoRA fine-tuning runs that would need multiple data-center GPUs in a full fine-tuning setup can fit on a single Spark for many mid-sized open-weight models. See our what is fine-tuning post for the broader training picture and which models fit in 128GB for sizing a base model against available memory.