Data, Pre-Training & Post-Training · 6 min

Fine-Tuning on a Budget: LoRA and QLoRA

Why full fine-tuning burns memory you don't need to spend, and how low-rank adapters plus 4-bit quantization put a 65B tune on one GPU.

You have a base model that already writes fluent English and holds a conversation. You want it to do one narrow thing well: sort support tickets into your taxonomy, answer in your house voice, reason over your internal jargon. Full fine-tuning gets you there, but the bill is brutal and most of that bill buys you nothing. Parameter-efficient fine-tuning (PEFT) is the family of tricks that skips the waste. LoRA and its quantized cousin QLoRA are the two that won.

Why full fine-tuning costs so much

The parameter count is the least of it. The optimizer is what kills you.

Train a 7B model with Adam in mixed precision and every parameter drags a small entourage through memory: a 16-bit weight for the forward pass, a 32-bit master copy of that weight, Adam's two 32-bit moment estimates (m and v), and a gradient. That lands around 16 to 20 bytes per parameter before you spend a single byte on activations. So 7B parameters means roughly 112 to 140 GB of GPU memory just to hold the training state. A model that happily runs inference on a 16 GB card now wants two or three 80 GB A100s to train. Every checkpoint you save is another full copy of all 7B weights, tens of gigabytes apiece.

The waste is structural. You compute a gradient and carry optimizer state for every weight, when the adaptation you actually need is small.

The LoRA hypothesis: updates are low-rank

Hu et al. (2021) started from one observation. When you fine-tune, the change to a weight matrix, call it ΔW, has low intrinsic rank. The update lives in a much smaller subspace than the full matrix can hold. If that is true, you never need ΔW as a dense d × k block. You can factor it into two skinny pieces and train only those.

LoRA freezes the pretrained weight W₀ and represents the update as a product of two low-rank matrices:

   W₀ (frozen)            output
   d × k                    ▲
   ┌────────────┐           │  h = W₀x  +  (α/r)·B·A·x
   │            │           │            └──┬──┘ └┬┘
   │  no grad   │──────────►│              B·A    only these train
   │            │           │
   └────────────┘           │
        x ──────────────────┘
                B: d × r   A: r × k     r ≪ min(d,k)

A is r × k, B is d × r, and the rank r is tiny, usually 4, 8, or 16. Take a 4096×4096 attention matrix, about 16.8M parameters. A rank-8 adapter is 4096·8 + 8·4096 = 65,536 parameters, roughly 0.4% of the original. W₀ stays frozen, gradients flow only into A and B, and Adam only tracks moments for those two. The optimizer memory that dominated your budget shrinks by the same ratio.

Two details trip people up.

Initialization is not cosmetic. A gets small random Gaussian values, B starts at zero. So BA = 0 at step one and training begins at exactly the base model's behavior, with no random perturbation to claw back from.

The α/r term is a fixed scale, not a learned one. It decouples your effective learning rate from your choice of rank, so you can raise r without re-tuning everything downstream. A common convention sets α = 2r.

LoRA usually adapts the attention projection matrices. Hu et al. get their headline results tuning just the query and value weights, though modern practice often adapts every linear layer. On GPT-3 175B the numbers are the pitch: 10,000× fewer trainable parameters and 3× less GPU memory than full Adam fine-tuning, with quality on par or better. The trained adapter drops from hundreds of gigabytes to tens of megabytes.

Because BA is just a matrix, you can fold it back in after training as W = W₀ + (α/r)BA. The merged model runs at zero extra inference latency, unlike the older adapter-layer methods that bolted extra sequential compute onto every forward pass.

Builder: If you serve many tasks, keep adapters unmerged in production. One frozen base in VRAM plus a stack of few-MB adapters, hot-swapped per request, beats loading N separate full fine-tunes. That is the whole premise of multi-LoRA serving.

QLoRA: shrink the frozen part too

LoRA cuts gradient and optimizer memory. But the frozen base still sits in memory at 16 bits: 14 GB for 7B, about 130 GB for 65B. You cannot tune a 65B model on one GPU because the frozen weights alone will not fit.

Dettmers et al. (2023) closed that gap with a simple lever. The base is frozen, so during training its weights are read-only. You never take a gradient step on them, so they never need training-grade precision. QLoRA stores the frozen base at 4 bits, keeps the LoRA adapters in higher precision, and backpropagates through the quantized base into those adapters. Three pieces keep the quality:

NF4 (4-bit NormalFloat) places its 16 quantization levels at the quantiles of a normal distribution, which the paper argues is information-theoretically optimal for weights that are roughly Gaussian. It spends resolution where weights actually cluster instead of spreading levels evenly across a range that is mostly empty.

Double quantization quantizes the quantization constants themselves, which are numerous enough to matter. It saves about 0.37 bits per parameter, roughly 3 GB on a 65B model, effectively for free.

Paged optimizers use NVIDIA unified memory to spill optimizer state to CPU RAM during the memory spikes (long sequences, gradient checkpointing) that would otherwise crash the run.

The payoff: fine-tune a 65B model on a single 48 GB GPU while matching 16-bit fine-tuning quality. Their Guanaco models reach 99.3% of ChatGPT's score on the Vicuna benchmark after 24 hours on one GPU. That result is a big reason "just fine-tune it yourself" became something individuals do, not only labs.

The cost is time. Forward passes de-quantize NF4 back to bf16 to compute, so each step is slower than plain LoRA on a full-precision base. You trade wall-clock for fitting on hardware you can actually rent.

What actually goes wrong

Rank set too low starves the task. A new domain (not just a new style) genuinely needs more capacity, and r=4 will underfit it. Sweep 8, 16, 32 and watch eval loss instead of copying a number off a blog post.

Learning rate is the quiet killer. Adapters tolerate, and usually want, a higher LR than full fine-tuning, often in the 1e-4 range. Reuse a full-FT schedule and they come out undertrained.

Quantization is lossy. The 4-bit base is not the fp16 base. QLoRA recovers the gap on most tasks, but for precision-sensitive work you should measure against a LoRA-on-fp16 baseline before assuming they are interchangeable.

Merging into a quantized base is a trap. Adapters are trained against 4-bit weights, so merge them into the fp16 base and keep the adapter separate. Bake them into the quantized weights and you lock in the quantization error.

The security angle

Adapters are small, portable, and increasingly traded on model hubs, which turns them into a supply-chain surface. A LoRA adapter is a weight delta, and applying an untrusted one is running untrusted code against your model's behavior. A poisoned adapter can carry a backdoor: benign on normal input, attacker-controlled on a trigger phrase, packed into a few-MB file no signature scanner understands. No diff will tell you what a rank-16 delta actually does.

Defender: Treat third-party adapters like unsigned binaries. Pin provenance, evaluate on a red-team trigger set before you deploy, and prefer training your own from a base whose hash you verified.
Researcher: Quantization can move safety behavior on its own. Refusal boundaries learned in fp16 do not always survive a 4-bit round-trip. If you QLoRA-tune, re-run your alignment and refusal evals on the quantized model, not the fp16 one you started from.

This loops back to the pretraining lesson. PEFT only steers what pretraining already put there. It reweights and redirects existing capability cheaply; it does not install knowledge the base never learned. When a LoRA fine-tune "won't learn" your domain, the base usually lacks the substrate, and no rank you turn up will conjure it. That is a data-and-pretraining problem wearing a fine-tuning costume.

Sources