Fine-tuning LLMs in Practice
Page 10 covers what SFT, LoRA and DPO are. This page is the part that gets you a working adapter: the memory arithmetic that decides the method, what rank and alpha do, what data to feed it, what it quietly costs you, and how to tell whether any of it helped.
Most fine-tuning projects should not have been fine-tuning projects. Work up the ladder first: prompt, few-shot, then retrieval. Fine-tune when you need behaviour the base model will not reliably produce, such as a house style, a rigid output format, or a domain's idiom. Once you commit, your VRAM chooses the method: QLoRA on a quantized base is what fits on one consumer GPU.
The arithmetic behind that: mixed-precision AdamW costs 16 bytes per trainable parameter, so a full 7B fine-tune is 108 GB while a rank-16 adapter on a 4-bit base is under 4. Three of those four bytes-per-parameter lines vanish when you freeze the base, which is the entire mechanism.
Your dataset decides the outcome far more than any hyperparameter. Match the base model's chat template exactly, hold out a real eval set before you start, and remember that falling training loss is evidence of memorisation and says nothing yet about usefulness. One sentence for an interview: fine-tuning reliably changes how a model behaves and unreliably changes what it knows.
01The decision before the decision
Fine-tuning is rarely the first correct move. It is the most expensive rung on a ladder, and most teams reach for it two rungs too early.
Climb in this order. Stop at the first rung that solves your problem.
- Prompting. Free, instant, and revisable in seconds. Exhaust it properly before concluding it failed.
- Few-shot examples. Still just prompting. Costs context window, but often fixes format problems outright.
- Retrieval (RAG). The right answer whenever the issue is knowledge: updating a fact means updating a document, with no retraining. See 11.
- Fine-tuning. Changes the weights. Slowest loop, hardest to undo, and the only option that reshapes default behaviour.
What fine-tuning earns its cost on:
- Style and tone. Making every response sound like your organisation wrote it.
- Format adherence. Producing valid JSON, or a fixed schema, without a paragraph of instructions each time.
- Domain idiom. Absorbing how a specialist field phrases things, which prompting mimics only shallowly.
- Cost and latency. Teaching a small model one narrow job well, so you stop paying for a large one.
What it does not fix. It will not make a weak base model competent, and no amount of adapter training rescues a poor architecture choice. It is also a bad mechanism for facts that change.
| Prompting | Few-shot | Retrieval | Fine-tuning | |
|---|---|---|---|---|
| Time to try a change | Seconds | Seconds | Minutes — reindex | Hours to days — retrain, then re-evaluate |
| Cost of fixing one wrong fact | Edit a sentence | Edit an example | Edit a document | Rebuild the dataset and train again |
| Can point at its source | No | No | Yes | No |
| Per-request token cost | Grows with the instruction | Grows fastest | Grows with retrieved context | Near zero — the behaviour is in the weights |
| What it fixes | Most things, if you try properly | Output format, usually | Knowledge, freshness, provenance | Default behaviour: style, format, idiom |
Read the bottom row against the rightmost column. Fine-tuning's one real advantage over everything to its left is that the behaviour becomes the default and stops costing tokens on every call. Every other row in that table is an argument for trying a cheaper rung first.
Would a perfect prompt fix this? If the model produces the right output when you ask carefully and fails when you ask casually, the problem is consistency — which is a genuine fine-tuning case. If it cannot produce the output at all, fine-tuning will not teach it to.
Will the right answer be different next quarter? Then it belongs in a document. Weights give you no way to find what you baked in, no way to correct one fact, and no citation when someone asks where an answer came from.
Could you write two hundred examples you would be happy to ship? If not, you do not have a dataset yet. The run will faithfully learn whatever inconsistency you have, and you will read the result as a modelling failure.
Unsloth's own fine-tuning guide argues that fine-tuning can replicate everything RAG does, since it changes the weights and RAG cannot. The argument is technically defensible, because weights do store knowledge.
The practitioner objection is operational. A fact stored in weights cannot be updated without another training run, carries no citation, and gives you no way to check where an answer came from, while a fact stored in a document is edited in seconds and cites itself.
So the useful rule is not "fine-tuning cannot learn facts." It is: put the parts that change in retrieval, and the parts that never change in the weights.
02Choosing the method: your VRAM decides
Three options, and in practice the hardware sitting in front of you usually makes the choice. Page 10 covers why LoRA works mathematically. Here is when to use which.
| Full fine-tune | LoRA | QLoRA | |
|---|---|---|---|
| What updates | Every weight | Small low-rank adapters, base frozen | Same adapters, on a 4-bit quantized base |
| Relative memory | Highest: weights, gradients and optimizer state for every parameter | Much lower: optimizer state only for adapter weights | Lowest: adapters plus a quantized frozen base |
| Reach for it when | You have multi-GPU budget and are changing the model deeply | The full-precision base fits comfortably in VRAM | The base only fits when quantized, which is the common case on one GPU |
| Main cost | Expensive, and you get one model per task | Slight quality gap versus full fine-tuning on some tasks | Quantizing the base adds its own small quality cost |
Adapters have a property worth noticing. Because the base stays frozen, one base model can serve many tasks by swapping small adapter files. Full fine-tuning gives you a whole separate model per task.
BA can be folded into W, so the served model has exactly the original shape and no extra latency.The fourth family: prompt and prefix tuning
LoRA is not the only way to train a small number of parameters, and the alternatives are worth studying because they fail in an instructive direction.
- Prompt tuning (Lester et al., 2021) freezes the entire network and learns a short sequence of continuous vectors prepended to the input embeddings. Nothing inside the model moves. Their finding: it only caught up with full model tuning at T5-XXL, 11B parameters.
- Prefix tuning (Li & Liang, 2021) learns vectors at every layer and not only at the input, which is why it holds up at smaller scales. The paper reports comparable performance while training about 0.1% of the parameters.
Both spend context window at inference, and neither can be folded back into the weights — so they add latency that LoRA does not. They earn their place when the base must stay literally untouched, or when you cannot change the serving stack at all.
Starting hyperparameters
These are starting points, and you should expect to move them. They come from Unsloth's LoRA hyperparameters guide.
- Rank
r. Start at 16 or 32. Common values run 8 to 128. Higher rank adds capacity, and also adds overfitting risk. lora_alpha. Set it equal tor, or to2rto learn more aggressively.lora_dropout. Zero by default. It does little on short runs. Raise it only if you see overfitting.target_modules. Target all seven linear layers. Omitting some saves very little memory and measurably costs quality.- Learning rate.
2e-4is the usual starting point for LoRA and QLoRA. - Epochs. One to three. Beyond three, the guide reports diminishing returns and rising overfitting risk.
03The memory arithmetic, line by line
One 24 GB card, one 7B model. Full fine-tuning needs about 108 GB before a single activation exists; while QLoRA needs under 4. Both numbers fall out of the same five-line ledger, and it is worth walking through once instead of memorising the conclusion.
Worked example Where the 108 GB in a full 7B fine-tune actually goes
Llama-2-7B has 6,738,415,616 parameters. Assume the standard recipe: bf16 forward and backward, AdamW, no sharding. Every line is bytes per parameter.
-
$$\text{weights} = 2 \times 6.74\times10^{9} = 13.5\ \text{GB}$$Two bytes per parameter, because the matmuls run in bf16 to reach the tensor cores. This is the only line most people count, and it is an eighth of the bill.
-
$$\text{gradients} = 2 \times 6.74\times10^{9} = 13.5\ \text{GB}$$One gradient per trainable parameter, in the dtype the backward pass produced it in. This line is what LoRA deletes: a frozen weight has no gradient.
-
$$\text{fp32 master copy} = 4 \times 6.74\times10^{9} = 27.0\ \text{GB}$$bf16 carries seven explicit mantissa bits, so its spacing near a value is about one part in 256. Add a relative update of $10^{-4}$ to a bf16 weight and it rounds back to where it was. So the optimizer keeps an fp32 copy, updates that, and casts down for the next forward pass.
-
$$\text{Adam } m + v = (4 + 4) \times 6.74\times10^{9} = 53.9\ \text{GB}$$A running first and second moment per parameter, both fp32. Four times the size of the weights they are optimising, and the single largest line in the ledger.
-
$$(2 + 2 + 4 + 8) \times 6{,}738{,}415{,}616 = 107.8\ \text{GB}$$Sixteen bytes per parameter, the same accounting page 28 uses for distributed training. Adding the four rounded lines above gives 107.9 rather than 107.8 — each one rounds up, and four of those add a tenth. The exact figure is 107.81 GB. Activations are on top of this and depend on batch and sequence length.
So LoRA's memory saving doesn't come from the adapter being small. It comes from three of the four lines being indexed by trainable parameters and not by total parameters, and a rank-16 adapter makes that number about 0.6% of the model.
How small exactly? Each targeted matrix of shape $d_{\text{out}} \times d_{\text{in}}$ gets an $A$ of shape $r \times d_{\text{in}}$ and a $B$ of shape $d_{\text{out}} \times r$, so it costs $r(d_{\text{in}} + d_{\text{out}})$ parameters instead of $d_{\text{in}} d_{\text{out}}$. Summed over Llama-2-7B's seven projections and 32 layers, at $r = 16$, that is 40.0M.
Try it The whole ledger, for five configurations of the same model
# Llama-2-7B geometry, straight from its config.json.
L, H, FF = 32, 4096, 11008
BASE = 6_738_415_616 # its real parameter count
SHAPE = {"q": (H, H), "k": (H, H), "v": (H, H), "o": (H, H),
"gate": (H, FF), "up": (H, FF), "down": (FF, H)}
def lora_params(r, targets):
"""A is r x d_in, B is d_out x r, for every targeted matrix."""
return L * sum(r * (SHAPE[t][0] + SHAPE[t][1]) for t in targets)
ALL7 = list(SHAPE)
QV = ["q", "v"]
print("%-22s %9s %8s %7s %7s %8s"
% ("", "trainable", "weights", "grads", "optim", "total"))
for name, train, bytes_per_w in [
("full fine-tune bf16", BASE, 2),
("LoRA r=16, q+v", lora_params(16, QV), 2),
("LoRA r=16, all 7", lora_params(16, ALL7), 2),
("LoRA r=64, all 7", lora_params(64, ALL7), 2),
("QLoRA r=16, all 7", lora_params(16, ALL7), 0.5)]:
w = BASE * bytes_per_w / 1e9
g = train * 2 / 1e9 # bf16 gradient per trainable
o = train * 12 / 1e9 # fp32 master + Adam m + Adam v
print("%-22s %8.1fM %7.1fG %6.2fG %6.2fG %7.1fG"
% (name, train / 1e6, w, g, o, w + g + o))
print("\nactivations are on top of all of these and depend on"
"\nbatch, sequence length and gradient checkpointing.")
grads and optim columns down the rows. Full fine-tuning spends 13.48 and 80.86 GB; every LoRA row spends under half a gigabyte on both, because those columns are priced by trainable parameters. Now compare weights across the last two rows: identical adapters, 13.5 against 3.4 GB, purely from quantizing the frozen base. Two independent levers, and QLoRA is both of them at once.One more thing that table settles: rank is cheap. Going from r=16 to r=64 costs 1.7 GB, which you can afford on a 24 GB card, so the reason to keep rank low is overfitting risk and not memory.
Does it fit in 24 GB?
States are 3.9 GB for QLoRA — 3.4 for the 4-bit base plus roughly 0.1 GB of quantization constants, and 0.56 for the adapters, their gradients and their optimizer state. Then activations, which is the term people forget and the term that scales with your sequence length.
With gradient checkpointing on, you keep one hidden state per layer boundary: $32 \times 2048 \times 4096 \times 2 = 0.54$ GB at batch 1. Add the logits and their fp32 copy for the loss, about 0.39 GB, plus the recomputation peak inside one layer, and call it 1.2 GB.
Five gigabytes, on a 24 GB card. That headroom is what you spend on batch size and sequence length, both of which multiply the activation term and leave the other 3.9 GB untouched. LoRA on a bf16 base comes to about 15 GB by the same arithmetic, which also fits — with far less room to grow the sequence.
04What LoRA is doing
Page 10 states the equation. This section is about what each symbol in it controls, because r and lora_alpha are the two knobs you will turn and they do completely different things.
- W the pretrained weight, frozen: no gradient, no optimizer state, and after quantization not even full precision.
- A initialised from a Gaussian. It projects the input down to $r$ dimensions.
- B initialised to exactly zero, so $BA = 0$ at step 0 and training starts exactly at the pretrained model.
- α/r a fixed scalar that is not learned. It scales the whole update.
The zero initialisation of $B$ is a small detail with a large consequence: the first forward pass of fine-tuning is bit-identical to the base model's. There is no warm-up period where the adapter is random noise degrading a working model.
Why a low rank is enough
The claim is about $\Delta W$, the change that adapting to one task requires. $W$ itself is plainly not low rank.
The evidence predates LoRA. Aghajanyan et al. (2020) measured the intrinsic dimension of fine-tuning directly: RoBERTa-Large reaches 90% of full fine-tuning performance on MRPC while optimising 200 parameters inside a random low-dimensional subspace, and 774 on QQP. That is a 355M-parameter model adapted by a few hundred numbers.
LoRA's bet is that if the useful update lives in a subspace that small, you can just parameterise a low-rank one directly and skip the random projection. You can watch the mechanism on a matrix small enough to take the SVD of.
Try it Train a full weight matrix on a narrow task, then look at the spectrum of what changed
import numpy as np
rng = np.random.default_rng(0)
d, k = 128, 3
W0 = rng.normal(0, 0.05, (d, d)) # the "pretrained" weight
# The new task only exercises a k-dimensional slice of the input
# and only moves a k-dimensional slice of the output.
U = np.linalg.qr(rng.normal(size=(d, k)))[0]
V = np.linalg.qr(rng.normal(size=(d, k)))[0]
W_task = W0 + U @ rng.normal(size=(k, k)) @ V.T
W = W0.copy()
for _ in range(600): # plain full-matrix SGD
X = rng.normal(size=(64, d))
Y = X @ W_task + rng.normal(0, 0.02, (64, d)) # noisy labels
W -= 0.05 * X.T @ (X @ W - Y) / 64
s = np.linalg.svd(W - W0, compute_uv=False)
energy = np.cumsum(s ** 2) / np.sum(s ** 2)
print("delta W is a full %d x %d matrix (%d numbers)"
% (d, d, d * d))
print("its top 8 singular values:", np.round(s[:8], 3))
print("energy captured by rank 1..5:", np.round(energy[:5], 4))
print("rank needed for 99% of it:",
int(np.searchsorted(energy, 0.99)) + 1,
" -- LoRA would store %d numbers at that rank"
% (2 * d * (int(np.searchsorted(energy, 0.99)) + 1)))
0.435, then 0.009. Nothing constrained this optimisation — all 16,384 entries of a dense matrix, noisy labels, ordinary gradient descent — and the update it found is still rank 3, with 768 numbers reproducing 99.98% of it. Be clear on what that shows: the task is rank 3 by construction, so this is the mechanism. The evidence is Aghajanyan's measurement above.What r and alpha each control
They are routinely confused, including in tutorials. r is a capacity knob: it sets the rank of $\Delta W$ and therefore how many independent directions the update can move in. alpha is a multiplier and adds no capacity.
Try it Confirm which knob changes the rank and which one only changes the scale
import numpy as np
d = 256
W = np.random.default_rng(0).normal(0, 0.02, (d, d))
def delta_w(r, alpha, seed, zero_B=False):
g = np.random.default_rng(seed)
A = g.normal(0, 0.01, (r, d)) # Gaussian init
B = np.zeros((d, r)) if zero_B else g.normal(0, 0.01, (d, r))
return (alpha / r) * (B @ A) # the effective dW
# 1. B starts at exactly zero, so a fresh adapter is a no-op.
print("fresh adapter leaves W untouched:",
np.array_equal(W + delta_w(16, 32, 1, zero_B=True), W))
# 2. r is the capacity knob: it sets the RANK of the update.
for r in (1, 4, 16, 64):
dw = delta_w(r, 16, 2)
print("r=%-3d rank(dW)=%-3d stored %6d numbers (%5.2f%% of W)"
% (r, np.linalg.matrix_rank(dw), 2 * r * d,
100 * 2 * r * d / d ** 2))
# 3. alpha is not a capacity knob. It is a plain multiplier.
a16, a32 = delta_w(16, 16, 3), delta_w(16, 32, 3)
print("alpha=32 is exactly 2x alpha=16:", np.allclose(a32, 2 * a16))
print("same rank either way:", np.linalg.matrix_rank(a16),
np.linalg.matrix_rank(a32))
alpha=32 is exactly 2x alpha=16, at identical rank. Doubling alpha is arithmetically the same as doubling $B$, so it acts as a learning-rate multiplier on the adapter, not as extra capacity. Setting $\alpha = r$ or $2r$ holds the effective scale $\alpha/r$ fixed as rank changes, so the learning rate survives. And note r=64 storing half of a 256-wide $W$ — the same rank on a 4096-wide layer is 3.1%.Which layers to target
Two published findings that look contradictory and are not.
- The LoRA paper's ablation, at a fixed parameter budget, found that spending it on $W_q$ and $W_v$ beat spending it on any single projection — and that a rank as low as 1 sufficed in that setting.
- The QLoRA paper found that matching full fine-tuning required applying LoRA to all linear layers in the transformer block, beyond the attention projections.
They answer different questions. If the budget is binding, attention's query and value projections are where it buys the most. But the ledger in section 03 shows the budget usually is not binding: all seven projections at $r=16$ cost 0.56 GB of gradients and optimizer state on a 24 GB card.
So the practical default is the second finding. Target everything, keep the rank modest, and treat the attention-only configuration as what you fall back to when memory is tight — which, after section 03, it rarely is.
Because $\Delta W = \frac{\alpha}{r}BA$ is a matrix of exactly $W$'s shape, it can be added into $W$ once and thrown away. The served model has the original architecture, the original parameter count and the original latency. Prompt and prefix tuning cannot do this.
Or you keep it separate: a 40M-parameter adapter is roughly 80 MB in bf16, so one base model in memory can serve many tasks by swapping files. That also means you can turn a fine-tune off, which is the cheapest catastrophic-forgetting insurance available.
05Data
Hyperparameters move results by a little and the dataset moves them by a lot, so budget your time accordingly.
A supervised fine-tuning example is a conversation: an instruction, optionally a system message, and the response you wish the model had given. The model trains to predict the response tokens. The objective is the same next-token objective from 10, pointed at a much smaller and much more deliberate pile of text.
The chat template trap
Every instruct model was post-trained with a specific chat template: the exact special tokens and markers that separate system, user and assistant turns. Llama's template differs from ChatML, and ChatML differs from Gemma's.
If you train with a template the base model has never seen, you teach it a new formatting convention at the same time as your actual task, and both suffer. It is the most common way a first fine-tune quietly underperforms.
Symptoms of a template mismatch include the model rambling past where it should stop, leaking role markers into its output, or ignoring the system prompt at inference. People usually respond by collecting more data, but more data doesn't fix it. Matching the template does.
How much data
There is no universal minimum, and any specific number you see quoted is a rule of thumb and not a finding. The honest guidance from Unsloth's datasets guide is that quality and quantity together determine the result.
The useful heuristic is directional. Teaching style or format needs far fewer examples than teaching a domain, because you are nudging a behaviour the model can already produce and not installing something new. Start small, evaluate, and add data only where evaluation shows a gap.
Three anchors that are published results and not folklore:
- One thousand examples. LIMA (Zhou et al., 2023) fine-tuned a 65B base on exactly 1,000 hand-curated prompt-response pairs, with no RLHF at all, and reported performance competitive with models trained on vastly more. It is the strongest single piece of evidence that curation beats volume.
- A few hundred. Enough to move tone or lock an output format, where you are steering a behaviour the base can already produce on a good day.
- Tens of thousands upward. What a genuine domain shift takes — and the point at which the question from section 01 comes back, because retrieval may well be doing that job better.
"Quality over quantity" is a slogan until you say what quality means. Operationally it means six specific things, and you can check every one of them by reading fifty examples.
- Every response is one you would ship. The model reproduces the average of what you give it, so examples that are merely acceptable or nearly right pull that average down.
- One convention throughout. If half the responses open with a greeting and half do not, you are training a coin flip.
- Lengths match the target. Response length is one of the easiest things to learn and one of the easiest to teach wrong.
- The prompts look like production. Clean, well-formed instructions teach the model to expect clean, well-formed instructions.
- Failure cases are present. Include the examples where the right answer is "I don't know" or a refusal, or the model will learn that an answer always exists.
- No near-duplicates. They inflate the apparent dataset size, get memorised first, and leak across a random train/eval split.
Generating training data with a larger model is now the default, and it works. The standard shape: prompt a frontier model for candidate responses, filter them against an explicit written rubric, then read a random sample by hand before anything trains.
The filter is the part people skip, and skipping it produces a dataset with one voice, one sentence rhythm and one way of being wrong. You then fine-tune a smaller model to imitate exactly that. The symptom is a model that is fluent, plausible and strangely uniform — and no aggregate metric will tell you, because the eval set was generated the same way.
The unglamorous checklist
- Deduplicate. Near-duplicates inflate your apparent dataset size and quietly become memorised.
- Hold out a real eval set before training, and split by source or by time where you can, so near-duplicates cannot straddle the split.
- Read a random sample by hand. Fifty examples read properly will find problems no aggregate statistic surfaces.
- Be consistent. If half your responses are terse and half are chatty, you are training the model to be unpredictable.
- Mask the prompt. Train on the response tokens only, so the model learns to answer your questions and does not learn to reproduce them.
06Doing it with Unsloth
Unsloth is a drop-in optimisation layer for fine-tuning on a single GPU. It is not a new algorithm. It rewrites the model's PyTorch modules as Triton kernels and derives the backward pass manually, so you run the same LoRA or QLoRA training with less memory and more speed.
The property that matters most is easy to skip past: because it makes no approximations, the result is mathematically the same training you would have run anyway. You are buying throughput and giving up no accuracy.
On measured gains, be careful with numbers you see quoted. Hugging Face's benchmark write-up reported roughly 2x speedups and large memory reductions, but it was measured against Transformers 4.36 on specific small models. Read it as directional evidence that the gains are real and not as a current spec sheet.
The pipeline, end to end
Load a quantized base and attach adapters. load_in_4bit=True is the setting that makes this QLoRA.
from unsloth import FastLanguageModel
max_seq_length = 2048
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = "unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit",
max_seq_length = max_seq_length,
load_in_4bit = True, # QLoRA: 4-bit frozen base
dtype = None, # autodetect bf16 / fp16
)
model = FastLanguageModel.get_peft_model(
model,
r = 16,
lora_alpha = 16,
lora_dropout = 0,
target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
use_gradient_checkpointing = "unsloth",
)
Apply the chat template that matches your base model. Getting this wrong is the trap from the previous section.
from unsloth.chat_templates import get_chat_template
tokenizer = get_chat_template(tokenizer, chat_template = "llama-3.1")
def formatting_func(examples):
convos = examples["conversations"]
texts = [tokenizer.apply_chat_template(c, tokenize = False,
add_generation_prompt = False)
for c in convos]
return {"text": texts}
dataset = dataset.map(formatting_func, batched = True)
Train. Note train_on_responses_only, which masks the instruction tokens so loss is computed on the assistant's turn alone.
from trl import SFTTrainer, SFTConfig
from unsloth.chat_templates import train_on_responses_only
trainer = SFTTrainer(
model = model,
processing_class = tokenizer, # older examples call this `tokenizer=`
train_dataset = dataset,
args = SFTConfig(
dataset_text_field = "text",
max_length = max_seq_length,
per_device_train_batch_size = 2,
gradient_accumulation_steps = 4, # effective batch size 8
num_train_epochs = 2,
learning_rate = 2e-4,
warmup_steps = 10,
optim = "adamw_8bit",
weight_decay = 0.01,
lr_scheduler_type = "linear",
logging_steps = 1,
output_dir = "outputs",
seed = 3407,
),
)
trainer = train_on_responses_only(
trainer,
instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n",
response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n",
)
trainer.train()
Save. Which export you want depends entirely on where the model will run.
# Just the adapter. Small, and swappable against the same base.
model.save_pretrained("lora_model")
tokenizer.save_pretrained("lora_model")
# Adapter folded into the base, for a normal serving stack.
model.save_pretrained_merged("merged_model", tokenizer,
save_method = "merged_16bit")
# GGUF, for llama.cpp and Ollama. See page 30.
model.save_pretrained_gguf("gguf_model", tokenizer,
quantization_method = "q4_k_m")
That last export is the handoff to 30. A q4_k_m GGUF is the format most people run a fine-tune in locally.
07Did it work?
Training loss going down tells you the model memorised your data. It says nothing about whether the model got better.
The only measurement that counts is the same held-out eval set, scored before and after, ideally with outputs you read yourself. Set that up before training starts. Building an eval after you have a result is how you end up grading your own homework.
Three failures, and how each announces itself
- Overfitting. Training loss keeps dropping while held-out quality stalls or declines. Unsloth's guide flags training loss below roughly 0.2 as a sign of memorisation. Fix by cutting epochs, lowering the learning rate, or adding data.
- Catastrophic forgetting. Your task improves and unrelated general ability quietly rots. It is invisible unless you deliberately test capabilities you never trained on. Keep a small general-purpose probe set and run it every time. Section 08 is about how much of it you can prevent.
- Underfitting. Barely distinguishable from the base model. Raise the learning rate, raise rank, or train longer, in that order.
Build one habit: keep base and fine-tuned outputs side by side for the same prompts, and read them. Aggregate scores hide the regressions that a human notices in ten seconds.
08Catastrophic forgetting, and what helps
Your support model now answers in house style, and it has quietly stopped being able to write a Python function. Nothing in your evaluation could have told you, because your evaluation is made of support tickets.
The mechanism is not mysterious. General ability and your new behaviour ride on the same weights. Gradient descent on a narrow objective moves those weights toward whatever serves that objective, and it has no term that cares what else they were doing.
It is invisible by construction, which is the actual problem. Every metric you built points at the task you trained on, and on that task the numbers improve. Here is the shape of it, on a model small enough to train while you read this.
Try it Forget task A by learning task B — then try four ways of not doing that
import numpy as np
rng = np.random.default_rng(0)
D, HID = 16, 4 # a deliberately tight shared trunk
def task(w, n):
X = rng.normal(size=(n, D))
return X, (X @ w > 0).astype(float)
wA, wB = rng.normal(size=D), rng.normal(size=D)
XA, yA = task(wA, 3000); XB, yB = task(wB, 3000)
TA, tA = task(wA, 2000); TB, tB = task(wB, 2000)
P0 = [rng.normal(0, .3, (D, HID)), np.zeros(HID),
rng.normal(0, .3, (HID, 2)), np.zeros(2)]
def fwd(P, X):
h = np.maximum(X @ P[0] + P[1], 0) # shared trunk
return h, h @ P[2] + P[3] # two task heads
def train(P, X, y, head, steps=600, lr=.5, trunk=True):
P = [q.copy() for q in P]
i = np.arange(len(X))
for _ in range(steps):
h, z = fwd(P, X)
gz = np.zeros_like(z)
gz[i, head] = (1 / (1 + np.exp(-z[i, head])) - y) / len(X)
gh = (gz @ P[2].T) * (h > 0)
P[2] -= lr * h.T @ gz; P[3] -= lr * gz.sum(0)
if trunk:
P[0] -= lr * X.T @ gh; P[1] -= lr * gh.sum(0)
return P
def show(tag, P):
a = ((fwd(P, TA)[1][:, 0] > 0) == (tA > 0)).mean()
b = ((fwd(P, TB)[1][:, 1] > 0) == (tB > 0)).mean()
print("%-27s A=%.3f B=%.3f" % (tag, a, b))
old, new = np.zeros(3000, int), np.ones(3000, int)
P = train(P0, XA, yA, old)
show("after task A", P)
show("then task B, nothing else", train(P, XB, yB, new))
show(" same, 1/10 learning rate", train(P, XB, yB, new, lr=.05))
show(" same, trunk frozen", train(P, XB, yB, new, trunk=False))
k = 150 # 5% replay
Xm = np.vstack([XB, XA[:k]]); ym = np.r_[yB, yA[:k]]
hm = np.r_[new, np.zeros(k, int)]
show(" same, 5% task-A replay", train(P, Xm, ym, hm))
0.996 to 0.759 without the model ever seeing a task-A example — the shared trunk moved. Freezing that trunk protects A perfectly and fails to learn B at all (0.680), which is what over-constraining looks like. A tenth of the learning rate is worse at both than the last row. 5% replay recovers A to 0.941 and keeps B at 0.994.Ordered by what they cost you, these are the levers that work.
- Train less. Use fewer epochs, a lower learning rate, and early stopping judged on a general probe and not on training loss. This is the most common real fix and it is free, because most forgetting comes from runs that went on too long.
- Constrain what can move. LoRA is already this: the base is untouched and the update is confined to rank $r$ per matrix. Lowering rank and targeting fewer modules moves less of the model, at the cost of how much your task can change.
- Replay. Mix a few percent of general instruction data into the fine-tuning set. The demo above shows it working: 5% recovered most of what was lost and cost nothing at the new task.
- Do not merge the adapter. Keep it as a separate file and the base model is still there, unchanged, one flag away. You can serve both, compare them on the same prompts, and roll back in seconds.
- Measure it deliberately. Keep a small fixed probe set of capabilities you never trained on — code, arithmetic, a different language, instruction following — and run it at every checkpoint. Forgetting you cannot see is forgetting you will ship.
Every row of the demo trades the same two quantities. The frozen-trunk row preserves the old behaviour perfectly and learns nothing; training hard learns the new task perfectly and damages the old one. Replay and restraint move you along the curve, they do not escape it.
Which means the honest question is not "how do I avoid forgetting" but "how much general ability is this behaviour worth, and have I measured what I actually gave up?"
09Where you meet this in the wild
The pattern across all of these: fine-tuning wins where the requirement is a consistent behaviour that would otherwise cost a long prompt on every single call.
- You need consistent behaviour that prompting produces only sometimes.
- The instruction has become a long preamble on every call, and you are paying for it repeatedly.
- You want a smaller model to do one job as well as a larger one does it generally.
- The target is style, format or idiom, which is exactly what weight updates encode well.
- You have, or can generate, consistent examples of the output you want.
- The problem is missing or changing facts, which is a retrieval problem, and retrieval updates in seconds.
- You have not exhausted prompting. Most reported fine-tuning wins were available from a better prompt.
- You need citations or provenance. Weights cannot tell you where an answer came from.
- Your data is small, inconsistent or unread. You will train the model on your inconsistency.
- The base model is not good enough. Fine-tuning shapes competence; it does not manufacture it.
10Interview questions
BeginnerWhen would you fine-tune instead of using RAG?
Fine-tune when the gap is behavioural: style, tone, output format, or a domain's idiom that prompting only approximates. Use RAG when the gap is knowledge, especially knowledge that changes or needs to be cited. The operational argument is decisive. A fact in a document is updated by editing the document; a fact in the weights requires another training run and still gives you no provenance. The two compose well, and a common production shape is a fine-tuned model that formats and reasons over retrieved context.
BeginnerWhat is the practical difference between LoRA and QLoRA?
LoRA freezes the base model and trains small low-rank adapter matrices, so optimizer state only exists for the adapters. QLoRA does the same thing but quantizes the frozen base to 4-bit first, which cuts the memory needed to merely hold the base. The distinction matters when the full-precision base does not fit in your VRAM at all. QLoRA is what makes fine-tuning a mid-sized model feasible on a single consumer GPU, at the cost of a small quality hit from quantizing the base.
IntermediateYour fine-tuned model rambles and leaks role markers like "assistant" into its output. What went wrong?
Almost certainly a chat template mismatch. Every instruct model is post-trained with a specific set of special tokens delimiting system, user and assistant turns. If you format training data with a different template, the model is simultaneously learning a new formatting convention and your task, and it does neither cleanly. The symptoms are exactly this: failing to stop in the right place, and emitting role markers as ordinary text. The fix is to apply the base model's own template, not to collect more data.
IntermediateWhy mask the instruction tokens during SFT?
If loss is computed over the whole sequence, the model spends capacity learning to predict your prompts as well as the responses. You do not want a model that is good at generating user questions. Masking the instruction portion, which Unsloth exposes as train_on_responses_only, restricts the loss to assistant tokens so all the gradient signal goes toward producing better answers. It matters most when prompts are long relative to responses, since otherwise most of the loss is being spent on text you will never need generated.
IntermediateTraining loss dropped steadily. Is the fine-tune working?
Unknown, and the question is a trap. Falling training loss shows the model is fitting the training set, which is also exactly what memorisation looks like. The measurement that matters is held-out performance, scored before and after on a set you split off before training. Unsloth's guide treats training loss below roughly 0.2 as a memorisation warning rather than a success signal. You should also probe capabilities you never trained on, because catastrophic forgetting is invisible to any metric computed only on your task.
IntermediateWhy does a full 7B fine-tune need over 100GB when QLoRA fits on a 24GB card?
Count bytes per parameter under mixed-precision AdamW. Two for the bf16 weights, two for the gradients, four for the fp32 master copy the optimizer updates — bf16 has too few mantissa bits to absorb small updates — and four each for Adam's first and second moments. Sixteen bytes per parameter, so 6.74 billion parameters is 107.8 GB before a single activation. The key structural fact is that only the first of those four lines is indexed by total parameters; the other twelve bytes are per trainable parameter. Freezing the base and training a rank-16 adapter across all seven projections leaves about 40M trainable parameters, roughly 0.6% of the model, so gradients and optimizer state collapse from 94 GB to about half a gigabyte. That still leaves 13.5 GB of frozen weights, which is what the 4-bit quantization in QLoRA attacks, taking it to about 3.4 GB. Total states under 4 GB, plus roughly 1 GB of activations at sequence 2048 with gradient checkpointing.
IntermediateWhat do rank and alpha control in LoRA, and which modules should you target?
The update is ΔW = (α/r)·BA, added to a frozen W. Rank r is the capacity knob: it is literally the rank of ΔW, so it sets how many independent directions the update can move in, and it determines the parameter count, r(d_in + d_out) per matrix. Alpha is not capacity — it is a plain scalar multiplier, so doubling alpha is arithmetically identical to doubling B, and it behaves like a learning-rate multiplier on the adapter. The convention of setting α = r or 2r keeps the effective scale α/r constant as you vary rank, so you do not re-tune the learning rate every time. On modules: the LoRA paper's ablation, at a fixed parameter budget, found query and value projections the best place to spend it. The QLoRA paper found that matching full fine-tuning required all linear layers. Those answer different questions, and since all seven projections at r=16 cost well under a gigabyte of optimizer state, the budget is usually not the binding constraint — so target everything and keep the rank modest.
DeepHow would you detect catastrophic forgetting, and what actually mitigates it?
You cannot detect it with your task metrics, because by construction they only measure the thing you trained on, and that improves. Detection requires a deliberate probe: a small fixed set of capabilities you never trained on — code, arithmetic, a second language, plain instruction following — scored at every checkpoint against the base model's own scores on the same set. Mitigations, cheapest first: train less, since most forgetting comes from too many epochs or too high a learning rate, with early stopping judged on the probe rather than on training loss. Then constrain what can move — LoRA at a modest rank is exactly that, since the base is frozen and the update is rank-limited per matrix. Then replay: mixing a few percent of general instruction data into the fine-tuning set recovers most of the loss for almost no cost to the new task. And keep the adapter unmerged, so the unchanged base is one flag away and you can roll back or A/B in seconds. None of these is free. Freezing enough to protect the old behaviour perfectly also prevents learning the new one, so the real question is how much general ability the behaviour is worth, and whether you measured what you gave up.
DeepYou have one 24GB GPU and need a customer-support model in a specialised domain. Walk through your plan.
Start by not fine-tuning. Establish a prompted baseline with retrieval over the support corpus, and measure it, because that may be the whole answer and it is the cheapest thing to maintain. Assume it leaves a gap in tone and format consistency rather than in knowledge, since knowledge is what retrieval already handles.
Then: pick a small instruct base that fits in 24GB when 4-bit quantized, so QLoRA rather than LoRA. Build a few thousand curated examples from real resolved tickets, deduplicated, with a held-out split made by time so near-duplicates cannot leak across it. Apply the base model's own chat template and mask instruction tokens. Start at rank 16, alpha 16, learning rate 2e-4, two epochs, targeting all seven linear modules.
Evaluate against the prompted baseline on held-out tickets, reading outputs rather than only scoring them, and separately probe general ability to catch forgetting. Keep retrieval in the final system regardless: the fine-tune supplies the voice and the format, retrieval supplies the facts and the citations.
11Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.