←Home KnowML
Hands-onChapter 29

Fine-tuning LLMs in Practice

Page 10 covers what SFT, LoRA and DPO are. This page is the part that gets you a working adapter: the memory arithmetic that decides the method, what rank and alpha do, what data to feed it, what it quietly costs you, and how to tell whether any of it helped.

36 min read Assumes: LoRA and SFT theory (10)
Start reading
TL;DR

Most fine-tuning projects should not have been fine-tuning projects. Work up the ladder first: prompt, few-shot, then retrieval. Fine-tune when you need behaviour the base model will not reliably produce, such as a house style, a rigid output format, or a domain's idiom. Once you commit, your VRAM chooses the method: QLoRA on a quantized base is what fits on one consumer GPU.

The arithmetic behind that: mixed-precision AdamW costs 16 bytes per trainable parameter, so a full 7B fine-tune is 108 GB while a rank-16 adapter on a 4-bit base is under 4. Three of those four bytes-per-parameter lines vanish when you freeze the base, which is the entire mechanism.

Your dataset decides the outcome far more than any hyperparameter. Match the base model's chat template exactly, hold out a real eval set before you start, and remember that falling training loss is evidence of memorisation and says nothing yet about usefulness. One sentence for an interview: fine-tuning reliably changes how a model behaves and unreliably changes what it knows.

01The decision before the decision

Fine-tuning is rarely the first correct move. It is the most expensive rung on a ladder, and most teams reach for it two rungs too early.

Climb in this order. Stop at the first rung that solves your problem.

  • Prompting. Free, instant, and revisable in seconds. Exhaust it properly before concluding it failed.
  • Few-shot examples. Still just prompting. Costs context window, but often fixes format problems outright.
  • Retrieval (RAG). The right answer whenever the issue is knowledge: updating a fact means updating a document, with no retraining. See 11.
  • Fine-tuning. Changes the weights. Slowest loop, hardest to undo, and the only option that reshapes default behaviour.

What fine-tuning earns its cost on:

  • Style and tone. Making every response sound like your organisation wrote it.
  • Format adherence. Producing valid JSON, or a fixed schema, without a paragraph of instructions each time.
  • Domain idiom. Absorbing how a specialist field phrases things, which prompting mimics only shallowly.
  • Cost and latency. Teaching a small model one narrow job well, so you stop paying for a large one.

What it does not fix. It will not make a weak base model competent, and no amount of adapter training rescues a poor architecture choice. It is also a bad mechanism for facts that change.

PromptingFew-shotRetrievalFine-tuning
Time to try a changeSecondsSecondsMinutes — reindexHours to days — retrain, then re-evaluate
Cost of fixing one wrong factEdit a sentenceEdit an exampleEdit a documentRebuild the dataset and train again
Can point at its sourceNoNoYesNo
Per-request token costGrows with the instructionGrows fastestGrows with retrieved contextNear zero — the behaviour is in the weights
What it fixesMost things, if you try properlyOutput format, usuallyKnowledge, freshness, provenanceDefault behaviour: style, format, idiom

Read the bottom row against the rightmost column. Fine-tuning's one real advantage over everything to its left is that the behaviour becomes the default and stops costing tokens on every call. Every other row in that table is an argument for trying a cheaper rung first.

Three questions that settle it before you write any code

Would a perfect prompt fix this? If the model produces the right output when you ask carefully and fails when you ask casually, the problem is consistency — which is a genuine fine-tuning case. If it cannot produce the output at all, fine-tuning will not teach it to.

Will the right answer be different next quarter? Then it belongs in a document. Weights give you no way to find what you baked in, no way to correct one fact, and no citation when someone asks where an answer came from.

Could you write two hundred examples you would be happy to ship? If not, you do not have a dataset yet. The run will faithfully learn whatever inconsistency you have, and you will read the result as a modelling failure.

Where the vendor guidance and the practitioner consensus disagree

Unsloth's own fine-tuning guide argues that fine-tuning can replicate everything RAG does, since it changes the weights and RAG cannot. The argument is technically defensible, because weights do store knowledge.

The practitioner objection is operational. A fact stored in weights cannot be updated without another training run, carries no citation, and gives you no way to check where an answer came from, while a fact stored in a document is edited in seconds and cites itself.

So the useful rule is not "fine-tuning cannot learn facts." It is: put the parts that change in retrieval, and the parts that never change in the weights.

02Choosing the method: your VRAM decides

Three options, and in practice the hardware sitting in front of you usually makes the choice. Page 10 covers why LoRA works mathematically. Here is when to use which.

Full fine-tuneLoRAQLoRA
What updatesEvery weightSmall low-rank adapters, base frozenSame adapters, on a 4-bit quantized base
Relative memoryHighest: weights, gradients and optimizer state for every parameterMuch lower: optimizer state only for adapter weightsLowest: adapters plus a quantized frozen base
Reach for it whenYou have multi-GPU budget and are changing the model deeplyThe full-precision base fits comfortably in VRAMThe base only fits when quantized, which is the common case on one GPU
Main costExpensive, and you get one model per taskSlight quality gap versus full fine-tuning on some tasksQuantizing the base adds its own small quality cost

Adapters have a property worth noticing. Because the base stays frozen, one base model can serve many tasks by swapping small adapter files. Full fine-tuning gives you a whole separate model per task.

Why LoRA trains under 1% of the parameters — the shape of the update
W 4096 × 4096 16.8M params FROZEN ❄ no gradients no optimizer state + B 4096 × 16 × A 16 × 4096 TRAINABLE 131k params · r = 16 0.78% of W = W + BA same 4096 × 4096 shape at inference EFFECTIVE gradients flow only here nothing to store
A rank-16 adapter on one 4096×4096 layer replaces 16.8M trainable parameters with 131k — under 1%. The saving compounds: optimizer state for Adam is roughly twice the trainable parameter count, so it shrinks by the same factor, and that is what moves fine-tuning from a multi-GPU job to a single card. At inference BA can be folded into W, so the served model has exactly the original shape and no extra latency.

The fourth family: prompt and prefix tuning

LoRA is not the only way to train a small number of parameters, and the alternatives are worth studying because they fail in an instructive direction.

  • Prompt tuning (Lester et al., 2021) freezes the entire network and learns a short sequence of continuous vectors prepended to the input embeddings. Nothing inside the model moves. Their finding: it only caught up with full model tuning at T5-XXL, 11B parameters.
  • Prefix tuning (Li & Liang, 2021) learns vectors at every layer and not only at the input, which is why it holds up at smaller scales. The paper reports comparable performance while training about 0.1% of the parameters.

Both spend context window at inference, and neither can be folded back into the weights — so they add latency that LoRA does not. They earn their place when the base must stay literally untouched, or when you cannot change the serving stack at all.

Starting hyperparameters

These are starting points, and you should expect to move them. They come from Unsloth's LoRA hyperparameters guide.

  • Rank r. Start at 16 or 32. Common values run 8 to 128. Higher rank adds capacity, and also adds overfitting risk.
  • lora_alpha. Set it equal to r, or to 2r to learn more aggressively.
  • lora_dropout. Zero by default. It does little on short runs. Raise it only if you see overfitting.
  • target_modules. Target all seven linear layers. Omitting some saves very little memory and measurably costs quality.
  • Learning rate. 2e-4 is the usual starting point for LoRA and QLoRA.
  • Epochs. One to three. Beyond three, the guide reports diminishing returns and rising overfitting risk.

03The memory arithmetic, line by line

One 24 GB card, one 7B model. Full fine-tuning needs about 108 GB before a single activation exists; while QLoRA needs under 4. Both numbers fall out of the same five-line ledger, and it is worth walking through once instead of memorising the conclusion.

Worked example Where the 108 GB in a full 7B fine-tune actually goes

Llama-2-7B has 6,738,415,616 parameters. Assume the standard recipe: bf16 forward and backward, AdamW, no sharding. Every line is bytes per parameter.

  1. $$\text{weights} = 2 \times 6.74\times10^{9} = 13.5\ \text{GB}$$
    Two bytes per parameter, because the matmuls run in bf16 to reach the tensor cores. This is the only line most people count, and it is an eighth of the bill.
  2. $$\text{gradients} = 2 \times 6.74\times10^{9} = 13.5\ \text{GB}$$
    One gradient per trainable parameter, in the dtype the backward pass produced it in. This line is what LoRA deletes: a frozen weight has no gradient.
  3. $$\text{fp32 master copy} = 4 \times 6.74\times10^{9} = 27.0\ \text{GB}$$
    bf16 carries seven explicit mantissa bits, so its spacing near a value is about one part in 256. Add a relative update of $10^{-4}$ to a bf16 weight and it rounds back to where it was. So the optimizer keeps an fp32 copy, updates that, and casts down for the next forward pass.
  4. $$\text{Adam } m + v = (4 + 4) \times 6.74\times10^{9} = 53.9\ \text{GB}$$
    A running first and second moment per parameter, both fp32. Four times the size of the weights they are optimising, and the single largest line in the ledger.
  5. $$(2 + 2 + 4 + 8) \times 6{,}738{,}415{,}616 = 107.8\ \text{GB}$$
    Sixteen bytes per parameter, the same accounting page 28 uses for distributed training. Adding the four rounded lines above gives 107.9 rather than 107.8 — each one rounds up, and four of those add a tenth. The exact figure is 107.81 GB. Activations are on top of this and depend on batch and sequence length.
Lines 2, 3 and 4 exist only for parameters you are training. That is 94.3 GB of the 107.8, and it is the entire reason parameter-efficient fine-tuning works. Freeze the base and those three lines shrink by whatever fraction of the model you left trainable — not by a little, by the same factor. Line 1 is the only one that stays, and quantization is what shrinks that.

So LoRA's memory saving doesn't come from the adapter being small. It comes from three of the four lines being indexed by trainable parameters and not by total parameters, and a rank-16 adapter makes that number about 0.6% of the model.

How small exactly? Each targeted matrix of shape $d_{\text{out}} \times d_{\text{in}}$ gets an $A$ of shape $r \times d_{\text{in}}$ and a $B$ of shape $d_{\text{out}} \times r$, so it costs $r(d_{\text{in}} + d_{\text{out}})$ parameters instead of $d_{\text{in}} d_{\text{out}}$. Summed over Llama-2-7B's seven projections and 32 layers, at $r = 16$, that is 40.0M.

Try it The whole ledger, for five configurations of the same model
# Llama-2-7B geometry, straight from its config.json.
L, H, FF = 32, 4096, 11008
BASE = 6_738_415_616                  # its real parameter count

SHAPE = {"q": (H, H), "k": (H, H), "v": (H, H), "o": (H, H),
         "gate": (H, FF), "up": (H, FF), "down": (FF, H)}

def lora_params(r, targets):
    """A is r x d_in, B is d_out x r, for every targeted matrix."""
    return L * sum(r * (SHAPE[t][0] + SHAPE[t][1]) for t in targets)

ALL7 = list(SHAPE)
QV = ["q", "v"]

print("%-22s %9s %8s %7s %7s %8s"
      % ("", "trainable", "weights", "grads", "optim", "total"))
for name, train, bytes_per_w in [
        ("full fine-tune bf16", BASE, 2),
        ("LoRA r=16, q+v", lora_params(16, QV), 2),
        ("LoRA r=16, all 7", lora_params(16, ALL7), 2),
        ("LoRA r=64, all 7", lora_params(64, ALL7), 2),
        ("QLoRA r=16, all 7", lora_params(16, ALL7), 0.5)]:
    w = BASE * bytes_per_w / 1e9
    g = train * 2 / 1e9               # bf16 gradient per trainable
    o = train * 12 / 1e9              # fp32 master + Adam m + Adam v
    print("%-22s %8.1fM %7.1fG %6.2fG %6.2fG %7.1fG"
          % (name, train / 1e6, w, g, o, w + g + o))
print("\nactivations are on top of all of these and depend on"
      "\nbatch, sequence length and gradient checkpointing.")
Compare the grads and optim columns down the rows. Full fine-tuning spends 13.48 and 80.86 GB; every LoRA row spends under half a gigabyte on both, because those columns are priced by trainable parameters. Now compare weights across the last two rows: identical adapters, 13.5 against 3.4 GB, purely from quantizing the frozen base. Two independent levers, and QLoRA is both of them at once.

One more thing that table settles: rank is cheap. Going from r=16 to r=64 costs 1.7 GB, which you can afford on a 24 GB card, so the reason to keep rank low is overfitting risk and not memory.

The same 7B model, three ways — bars are linear in gigabytes
24 GB — one RTX 4090 80 GB — one H100 full fine-tune 107.8 GB LoRA r=16 14.0 GB — bf16 base, 40M trainable QLoRA r=16 3.9 GB — 4-bit base, same 40M trainable 0 40 80 120 GB weights + gradients + optimizer state only · activations add roughly 1.2 GB at batch 1, sequence 2048, with gradient checkpointing
Bar lengths are the totals the Try it above prints, drawn against the same axis. The gap between the top bar and the 24 GB line is the reason the choice in section 02 is made by hardware: full fine-tuning a 7B model overshoots a 24 GB card four and a half times over, and needs at least two 80 GB cards. No amount of batch-size tuning changes that, because not one of those 107.8 GB depends on batch size.

Does it fit in 24 GB?

States are 3.9 GB for QLoRA — 3.4 for the 4-bit base plus roughly 0.1 GB of quantization constants, and 0.56 for the adapters, their gradients and their optimizer state. Then activations, which is the term people forget and the term that scales with your sequence length.

With gradient checkpointing on, you keep one hidden state per layer boundary: $32 \times 2048 \times 4096 \times 2 = 0.54$ GB at batch 1. Add the logits and their fp32 copy for the loss, about 0.39 GB, plus the recomputation peak inside one layer, and call it 1.2 GB.

Five gigabytes, on a 24 GB card. That headroom is what you spend on batch size and sequence length, both of which multiply the activation term and leave the other 3.9 GB untouched. LoRA on a bf16 base comes to about 15 GB by the same arithmetic, which also fits — with far less room to grow the sequence.

04What LoRA is doing

Page 10 states the equation. This section is about what each symbol in it controls, because r and lora_alpha are the two knobs you will turn and they do completely different things.

$$h = Wx + \frac{\alpha}{r}\,BAx, \qquad B \in \mathbb{R}^{d_{\text{out}} \times r},\ A \in \mathbb{R}^{r \times d_{\text{in}}}$$
  • W the pretrained weight, frozen: no gradient, no optimizer state, and after quantization not even full precision.
  • A initialised from a Gaussian. It projects the input down to $r$ dimensions.
  • B initialised to exactly zero, so $BA = 0$ at step 0 and training starts exactly at the pretrained model.
  • α/r a fixed scalar that is not learned. It scales the whole update.

The zero initialisation of $B$ is a small detail with a large consequence: the first forward pass of fine-tuning is bit-identical to the base model's. There is no warm-up period where the adapter is random noise degrading a working model.

Why a low rank is enough

The claim is about $\Delta W$, the change that adapting to one task requires. $W$ itself is plainly not low rank.

The evidence predates LoRA. Aghajanyan et al. (2020) measured the intrinsic dimension of fine-tuning directly: RoBERTa-Large reaches 90% of full fine-tuning performance on MRPC while optimising 200 parameters inside a random low-dimensional subspace, and 774 on QQP. That is a 355M-parameter model adapted by a few hundred numbers.

LoRA's bet is that if the useful update lives in a subspace that small, you can just parameterise a low-rank one directly and skip the random projection. You can watch the mechanism on a matrix small enough to take the SVD of.

Try it Train a full weight matrix on a narrow task, then look at the spectrum of what changed
import numpy as np
rng = np.random.default_rng(0)

d, k = 128, 3
W0 = rng.normal(0, 0.05, (d, d))         # the "pretrained" weight

# The new task only exercises a k-dimensional slice of the input
# and only moves a k-dimensional slice of the output.
U = np.linalg.qr(rng.normal(size=(d, k)))[0]
V = np.linalg.qr(rng.normal(size=(d, k)))[0]
W_task = W0 + U @ rng.normal(size=(k, k)) @ V.T

W = W0.copy()
for _ in range(600):                     # plain full-matrix SGD
    X = rng.normal(size=(64, d))
    Y = X @ W_task + rng.normal(0, 0.02, (64, d))   # noisy labels
    W -= 0.05 * X.T @ (X @ W - Y) / 64

s = np.linalg.svd(W - W0, compute_uv=False)
energy = np.cumsum(s ** 2) / np.sum(s ** 2)
print("delta W is a full %d x %d matrix (%d numbers)"
      % (d, d, d * d))
print("its top 8 singular values:", np.round(s[:8], 3))
print("energy captured by rank 1..5:", np.round(energy[:5], 4))
print("rank needed for 99% of it:",
      int(np.searchsorted(energy, 0.99)) + 1,
      " -- LoRA would store %d numbers at that rank"
      % (2 * d * (int(np.searchsorted(energy, 0.99)) + 1)))
Look at the cliff between the third and fourth singular value: 0.435, then 0.009. Nothing constrained this optimisation — all 16,384 entries of a dense matrix, noisy labels, ordinary gradient descent — and the update it found is still rank 3, with 768 numbers reproducing 99.98% of it. Be clear on what that shows: the task is rank 3 by construction, so this is the mechanism. The evidence is Aghajanyan's measurement above.

What r and alpha each control

They are routinely confused, including in tutorials. r is a capacity knob: it sets the rank of $\Delta W$ and therefore how many independent directions the update can move in. alpha is a multiplier and adds no capacity.

Try it Confirm which knob changes the rank and which one only changes the scale
import numpy as np

d = 256
W = np.random.default_rng(0).normal(0, 0.02, (d, d))

def delta_w(r, alpha, seed, zero_B=False):
    g = np.random.default_rng(seed)
    A = g.normal(0, 0.01, (r, d))                  # Gaussian init
    B = np.zeros((d, r)) if zero_B else g.normal(0, 0.01, (d, r))
    return (alpha / r) * (B @ A)                   # the effective dW

# 1. B starts at exactly zero, so a fresh adapter is a no-op.
print("fresh adapter leaves W untouched:",
      np.array_equal(W + delta_w(16, 32, 1, zero_B=True), W))

# 2. r is the capacity knob: it sets the RANK of the update.
for r in (1, 4, 16, 64):
    dw = delta_w(r, 16, 2)
    print("r=%-3d rank(dW)=%-3d  stored %6d numbers (%5.2f%% of W)"
          % (r, np.linalg.matrix_rank(dw), 2 * r * d,
             100 * 2 * r * d / d ** 2))

# 3. alpha is not a capacity knob. It is a plain multiplier.
a16, a32 = delta_w(16, 16, 3), delta_w(16, 32, 3)
print("alpha=32 is exactly 2x alpha=16:", np.allclose(a32, 2 * a16))
print("same rank either way:", np.linalg.matrix_rank(a16),
      np.linalg.matrix_rank(a32))
alpha=32 is exactly 2x alpha=16, at identical rank. Doubling alpha is arithmetically the same as doubling $B$, so it acts as a learning-rate multiplier on the adapter, not as extra capacity. Setting $\alpha = r$ or $2r$ holds the effective scale $\alpha/r$ fixed as rank changes, so the learning rate survives. And note r=64 storing half of a 256-wide $W$ — the same rank on a 4096-wide layer is 3.1%.

Which layers to target

Two published findings that look contradictory and are not.

  • The LoRA paper's ablation, at a fixed parameter budget, found that spending it on $W_q$ and $W_v$ beat spending it on any single projection — and that a rank as low as 1 sufficed in that setting.
  • The QLoRA paper found that matching full fine-tuning required applying LoRA to all linear layers in the transformer block, beyond the attention projections.

They answer different questions. If the budget is binding, attention's query and value projections are where it buys the most. But the ledger in section 03 shows the budget usually is not binding: all seven projections at $r=16$ cost 0.56 GB of gradients and optimizer state on a 24 GB card.

So the practical default is the second finding. Target everything, keep the rank modest, and treat the attention-only configuration as what you fall back to when memory is tight — which, after section 03, it rarely is.

The property that makes adapters an operational choice, not just a memory one

Because $\Delta W = \frac{\alpha}{r}BA$ is a matrix of exactly $W$'s shape, it can be added into $W$ once and thrown away. The served model has the original architecture, the original parameter count and the original latency. Prompt and prefix tuning cannot do this.

Or you keep it separate: a 40M-parameter adapter is roughly 80 MB in bf16, so one base model in memory can serve many tasks by swapping files. That also means you can turn a fine-tune off, which is the cheapest catastrophic-forgetting insurance available.

05Data

Hyperparameters move results by a little and the dataset moves them by a lot, so budget your time accordingly.

A supervised fine-tuning example is a conversation: an instruction, optionally a system message, and the response you wish the model had given. The model trains to predict the response tokens. The objective is the same next-token objective from 10, pointed at a much smaller and much more deliberate pile of text.

The chat template trap

Every instruct model was post-trained with a specific chat template: the exact special tokens and markers that separate system, user and assistant turns. Llama's template differs from ChatML, and ChatML differs from Gemma's.

If you train with a template the base model has never seen, you teach it a new formatting convention at the same time as your actual task, and both suffer. It is the most common way a first fine-tune quietly underperforms.

The failure that looks like a data problem

Symptoms of a template mismatch include the model rambling past where it should stop, leaking role markers into its output, or ignoring the system prompt at inference. People usually respond by collecting more data, but more data doesn't fix it. Matching the template does.

How much data

There is no universal minimum, and any specific number you see quoted is a rule of thumb and not a finding. The honest guidance from Unsloth's datasets guide is that quality and quantity together determine the result.

The useful heuristic is directional. Teaching style or format needs far fewer examples than teaching a domain, because you are nudging a behaviour the model can already produce and not installing something new. Start small, evaluate, and add data only where evaluation shows a gap.

Three anchors that are published results and not folklore:

  • One thousand examples. LIMA (Zhou et al., 2023) fine-tuned a 65B base on exactly 1,000 hand-curated prompt-response pairs, with no RLHF at all, and reported performance competitive with models trained on vastly more. It is the strongest single piece of evidence that curation beats volume.
  • A few hundred. Enough to move tone or lock an output format, where you are steering a behaviour the base can already produce on a good day.
  • Tens of thousands upward. What a genuine domain shift takes — and the point at which the question from section 01 comes back, because retrieval may well be doing that job better.

"Quality over quantity" is a slogan until you say what quality means. Operationally it means six specific things, and you can check every one of them by reading fifty examples.

  • Every response is one you would ship. The model reproduces the average of what you give it, so examples that are merely acceptable or nearly right pull that average down.
  • One convention throughout. If half the responses open with a greeting and half do not, you are training a coin flip.
  • Lengths match the target. Response length is one of the easiest things to learn and one of the easiest to teach wrong.
  • The prompts look like production. Clean, well-formed instructions teach the model to expect clean, well-formed instructions.
  • Failure cases are present. Include the examples where the right answer is "I don't know" or a refusal, or the model will learn that an answer always exists.
  • No near-duplicates. They inflate the apparent dataset size, get memorised first, and leak across a random train/eval split.
The synthetic data loop, and the way it goes wrong

Generating training data with a larger model is now the default, and it works. The standard shape: prompt a frontier model for candidate responses, filter them against an explicit written rubric, then read a random sample by hand before anything trains.

The filter is the part people skip, and skipping it produces a dataset with one voice, one sentence rhythm and one way of being wrong. You then fine-tune a smaller model to imitate exactly that. The symptom is a model that is fluent, plausible and strangely uniform — and no aggregate metric will tell you, because the eval set was generated the same way.

The unglamorous checklist

  • Deduplicate. Near-duplicates inflate your apparent dataset size and quietly become memorised.
  • Hold out a real eval set before training, and split by source or by time where you can, so near-duplicates cannot straddle the split.
  • Read a random sample by hand. Fifty examples read properly will find problems no aggregate statistic surfaces.
  • Be consistent. If half your responses are terse and half are chatty, you are training the model to be unpredictable.
  • Mask the prompt. Train on the response tokens only, so the model learns to answer your questions and does not learn to reproduce them.

06Doing it with Unsloth

Unsloth is a drop-in optimisation layer for fine-tuning on a single GPU. It is not a new algorithm. It rewrites the model's PyTorch modules as Triton kernels and derives the backward pass manually, so you run the same LoRA or QLoRA training with less memory and more speed.

The property that matters most is easy to skip past: because it makes no approximations, the result is mathematically the same training you would have run anyway. You are buying throughput and giving up no accuracy.

On measured gains, be careful with numbers you see quoted. Hugging Face's benchmark write-up reported roughly 2x speedups and large memory reductions, but it was measured against Transformers 4.36 on specific small models. Read it as directional evidence that the gains are real and not as a current spec sheet.

The pipeline, end to end

Load a quantized base and attach adapters. load_in_4bit=True is the setting that makes this QLoRA.

from unsloth import FastLanguageModel

max_seq_length = 2048

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name     = "unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit",
    max_seq_length = max_seq_length,
    load_in_4bit   = True,   # QLoRA: 4-bit frozen base
    dtype          = None,   # autodetect bf16 / fp16
)

model = FastLanguageModel.get_peft_model(
    model,
    r              = 16,
    lora_alpha     = 16,
    lora_dropout   = 0,
    target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
                      "gate_proj", "up_proj", "down_proj"],
    use_gradient_checkpointing = "unsloth",
)

Apply the chat template that matches your base model. Getting this wrong is the trap from the previous section.

from unsloth.chat_templates import get_chat_template

tokenizer = get_chat_template(tokenizer, chat_template = "llama-3.1")

def formatting_func(examples):
    convos = examples["conversations"]
    texts  = [tokenizer.apply_chat_template(c, tokenize = False,
                                            add_generation_prompt = False)
              for c in convos]
    return {"text": texts}

dataset = dataset.map(formatting_func, batched = True)

Train. Note train_on_responses_only, which masks the instruction tokens so loss is computed on the assistant's turn alone.

from trl import SFTTrainer, SFTConfig
from unsloth.chat_templates import train_on_responses_only

trainer = SFTTrainer(
    model            = model,
    processing_class = tokenizer,   # older examples call this `tokenizer=`
    train_dataset    = dataset,
    args = SFTConfig(
        dataset_text_field = "text",
        max_length         = max_seq_length,
        per_device_train_batch_size = 2,
        gradient_accumulation_steps = 4,   # effective batch size 8
        num_train_epochs   = 2,
        learning_rate      = 2e-4,
        warmup_steps       = 10,
        optim              = "adamw_8bit",
        weight_decay       = 0.01,
        lr_scheduler_type  = "linear",
        logging_steps      = 1,
        output_dir         = "outputs",
        seed               = 3407,
    ),
)

trainer = train_on_responses_only(
    trainer,
    instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n",
    response_part    = "<|start_header_id|>assistant<|end_header_id|>\n\n",
)

trainer.train()

Save. Which export you want depends entirely on where the model will run.

# Just the adapter. Small, and swappable against the same base.
model.save_pretrained("lora_model")
tokenizer.save_pretrained("lora_model")

# Adapter folded into the base, for a normal serving stack.
model.save_pretrained_merged("merged_model", tokenizer,
                             save_method = "merged_16bit")

# GGUF, for llama.cpp and Ollama. See page 30.
model.save_pretrained_gguf("gguf_model", tokenizer,
                           quantization_method = "q4_k_m")

That last export is the handoff to 30. A q4_k_m GGUF is the format most people run a fine-tune in locally.

07Did it work?

Training loss going down tells you the model memorised your data. It says nothing about whether the model got better.

The only measurement that counts is the same held-out eval set, scored before and after, ideally with outputs you read yourself. Set that up before training starts. Building an eval after you have a result is how you end up grading your own homework.

Three failures, and how each announces itself

  • Overfitting. Training loss keeps dropping while held-out quality stalls or declines. Unsloth's guide flags training loss below roughly 0.2 as a sign of memorisation. Fix by cutting epochs, lowering the learning rate, or adding data.
  • Catastrophic forgetting. Your task improves and unrelated general ability quietly rots. It is invisible unless you deliberately test capabilities you never trained on. Keep a small general-purpose probe set and run it every time. Section 08 is about how much of it you can prevent.
  • Underfitting. Barely distinguishable from the base model. Raise the learning rate, raise rank, or train longer, in that order.

Build one habit: keep base and fine-tuned outputs side by side for the same prompts, and read them. Aggregate scores hide the regressions that a human notices in ten seconds.

08Catastrophic forgetting, and what helps

Your support model now answers in house style, and it has quietly stopped being able to write a Python function. Nothing in your evaluation could have told you, because your evaluation is made of support tickets.

The mechanism is not mysterious. General ability and your new behaviour ride on the same weights. Gradient descent on a narrow objective moves those weights toward whatever serves that objective, and it has no term that cares what else they were doing.

It is invisible by construction, which is the actual problem. Every metric you built points at the task you trained on, and on that task the numbers improve. Here is the shape of it, on a model small enough to train while you read this.

Try it Forget task A by learning task B — then try four ways of not doing that
import numpy as np
rng = np.random.default_rng(0)
D, HID = 16, 4                    # a deliberately tight shared trunk

def task(w, n):
    X = rng.normal(size=(n, D))
    return X, (X @ w > 0).astype(float)

wA, wB = rng.normal(size=D), rng.normal(size=D)
XA, yA = task(wA, 3000);  XB, yB = task(wB, 3000)
TA, tA = task(wA, 2000);  TB, tB = task(wB, 2000)
P0 = [rng.normal(0, .3, (D, HID)), np.zeros(HID),
      rng.normal(0, .3, (HID, 2)), np.zeros(2)]

def fwd(P, X):
    h = np.maximum(X @ P[0] + P[1], 0)    # shared trunk
    return h, h @ P[2] + P[3]             # two task heads

def train(P, X, y, head, steps=600, lr=.5, trunk=True):
    P = [q.copy() for q in P]
    i = np.arange(len(X))
    for _ in range(steps):
        h, z = fwd(P, X)
        gz = np.zeros_like(z)
        gz[i, head] = (1 / (1 + np.exp(-z[i, head])) - y) / len(X)
        gh = (gz @ P[2].T) * (h > 0)
        P[2] -= lr * h.T @ gz;  P[3] -= lr * gz.sum(0)
        if trunk:
            P[0] -= lr * X.T @ gh;  P[1] -= lr * gh.sum(0)
    return P

def show(tag, P):
    a = ((fwd(P, TA)[1][:, 0] > 0) == (tA > 0)).mean()
    b = ((fwd(P, TB)[1][:, 1] > 0) == (tB > 0)).mean()
    print("%-27s A=%.3f  B=%.3f" % (tag, a, b))

old, new = np.zeros(3000, int), np.ones(3000, int)
P = train(P0, XA, yA, old)
show("after task A", P)
show("then task B, nothing else", train(P, XB, yB, new))
show("  same, 1/10 learning rate", train(P, XB, yB, new, lr=.05))
show("  same, trunk frozen", train(P, XB, yB, new, trunk=False))
k = 150                                          # 5% replay
Xm = np.vstack([XB, XA[:k]]);  ym = np.r_[yB, yA[:k]]
hm = np.r_[new, np.zeros(k, int)]
show("  same, 5% task-A replay", train(P, Xm, ym, hm))
Read it as one tradeoff curve with four points. Task B alone drags task A from 0.996 to 0.759 without the model ever seeing a task-A example — the shared trunk moved. Freezing that trunk protects A perfectly and fails to learn B at all (0.680), which is what over-constraining looks like. A tenth of the learning rate is worse at both than the last row. 5% replay recovers A to 0.941 and keeps B at 0.994.

Ordered by what they cost you, these are the levers that work.

  • Train less. Use fewer epochs, a lower learning rate, and early stopping judged on a general probe and not on training loss. This is the most common real fix and it is free, because most forgetting comes from runs that went on too long.
  • Constrain what can move. LoRA is already this: the base is untouched and the update is confined to rank $r$ per matrix. Lowering rank and targeting fewer modules moves less of the model, at the cost of how much your task can change.
  • Replay. Mix a few percent of general instruction data into the fine-tuning set. The demo above shows it working: 5% recovered most of what was lost and cost nothing at the new task.
  • Do not merge the adapter. Keep it as a separate file and the base model is still there, unchanged, one flag away. You can serve both, compare them on the same prompts, and roll back in seconds.
  • Measure it deliberately. Keep a small fixed probe set of capabilities you never trained on — code, arithmetic, a different language, instruction following — and run it at every checkpoint. Forgetting you cannot see is forgetting you will ship.
There is no free corner of that curve

Every row of the demo trades the same two quantities. The frozen-trunk row preserves the old behaviour perfectly and learns nothing; training hard learns the new task perfectly and damages the old one. Replay and restraint move you along the curve, they do not escape it.

Which means the honest question is not "how do I avoid forgetting" but "how much general ability is this behaviour worth, and have I measured what I actually gave up?"

09Where you meet this in the wild

The pattern across all of these: fine-tuning wins where the requirement is a consistent behaviour that would otherwise cost a long prompt on every single call.

Structured output
Getting valid JSON without a wall of instructions
A small model fine-tuned on a few thousand schema-conforming examples will emit that schema more reliably than a large model told to. You also stop paying for the instruction tokens on every request.
Cost reduction
Replacing a big general model with a small specialist
The common production pattern. Use a frontier model to generate and curate training data for one narrow task, fine-tune something much smaller on it, then serve the small one. Latency and cost both drop sharply.
Regulated writing
House style that is not negotiable
Legal, medical and financial writing have conventions that prompting approximates and fine-tuning internalises. The facts still belong in retrieval. The phrasing belongs in the weights.
Classification
A labelling job wearing a chat interface
Routing tickets, tagging content, triaging intent. Often a fine-tuned small model beats a prompted large one on both accuracy and cost. Sometimes a plain classifier beats both, which is worth checking first.
On-device
Fitting a capable model in a hard memory ceiling
Fine-tune a small model on one job, export to GGUF, run it locally with no API call and no data leaving the machine. Page 30 covers what fits.
Reach for fine-tuning when
  • You need consistent behaviour that prompting produces only sometimes.
  • The instruction has become a long preamble on every call, and you are paying for it repeatedly.
  • You want a smaller model to do one job as well as a larger one does it generally.
  • The target is style, format or idiom, which is exactly what weight updates encode well.
  • You have, or can generate, consistent examples of the output you want.
Look elsewhere when
  • The problem is missing or changing facts, which is a retrieval problem, and retrieval updates in seconds.
  • You have not exhausted prompting. Most reported fine-tuning wins were available from a better prompt.
  • You need citations or provenance. Weights cannot tell you where an answer came from.
  • Your data is small, inconsistent or unread. You will train the model on your inconsistency.
  • The base model is not good enough. Fine-tuning shapes competence; it does not manufacture it.

10Interview questions

BeginnerWhen would you fine-tune instead of using RAG?

Fine-tune when the gap is behavioural: style, tone, output format, or a domain's idiom that prompting only approximates. Use RAG when the gap is knowledge, especially knowledge that changes or needs to be cited. The operational argument is decisive. A fact in a document is updated by editing the document; a fact in the weights requires another training run and still gives you no provenance. The two compose well, and a common production shape is a fine-tuned model that formats and reasons over retrieved context.

BeginnerWhat is the practical difference between LoRA and QLoRA?

LoRA freezes the base model and trains small low-rank adapter matrices, so optimizer state only exists for the adapters. QLoRA does the same thing but quantizes the frozen base to 4-bit first, which cuts the memory needed to merely hold the base. The distinction matters when the full-precision base does not fit in your VRAM at all. QLoRA is what makes fine-tuning a mid-sized model feasible on a single consumer GPU, at the cost of a small quality hit from quantizing the base.

IntermediateYour fine-tuned model rambles and leaks role markers like "assistant" into its output. What went wrong?

Almost certainly a chat template mismatch. Every instruct model is post-trained with a specific set of special tokens delimiting system, user and assistant turns. If you format training data with a different template, the model is simultaneously learning a new formatting convention and your task, and it does neither cleanly. The symptoms are exactly this: failing to stop in the right place, and emitting role markers as ordinary text. The fix is to apply the base model's own template, not to collect more data.

IntermediateWhy mask the instruction tokens during SFT?

If loss is computed over the whole sequence, the model spends capacity learning to predict your prompts as well as the responses. You do not want a model that is good at generating user questions. Masking the instruction portion, which Unsloth exposes as train_on_responses_only, restricts the loss to assistant tokens so all the gradient signal goes toward producing better answers. It matters most when prompts are long relative to responses, since otherwise most of the loss is being spent on text you will never need generated.

IntermediateTraining loss dropped steadily. Is the fine-tune working?

Unknown, and the question is a trap. Falling training loss shows the model is fitting the training set, which is also exactly what memorisation looks like. The measurement that matters is held-out performance, scored before and after on a set you split off before training. Unsloth's guide treats training loss below roughly 0.2 as a memorisation warning rather than a success signal. You should also probe capabilities you never trained on, because catastrophic forgetting is invisible to any metric computed only on your task.

IntermediateWhy does a full 7B fine-tune need over 100GB when QLoRA fits on a 24GB card?

Count bytes per parameter under mixed-precision AdamW. Two for the bf16 weights, two for the gradients, four for the fp32 master copy the optimizer updates — bf16 has too few mantissa bits to absorb small updates — and four each for Adam's first and second moments. Sixteen bytes per parameter, so 6.74 billion parameters is 107.8 GB before a single activation. The key structural fact is that only the first of those four lines is indexed by total parameters; the other twelve bytes are per trainable parameter. Freezing the base and training a rank-16 adapter across all seven projections leaves about 40M trainable parameters, roughly 0.6% of the model, so gradients and optimizer state collapse from 94 GB to about half a gigabyte. That still leaves 13.5 GB of frozen weights, which is what the 4-bit quantization in QLoRA attacks, taking it to about 3.4 GB. Total states under 4 GB, plus roughly 1 GB of activations at sequence 2048 with gradient checkpointing.

IntermediateWhat do rank and alpha control in LoRA, and which modules should you target?

The update is ΔW = (α/r)·BA, added to a frozen W. Rank r is the capacity knob: it is literally the rank of ΔW, so it sets how many independent directions the update can move in, and it determines the parameter count, r(d_in + d_out) per matrix. Alpha is not capacity — it is a plain scalar multiplier, so doubling alpha is arithmetically identical to doubling B, and it behaves like a learning-rate multiplier on the adapter. The convention of setting α = r or 2r keeps the effective scale α/r constant as you vary rank, so you do not re-tune the learning rate every time. On modules: the LoRA paper's ablation, at a fixed parameter budget, found query and value projections the best place to spend it. The QLoRA paper found that matching full fine-tuning required all linear layers. Those answer different questions, and since all seven projections at r=16 cost well under a gigabyte of optimizer state, the budget is usually not the binding constraint — so target everything and keep the rank modest.

DeepHow would you detect catastrophic forgetting, and what actually mitigates it?

You cannot detect it with your task metrics, because by construction they only measure the thing you trained on, and that improves. Detection requires a deliberate probe: a small fixed set of capabilities you never trained on — code, arithmetic, a second language, plain instruction following — scored at every checkpoint against the base model's own scores on the same set. Mitigations, cheapest first: train less, since most forgetting comes from too many epochs or too high a learning rate, with early stopping judged on the probe rather than on training loss. Then constrain what can move — LoRA at a modest rank is exactly that, since the base is frozen and the update is rank-limited per matrix. Then replay: mixing a few percent of general instruction data into the fine-tuning set recovers most of the loss for almost no cost to the new task. And keep the adapter unmerged, so the unchanged base is one flag away and you can roll back or A/B in seconds. None of these is free. Freezing enough to protect the old behaviour perfectly also prevents learning the new one, so the real question is how much general ability the behaviour is worth, and whether you measured what you gave up.

DeepYou have one 24GB GPU and need a customer-support model in a specialised domain. Walk through your plan.

Start by not fine-tuning. Establish a prompted baseline with retrieval over the support corpus, and measure it, because that may be the whole answer and it is the cheapest thing to maintain. Assume it leaves a gap in tone and format consistency rather than in knowledge, since knowledge is what retrieval already handles.

Then: pick a small instruct base that fits in 24GB when 4-bit quantized, so QLoRA rather than LoRA. Build a few thousand curated examples from real resolved tickets, deduplicated, with a held-out split made by time so near-duplicates cannot leak across it. Apply the base model's own chat template and mask instruction tokens. Start at rank 16, alpha 16, learning rate 2e-4, two epochs, targeting all seven linear modules.

Evaluate against the prompted baseline on held-out tickets, reading outputs rather than only scoring them, and separately probe general ability to catch forgetting. Keep retrieval in the final system regardless: the fine-tune supplies the voice and the format, retrieval supplies the facts and the citations.

11Go deeper

📘
Docs
Fine-tuning LLMs Guide
Unsloth's own guide. Read it alongside the caveat in section 01 about its stance on fine-tuning versus RAG.
📘
Docs
LoRA Hyperparameters Guide
Unsloth — the source for the rank, alpha, learning-rate and epoch guidance in section 02.
📘
Docs
Datasets Guide
Unsloth — dataset formats and curation, the part of the job that decides your result.
📄
Paper
LoRA: Low-Rank Adaptation of Large Language Models
Hu et al., 2021 — arXiv:2106.09685
📄
Paper
QLoRA: Efficient Finetuning of Quantized LLMs
Dettmers et al., 2023 — arXiv:2305.14314
📄
Paper
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Aghajanyan et al., 2020 — arXiv:2012.13255. The measurement section 04 rests on, and it predates LoRA.
📄
Paper
LIMA: Less Is More for Alignment
Zhou et al., 2023 — arXiv:2305.11206. One thousand curated examples, no RLHF.
📄
Paper
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester et al., 2021 — arXiv:2104.08691. The soft-prompt alternative, and where it starts working.
📄
Article
Make LLM Fine-tuning 2x faster with Unsloth and TRL
Hugging Face — the benchmark behind the speed claims. Note it was measured against Transformers 4.36.
▶
Video
LLM Fine-Tuning Course – From Supervised FT to RLHF, LoRA, and Multimodal
freeCodeCamp.org — course length, and the broadest coverage of the three listed here.
▶
Video
Fine-Tuning Local LLMs with Unsloth & Ollama
NeuralNine — closest to this page's pipeline, and carries it through to running the result locally.
▶
Video
LLM Fine Tuning Crash Course: 1 Hour End-to-End Guide
AI Anytime — a single end-to-end run if you want one sitting rather than a course.

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 29 Fine-tuning LLMs in Practice

Free notes

Highlights on this page