RunLocalModel.com

Supervised Fine-Tuning Explained: Instructions, LoRA, QLoRA, and DoRA

By the RunLocalModel editorial team · Published September 30, 2026 · ~9 minute read

If you only read one paragraph Supervised fine-tuning is someone sitting with the bookworm and saying, over and over, “when a person says this, you say that.” A big public run uses on the order of hundreds of thousands of those pairs. Afterward the model answers instead of rambling. You can aim the examples at instructions, chat, one industry, long documents, or tool calls. Updating every weight costs a lot. LoRA, QLoRA, and DoRA train a small add-on instead. QLoRA is the one that fits a 7B model on a gaming-laptop GPU.
Quick answers
What actually changes?
Behavior. The model learns the shape of an answer: reply, then stop. The broad knowledge still comes from the reading phase.
What data does it need?
Examples of the behavior you want. A prompt and the reply, or a full chat, or a tool call followed by a final answer.
Full fine-tuning or LoRA?
Full fine-tuning if you can afford to touch every weight. LoRA, QLoRA, or DoRA if you want a small patch. QLoRA is the practical choice for a 7B model on one consumer GPU.
Do I need to fine-tune just to chat?
No. An instruct file already includes someone else’s fine-tune. Do it yourself when you have a format, a private domain, or a tool shape the public model keeps missing.

This is the first lesson in the teaching half. The map is in How LLMs Are Trained. The reading decided how much the model knows. This lesson decides whether it can do the work.

Same homework. Different words on the test.

The model is still guessing the next word. What changed is the worksheet, and which words count. A typical example is a user message plus a reply a person wrote. Training ignores the question, or at least does not grade it, and grades the reply. The weights move so that, given a request like that, those reply words become the likely ones.

A serious general set is on the order of hundreds of thousands of pairs. A narrow one — your support replies, your JSON shape, your lab’s note format — can be much smaller, because you are teaching a habit on top of knowledge the base model already has. FLAN showed that this kind of lesson lets a pretrained model follow tasks it was not separately trained for. InstructGPT put example answers in front of human preference training, and that order stuck.

After this step, a question is something to answer. Before it, a question is just more text to continue.

Five jobs, one kind of lesson

The optimizer does not care what the examples are about. You do. These are the piles people actually build.

Follow an instruction

One request, one reply. “Summarize this.” “Write a function that…” “Explain this error.” The target is a direct answer in a stable format. This is the pass that creates the everyday instruct file.

Hold a conversation

Full transcripts, with a user side and an assistant side. The model learns to use the history, to ask a clarifying question when the request is fuzzy, and to stay in the assistant seat across turns. The special markers that say who is talking are part of this data. If you fine-tune, you have to use the same markers the model expects when you chat with it later.

Stay in one field

The replies are written the way that field writes: clinical notes, contract language, a product’s support voice. This often comes after continued pretraining. The extra reading taught the model to read the field. These examples teach it to answer inside the field. A pile of domain text on its own still leaves you with a bookworm that is better at the jargon.

Use the document you just pasted

The right answer depends on a long passage in the prompt: a paper, a log, a policy. The examples punish replies that ignore the passage and recite a memorized fact instead. This is how a model gets reliable at “use what I just gave you,” which is what you want when you paste a document into the chat.

Call a tool

Sometimes the right next step is a structured call: a function name and arguments, usually as JSON. Then, after a tool result is dropped back into the conversation, a final reply that uses that result. The model learns when a tool is the right move and when it should just answer. The shape in your examples has to match the shape your app will actually send.

Full fine-tuning: every knob turns

Full fine-tuning updates the entire network. Nothing is frozen, so the ceiling is high. The bill is high for the same reason. The optimizer keeps extra notes for every number in the model, so memory is several times the size of the weight file. A 7B full fine-tune is already a large-GPU job. A 70B one is a cluster job. When you are done you store a whole new model, not a small patch.

The cheap way: train a small add-on

Parameter-efficient fine-tuning, or PEFT, freezes the pretrained weights and trains a small add-on. You keep one base on disk and swap add-ons per task. Three versions cover almost every run a normal person or a small lab will do.

Method What gets trained What it costs
Full fine-tuning Every weight Highest ceiling. Memory is several times the weight file, because the optimizer keeps extra state.
LoRA A small add-on beside frozen weights The patch is megabytes. The base still sits in memory at full precision, so a huge base still needs a big GPU.
QLoRA The same small add-on, with the base stored in 4-bit The practical path for a 7B model on about 8–12 GB. A 70B run is still a workstation job.
DoRA A small direction update, plus a learned size A bit more compute than LoRA. Closer to a full update. Still a patch, not a full retrain.

LoRA

LoRA writes the change to a weight as the product of two small grids of numbers. The rank might be 8, 16, or 64, against hidden sizes in the thousands, so the number of trained values collapses. The add-on usually sits on the attention layers, sometimes on the feed-forward layers too. When you are done you can fold it into the base weights, and chatting is not slower. Leave it separate and you can swap tasks without reloading the whole model.

QLoRA

QLoRA stores the frozen base in 4-bit and trains the LoRA add-on in 16-bit. That is the change that puts a 7B fine-tune on a gaming-laptop GPU of roughly 8–12 GB. The paper’s own large runs used a 48 GB card for models far above 7B. Quantizing the base saves memory. It does not make a 70B training run cheap. If you are picking a quant for chatting, not training, use the quantization guide. Q4 for chat and QLoRA for training are different jobs.

DoRA

DoRA splits each weight into a size and a direction. The small update mostly changes the direction, and the size stays its own learned vector. A full fine-tune tends to move size and direction differently. Plain LoRA ties them together. DoRA is still a patch — you are not rewriting the whole matrix — and it is the one to try when LoRA is not quite enough and a full fine-tune is out of reach.

When you should actually bother

Most people running models at home should download an instruct file and stop. That file already contains a large fine-tune, and usually a preference stage after it. Fine-tune when at least one of these is true:

The model will copy what is in the examples, including the mistakes. Tone, refusals, and “which of two fluent answers is better” are only partly solved here. That is the next lesson, preference alignment. For tasks with a checkable right answer, a later reward stage — reinforcement fine-tuning — can go past copying a written solution.

Related guides on this site