How LLMs Are Trained: Pretraining, Fine-Tuning, and Alignment
- What are the two steps?
- Reading, then teaching. The reading is pretraining. The teaching is post-training. “Instruct,” “chat,” and “reasoning” all happen in the teaching half.
- What do you have after the reading?
- A base model. Great at finishing a paragraph. Bad at being an assistant.
- What happens in the teaching half?
- First, example answers (supervised fine-tuning). Then a preference stage: RLHF, DPO, GRPO, or RLAIF. Reasoning models get one more round that rewards a correct final answer.
- Which file should I download?
- For chat, grab the instruct or chat file. The base file is the raw student. A reasoning or thinking file spends extra words thinking before it answers.
First it reads. Then someone teaches it the job.
The models you run at home — Llama, Qwen, Gemma, DeepSeek — are not the product of one training run. There are two jobs, and they feel completely different.
Pretraining is the reading. The model goes through web pages, code, and books, on the order of trillions of words, and the only assignment is “guess what comes next.” Do that for long enough and it picks up grammar, facts that show up a lot, and a feel for how code is written. The file you get at the end is called a base model.
Post-training is the teaching. The pile of data is much smaller. Think hundreds of thousands of carefully written examples, plus rankings of which answer is better, plus, for reasoning models, problems that have a checkable right answer. This is where a model learns to answer you, stay in a conversation, call a tool, turn down a bad request, and sometimes scribble on scratch paper before it commits.
After all that reading, it is still a bookworm
Hand a base model the start of a paragraph and it will write a believable next sentence. Hand it a question and it might answer. It might also write another question, or the opening of a blog post. Both of those show up constantly in the text it read, so both are fair guesses for “what comes next.” It has no built-in idea that you are a person waiting for help.
The knowledge is in there. The habit of being an assistant is not. Some labs then do continued pretraining, which is the same guessing game on a second pile of text from one field: medicine, law, finance, a company’s own docs. You get a specialist bookworm. You still do not have an assistant. That part comes later. The longer version is LLM Pretraining and Continued Pretraining.
The teaching, in the order it actually happens
1. Show it the answer you wanted
This step is called supervised fine-tuning, or SFT. You show the model finished examples: a question and a good answer, a chat, a tool call, a long document plus an answer that actually uses the document. It is still guessing the next word. The words it gets graded on are now the reply. After enough of these, it stops rambling and starts answering.
People sort these examples by the job: follow an instruction, hold a conversation, stay inside one field, read a long document, or call a tool. They also sort by how much of the model they update. Updating every weight is expensive. LoRA, QLoRA, and DoRA train a small add-on and leave the rest frozen. QLoRA is the version that fits a small model, around 7 billion parameters, on a normal gaming GPU. The plain-language tour is Supervised Fine-Tuning, LoRA, QLoRA, and DoRA.
2. Decide whose grade counts
Fine-tuning teaches the shape of a good answer. It does not settle which of two fluent answers a person would rather read. That is the alignment step. The idea is the same every time: make the better answer more likely. The methods differ in who holds the red pen.
| Method | Who gives the grade | The catch |
|---|---|---|
| RLHF + PPO | People rank answers. A second model learns to copy those rankings. | The classic setup. You train an extra model and run a heavy optimization loop. |
| DPO | The same human rankings, applied straight to the chatbot. | No extra grader model. This is what a lot of open models use now. |
| GRPO | Several answers to the same question, scored against each other. | No separate judge model. This is the method behind DeepSeek-R1. |
| RLAIF | An AI judge, often following a written list of rules. | You pay fewer people. You also inherit whatever the judge is biased toward. |
A walk-through of each one is in RLHF, DPO, GRPO, and RLAIF.
3. Reward the right final answer
Reasoning models get one more lesson. If a task has a checkable result — the math comes out right, the tests pass — you can reward the attempts that get there. The model learns to spend words on scratch work before the final answer, because the attempts that do that win more often. DeepSeek wrote this up for R1, using GRPO and rewards a program can check. OpenAI’s o-series is the other famous example of “think, then answer,” and OpenAI sells a version of the idea to customers as reinforcement fine-tuning. Same kind of result. Not the same published recipe. Read Reinforcement Fine-Tuning and Reasoning Models, and if you want to run one, the hardware picks are in Open-Weight Reasoning Models in 2026.
How to tell the files apart
You can see this whole story in the filename.
- Base — finished the reading, skipped the lessons. The right start if you plan to train it yourself. A poor chatbot.
- Instruct or chat — someone already showed it how to answer, and usually how to behave. This is the file to pull for everyday use.
- Reasoning or thinking — an extra lesson on top. Stronger on math, code, and anything with several steps. Slower, because it writes a long chain of thought first.
A medical or legal model usually sits in the middle: a general base model that read extra text from that field, then saw example answers from that field. The filename rarely says “continued pretraining.” The model card usually does.
The rest of this series
- LLM Pretraining and Continued Pretraining — the reading phase, and what a base model can and cannot do
- Supervised Fine-Tuning Explained — teaching it to answer, plus LoRA, QLoRA, and DoRA
- LLM Alignment Explained — RLHF, DPO, GRPO, and RLAIF
- Reinforcement Fine-Tuning Explained — why reasoning models think before they answer
- Open-Weight Reasoning Models in 2026 — which thinking models fit on a home computer