RunLocalModel.com

How LLMs Are Trained: Pretraining, Fine-Tuning, and Alignment

By the RunLocalModel editorial team · Published September 30, 2026 · ~8 minute read

If you only read one paragraph A large language model is trained in two big steps. First it reads a mountain of text and practices one trick: guess the next word. That produces a base model. It knows a lot. It only knows how to keep writing. Then people teach it the job. Fine-tuning shows it how to answer. Alignment and reasoning training push it toward answers you can actually trust, and, for the newest models, toward thinking before it speaks. One line: pretraining decides how much it knows, fine-tuning decides whether it can do the work, and the rest of post-training decides whether you can rely on it.
Quick answers
What are the two steps?
Reading, then teaching. The reading is pretraining. The teaching is post-training. “Instruct,” “chat,” and “reasoning” all happen in the teaching half.
What do you have after the reading?
A base model. Great at finishing a paragraph. Bad at being an assistant.
What happens in the teaching half?
First, example answers (supervised fine-tuning). Then a preference stage: RLHF, DPO, GRPO, or RLAIF. Reasoning models get one more round that rewards a correct final answer.
Which file should I download?
For chat, grab the instruct or chat file. The base file is the raw student. A reasoning or thinking file spends extra words thinking before it answers.

First it reads. Then someone teaches it the job.

The models you run at home — Llama, Qwen, Gemma, DeepSeek — are not the product of one training run. There are two jobs, and they feel completely different.

Pretraining is the reading. The model goes through web pages, code, and books, on the order of trillions of words, and the only assignment is “guess what comes next.” Do that for long enough and it picks up grammar, facts that show up a lot, and a feel for how code is written. The file you get at the end is called a base model.

Post-training is the teaching. The pile of data is much smaller. Think hundreds of thousands of carefully written examples, plus rankings of which answer is better, plus, for reasoning models, problems that have a checkable right answer. This is where a model learns to answer you, stay in a conversation, call a tool, turn down a bad request, and sometimes scribble on scratch paper before it commits.

Pretraining decides how much the model knows. Supervised fine-tuning decides whether it can do the work. Alignment and reasoning training decide whether you can rely on it.

After all that reading, it is still a bookworm

Hand a base model the start of a paragraph and it will write a believable next sentence. Hand it a question and it might answer. It might also write another question, or the opening of a blog post. Both of those show up constantly in the text it read, so both are fair guesses for “what comes next.” It has no built-in idea that you are a person waiting for help.

The knowledge is in there. The habit of being an assistant is not. Some labs then do continued pretraining, which is the same guessing game on a second pile of text from one field: medicine, law, finance, a company’s own docs. You get a specialist bookworm. You still do not have an assistant. That part comes later. The longer version is LLM Pretraining and Continued Pretraining.

The teaching, in the order it actually happens

1. Show it the answer you wanted

This step is called supervised fine-tuning, or SFT. You show the model finished examples: a question and a good answer, a chat, a tool call, a long document plus an answer that actually uses the document. It is still guessing the next word. The words it gets graded on are now the reply. After enough of these, it stops rambling and starts answering.

People sort these examples by the job: follow an instruction, hold a conversation, stay inside one field, read a long document, or call a tool. They also sort by how much of the model they update. Updating every weight is expensive. LoRA, QLoRA, and DoRA train a small add-on and leave the rest frozen. QLoRA is the version that fits a small model, around 7 billion parameters, on a normal gaming GPU. The plain-language tour is Supervised Fine-Tuning, LoRA, QLoRA, and DoRA.

2. Decide whose grade counts

Fine-tuning teaches the shape of a good answer. It does not settle which of two fluent answers a person would rather read. That is the alignment step. The idea is the same every time: make the better answer more likely. The methods differ in who holds the red pen.

Method Who gives the grade The catch
RLHF + PPO People rank answers. A second model learns to copy those rankings. The classic setup. You train an extra model and run a heavy optimization loop.
DPO The same human rankings, applied straight to the chatbot. No extra grader model. This is what a lot of open models use now.
GRPO Several answers to the same question, scored against each other. No separate judge model. This is the method behind DeepSeek-R1.
RLAIF An AI judge, often following a written list of rules. You pay fewer people. You also inherit whatever the judge is biased toward.

A walk-through of each one is in RLHF, DPO, GRPO, and RLAIF.

3. Reward the right final answer

Reasoning models get one more lesson. If a task has a checkable result — the math comes out right, the tests pass — you can reward the attempts that get there. The model learns to spend words on scratch work before the final answer, because the attempts that do that win more often. DeepSeek wrote this up for R1, using GRPO and rewards a program can check. OpenAI’s o-series is the other famous example of “think, then answer,” and OpenAI sells a version of the idea to customers as reinforcement fine-tuning. Same kind of result. Not the same published recipe. Read Reinforcement Fine-Tuning and Reasoning Models, and if you want to run one, the hardware picks are in Open-Weight Reasoning Models in 2026.

How to tell the files apart

You can see this whole story in the filename.

A medical or legal model usually sits in the middle: a general base model that read extra text from that field, then saw example answers from that field. The filename rarely says “continued pretraining.” The model card usually does.

The rest of this series