Reinforcement Fine-Tuning Explained: Why Reasoning Models Think First
- What is the grade?
- A check on the final answer. A known correct result, a unit test, or a rubric. The score cares about the outcome.
- How is this different from ordinary fine-tuning?
- Ordinary fine-tuning copies a written solution. This method reinforces whichever of the model’s own tries the grader accepts, so it can learn a strategy nobody demonstrated.
- Why the long thinking trace?
- On hard problems, drafts with scratch work pass the grader more often. Training makes that pattern more likely.
- Is this the same recipe at every lab?
- No. DeepSeek published GRPO plus rewards a program can check. OpenAI published a grader-based API and has not published the full o-series recipe.
This is the last lesson in the map from How LLMs Are Trained. The reading supplied the knowledge. Supervised fine-tuning supplied the habit of answering. This lesson is how a model learns to work the problem before it commits.
Copying a solution only goes so far
You can teach a chain of thought with ordinary fine-tuning. Write the scratch work into the example, and the model will imitate it. That helps, and then it stalls. The model can only echo strategies that appear in the files you labeled. On a contest math problem, or a bug with a failing test, the useful move is often one the model has to try, throw out, and replace. Copying has no signal for “this attempt is the one that actually worked.”
Reinforcement fine-tuning adds that signal. The training set is a pile of problems whose answers can be checked. The model takes a shot, often a long trace and then a final answer. A grader scores the outcome. Shots that score well get reinforced. Shots that score poorly do not. Over many problems, the traces that tend to come before a correct answer become the model’s default behavior.
What counts as a grader
The grader has to be trustworthy enough to train on. The clean cases are automatic:
- An exact answer — a number, a short string, a multiple-choice key.
- Tests — code that compiles and passes a hidden suite.
- A format check — the reply put the reasoning in the expected block and the answer in the expected block. DeepSeek-R1 uses format rewards next to correctness rewards.
OpenAI’s Reinforcement Fine-Tuning guide describes the same deal for customers: a task with an answer you can verify, and a grader — code, or a model rubric — that scores the candidates. A model can grade fuzzier tasks. Then you inherit the judge’s mistakes, which is the same caution as RLAIF. This method earns its reputation on problems where a wrong score is rare.
DeepSeek-R1: a recipe you can read
The DeepSeek-R1 technical report is the clearest public account of this lesson. The optimizer is GRPO: for one problem, sample a group of tries, score them, and update based on how each try compares with the group. There is no separate learned grader in the PPO sense. The scores come from rules. A math answer matches. A program passes tests. The trace follows the think-then-answer shape.
The striking result is that this loop, run at scale, produces long chains of thought — checking itself, backing up, doing the arithmetic in the middle — as a habit that just shows up. Nobody had to write every reasoning trace by hand. The traces that survived were the ones that ended in a checkable success. DeepSeek then copied those traces into smaller models. The small ones are what most people can run. The full mixture-of-experts model is a datacenter machine. The 14B and 32B distillations are the local ones. Picks by GPU size are in Open-Weight Reasoning Models in 2026.
The o-series, and OpenAI’s version of the same idea
OpenAI’s o-series is the other public face of “think, then answer.” The o1 system card describes models trained with large-scale reinforcement learning to reason before the reply you see. OpenAI later shipped Reinforcement Fine-Tuning so a customer can keep training a reasoning model against their own grader.
Those are related ideas. They are not a claim that both labs run the same code. DeepSeek published GRPO, the reward rules, and the step that copies the traces into smaller models. OpenAI has published the behavior and a customer API, and has not published the full o-series training procedure. When you hear that this is “the secret” behind both, read it as: both reinforce answers you can check, and that pressure produces a think-first habit. The optimizer, the scale, and the unpublished details differ.
What you notice when you run one
A reasoning model spends words before the answer you came for. A hard problem might be a short prompt, then one to several thousand words of scratch work, then a short final answer. That scratch work is the capability. It is also the cost: waiting, context window, and the memory that context needs. For email, translation, and a quick chat, the extra words buy little. For contest math, a stubborn bug, or a plan that has to stay consistent across steps, they are the reason the answer is right.
The practical setup — Ollama, llama.cpp, context length, and which quant fits 12 GB or 24 GB — lives in the reasoning-model hardware guide. This page is about why those models behave that way. If you are choosing a quant or a GPU for the weights themselves, start from Choosing Quantization and the VRAM versus unified memory comparison, then leave room for the thinking trace.
When this is the wrong tool
This method needs a grader you trust. Taste, tone, and open-ended writing rarely have one. For those, preference data — DPO, or an AI judge you have actually checked — is the matching tool. It is also a poor way to install basic “please just answer the question.” That is still ordinary fine-tuning’s job. The usual order stands: reading for knowledge, example answers for the ability to do the work, then a reward stage so the model’s own search lands on answers you can check.
Related guides on this site
- How LLMs Are Trained — the full pipeline
- LLM Alignment Explained — GRPO next to RLHF, DPO, and RLAIF
- Open-Weight Reasoning Models in 2026 — what fits on a home computer
- Supervised Fine-Tuning Explained
- Choosing Quantization