RunLocalModel.com

Reinforcement Fine-Tuning Explained: Why Reasoning Models Think First

By the RunLocalModel editorial team · Published September 30, 2026 · ~8 minute read

If you only read one paragraph Showing a model a written solution only gets you so far. It can copy the solution. It cannot practice until it finds a better one. Reinforcement fine-tuning grades the final answer: the math matches, the tests pass, the rubric score is high. Tries that pass become more likely, including tries that spend a long time thinking first. That is the plain version of how DeepSeek-R1 learned to think before it answers, using GRPO and a checker, and of what OpenAI’s o-series and Reinforcement Fine-Tuning API are doing in public. The published recipes are different. The habit you see is the same kind of result. Which of those models fit on a home GPU is in Open-Weight Reasoning Models in 2026.
Quick answers
What is the grade?
A check on the final answer. A known correct result, a unit test, or a rubric. The score cares about the outcome.
How is this different from ordinary fine-tuning?
Ordinary fine-tuning copies a written solution. This method reinforces whichever of the model’s own tries the grader accepts, so it can learn a strategy nobody demonstrated.
Why the long thinking trace?
On hard problems, drafts with scratch work pass the grader more often. Training makes that pattern more likely.
Is this the same recipe at every lab?
No. DeepSeek published GRPO plus rewards a program can check. OpenAI published a grader-based API and has not published the full o-series recipe.

This is the last lesson in the map from How LLMs Are Trained. The reading supplied the knowledge. Supervised fine-tuning supplied the habit of answering. This lesson is how a model learns to work the problem before it commits.

Copying a solution only goes so far

You can teach a chain of thought with ordinary fine-tuning. Write the scratch work into the example, and the model will imitate it. That helps, and then it stalls. The model can only echo strategies that appear in the files you labeled. On a contest math problem, or a bug with a failing test, the useful move is often one the model has to try, throw out, and replace. Copying has no signal for “this attempt is the one that actually worked.”

Reinforcement fine-tuning adds that signal. The training set is a pile of problems whose answers can be checked. The model takes a shot, often a long trace and then a final answer. A grader scores the outcome. Shots that score well get reinforced. Shots that score poorly do not. Over many problems, the traces that tend to come before a correct answer become the model’s default behavior.

The grade is “the final answer checks out.” The thinking-out-loud part shows up because those drafts win more often.

What counts as a grader

The grader has to be trustworthy enough to train on. The clean cases are automatic:

OpenAI’s Reinforcement Fine-Tuning guide describes the same deal for customers: a task with an answer you can verify, and a grader — code, or a model rubric — that scores the candidates. A model can grade fuzzier tasks. Then you inherit the judge’s mistakes, which is the same caution as RLAIF. This method earns its reputation on problems where a wrong score is rare.

DeepSeek-R1: a recipe you can read

The DeepSeek-R1 technical report is the clearest public account of this lesson. The optimizer is GRPO: for one problem, sample a group of tries, score them, and update based on how each try compares with the group. There is no separate learned grader in the PPO sense. The scores come from rules. A math answer matches. A program passes tests. The trace follows the think-then-answer shape.

The striking result is that this loop, run at scale, produces long chains of thought — checking itself, backing up, doing the arithmetic in the middle — as a habit that just shows up. Nobody had to write every reasoning trace by hand. The traces that survived were the ones that ended in a checkable success. DeepSeek then copied those traces into smaller models. The small ones are what most people can run. The full mixture-of-experts model is a datacenter machine. The 14B and 32B distillations are the local ones. Picks by GPU size are in Open-Weight Reasoning Models in 2026.

The o-series, and OpenAI’s version of the same idea

OpenAI’s o-series is the other public face of “think, then answer.” The o1 system card describes models trained with large-scale reinforcement learning to reason before the reply you see. OpenAI later shipped Reinforcement Fine-Tuning so a customer can keep training a reasoning model against their own grader.

Those are related ideas. They are not a claim that both labs run the same code. DeepSeek published GRPO, the reward rules, and the step that copies the traces into smaller models. OpenAI has published the behavior and a customer API, and has not published the full o-series training procedure. When you hear that this is “the secret” behind both, read it as: both reinforce answers you can check, and that pressure produces a think-first habit. The optimizer, the scale, and the unpublished details differ.

What you notice when you run one

A reasoning model spends words before the answer you came for. A hard problem might be a short prompt, then one to several thousand words of scratch work, then a short final answer. That scratch work is the capability. It is also the cost: waiting, context window, and the memory that context needs. For email, translation, and a quick chat, the extra words buy little. For contest math, a stubborn bug, or a plan that has to stay consistent across steps, they are the reason the answer is right.

The practical setup — Ollama, llama.cpp, context length, and which quant fits 12 GB or 24 GB — lives in the reasoning-model hardware guide. This page is about why those models behave that way. If you are choosing a quant or a GPU for the weights themselves, start from Choosing Quantization and the VRAM versus unified memory comparison, then leave room for the thinking trace.

When this is the wrong tool

This method needs a grader you trust. Taste, tone, and open-ended writing rarely have one. For those, preference data — DPO, or an AI judge you have actually checked — is the matching tool. It is also a poor way to install basic “please just answer the question.” That is still ordinary fine-tuning’s job. The usual order stands: reading for knowledge, example answers for the ability to do the work, then a reward stage so the model’s own search lands on answers you can check.

Related guides on this site