RunLocalModel.com

LLM Pretraining Explained: Base Models and Continued Pretraining

By the RunLocalModel editorial team · Published September 30, 2026 · ~9 minute read

If you only read one paragraph Pretraining is the long first chapter. The model reads the web, code, and books — trillions of words — and practices one move the whole time: guess the next word. What you get is a base model. It writes smoothly and it knows a lot of facts. Ask it a question and it often just keeps the document going. Continued pretraining is the same practice on a pile of text from one field, so the bookworm changes majors. It is still a bookworm. The assistant comes later.
Quick answers
What is it actually practicing?
Guessing the next word. See some text, predict what follows, get told the real next word, adjust a little. Repeat.
What do you have at the end?
A base model. It continues text. It has not been taught to be a chatbot.
What is continued pretraining?
The same guessing game, on text from one field. You get a specialist base model, still waiting for someone to teach it to answer.
Which file should I run?
Instruct or chat, if you want a conversation. Base, only if you are about to train it yourself.

This is the reading chapter of the series. The whole map is in How LLMs Are Trained.

The only homework: guess the next word

People say “token” in papers. A token is a chunk of text, usually a word or a piece of a word. For the rest of this page, “word” is close enough.

Pretraining shows the model a long stretch of text and, at every position, asks what comes next. If it guesses wrong, the weights nudge so that the real next word would have been more likely. Nobody labels “this is the question” and “this is the answer.” The text itself is the answer key for the word that follows it.

That sounds too small to produce something that can write code or explain a concept. The pile of text is what makes it work. A modern pretraining run is measured in trillions of words, taken from the public web, books, and big code collections. Guessing the next word on that mix forces the model to learn things that actually help: grammar, how an argument usually unfolds, facts that keep showing up, and the habits of programming languages. GPT-3 made this recipe famous. Later base models do the same homework at a much larger scale.

It is only allowed to look backward. It cannot peek at the words that have not been written yet. That is why a finished base model writes one word at a time, and each new word becomes part of the context for the next guess.

What all that reading sticks in its head

After pretraining, a base model is fluent. It can finish a sentence in a consistent voice, fill in a common fact, and continue a function in a language it has seen enough of. A huge amount of written culture has been squeezed into the weights.

It also picked up the habits of the internet. Web text includes questions followed by answers. It also includes questions followed by more questions, ads, arguments, and half-finished posts. A base model copies those habits. Start it with “The capital of France is” and the next words are often “Paris,” because that sentence is common. Start it with “What is the capital of France?” and a reasonable continuation might be an answer, or another quiz question, or the first line of an article. All of those are legal guesses. None of them means the model knows it is supposed to stop, be brief, and talk to you.

A base model is a well-read bookworm. The knowledge is in there. The job of answering you is not.

Why people call it a bookworm

The nickname fits. The model has “read” more text than any person, and when you talk to it, it does what autocomplete does: it keeps the document going. Ask for a summary and you may get a summary, because summaries exist in the data. You may also get a preamble, a heading, or a second document. There is no chat costume baked in, no system role, and no trained instinct to say no or to call a tool.

That is why labs do not hand you the raw base file as the thing you chat with. The product is the base model after supervised fine-tuning and, usually, a preference stage. The reading decides how much it knows. Those later lessons decide whether it will do the work.

Continued pretraining: a second library, one subject

A general base model has seen a little of everything, and a lot of whatever dominated the web. A hospital, a law firm, or a chip team often wants the opposite: less random internet, more of their field. Continued pretraining (people also say CPT, or domain-adaptive pretraining) is how they do that.

You start from the general weights and keep the same homework: guess the next word. The new reading list is narrow. Textbooks and papers. Statutes. Filings. Internal wikis. A programming language the general mix barely covered. Don’t Stop Pretraining showed that this second pass teaches a field better than jumping straight to a small stack of labeled questions and answers. The model’s guesses inside that field get sharper, because the field is now a large share of what it has just been reading.

The homework did not change. The model is still practicing “keep writing.” A medical CPT file is better at medical prose, and it is still happy to ramble, invent a heading, or finish a case note as if it were the next paragraph of a textbook. The assistant a clinician would actually talk to is a later file, after someone shows it medical questions and good answers. Labs often stack the two and call the result a domain model. The extra reading is the knowledge pass. The fine-tune is the behavior pass.

There is a limit. If the new pile of text is small, or narrow, or written in one house style, the model can drift toward that style and get worse at ordinary writing. Serious runs mix some general text back in, so it keeps the skills it already had. The model card is where you find out whether a “legal” or “medical” release did this, and how much text they used.

What this means when you download a file

On Hugging Face and in Ollama, the useful clue is the name, plus the model card.

Pretraining is also why a local model knows as much as it does before you upload anything. The public facts and the code patterns are already in the weights. Your private documents are not. If those documents matter, you can keep training on them, or you can paste them into the prompt. Pasting is what most people should try first. Extra pretraining needs a corpus and a training budget.

Related guides on this site