A model trained to predict text does not automatically follow instructions.

This final part of the AI basics series explains pretraining, instruction tuning, human feedback and the settings used when you call a chat model.

Stage 1: pretraining

This is the expensive part and the one the earlier posts described.

Take trillions of tokens of text: web crawl, books, code, papers, transcripts. At each position, predict the next token.

Measure the loss, backpropagate, update the weights.

Run on thousands of graphics cards for months.

The result is a base model, sometimes called a foundation model.

It has absorbed grammar, facts, styles and reasoning patterns as a side effect of getting good at prediction. It is a completion engine: give it the start of anything and it continues in the most likely way.

It has no notion of being asked a question or being helpful.

It is imitating the internet.

Base models are what researchers mean when they discuss what "the model knows". Everything after this stage is about steering that knowledge, not adding much to it.

Stage 2: supervised fine tuning

To make the model behave like an assistant, you show it what an assistant looks like.

Humans write thousands of example conversations: a user message, then an ideal reply. The model is trained on these with the same next token prediction, but now the text it learns to continue is formatted as a dialogue and the continuation is always a helpful answer.

User: What is the capital of France?
Assistant: The capital of France is Paris.

This is a tiny amount of data compared to pretraining, tens of thousands of examples rather than trillions of tokens, and it takes hours or days rather than months.

But it changes the behaviour completely. The model learns the format, learns to stop after answering, learns to say "I" and to address the user.

This is instruction tuning or supervised fine tuning.

Fine tuning in general means further training an already trained model on a smaller specialised dataset, and companies use the same technique to specialise a model on their own documents or style.

Stage 3: learning from human preferences

Writing perfect example answers is slow and does not cover everything.

It is much easier for a human to look at two answers and say which is better. That comparison is the basis of reinforcement learning from human feedback, RLHF.

It works in two steps:

Train a reward model. Generate several answers to many prompts. Have people rank them. Train a separate network to predict, for any answer, the score a human would give it. This reward model is a learned stand in for human judgement.

Optimise against the reward model. Let the chat model produce answers, score them with the reward model, and use reinforcement learning, the third kind of learning from part one, to nudge the weights toward answers that score higher. The model is no longer imitating examples. It is exploring and being rewarded.

This is where models learn to be polite, to refuse harmful requests, to admit uncertainty and to format answers with headings and lists.

It is also where some of their annoying habits come from. If raters slightly prefer confident, longer, agreeable answers, the model becomes more confident, longer and more agreeable than it should be.

The industry is still working on this.

Some variants replace human raters with rules or another model. Reinforcement learning can also reward correct answers to maths and code problems.

What you get: a chat model

The finished product is a chat model or instruct model.

When you use it, your conversation is formatted into a single text with special tokens marking who said what, and the model continues it from the assistant's turn. Under the hood it is still the same next token predictor.

The training has simply made "a helpful assistant's reply" the most likely continuation.

The system prompt

Before your first message, the product inserts a hidden block of text: the system prompt. It tells the model who it is, today's date, what tools it has, house rules about tone and topics.

When a chatbot knows the date or refuses a topic, that usually comes from the system prompt plus RLHF, not from the weights.

If you use a model through an API, you write the system prompt yourself, and you can use it to set instructions for that request.

Temperature and sampling

Part three ended with the model producing a probability for every token in the vocabulary.

Something has to pick one. The options:

Greedy. Always take the most likely token. Deterministic but tends to be repetitive and dull.

Sampling. Pick randomly in proportion to the probabilities. "Paris" at 92 percent is chosen 92 times out of 100.

Temperature. A knob that reshapes the probabilities before sampling. Temperature 0 is greedy. Temperature 1 samples as is. Higher temperatures flatten the distribution so unlikely tokens get chosen more often. Low temperature for code and data extraction where you want consistency. Higher for brainstorming where you want variety.

Sampling can produce different answers to the same question. Temperature 0 makes selection more repeatable, but it does not guarantee identical output across runs.

Why models hallucinate

Now the pieces are in place to explain the most important failure mode.

A hallucination is a fluent, confident, false statement: a citation that does not exist, a function that was never in the library, a date that is wrong.

It happens for reasons built into the training:

Pretraining rewards plausibility. The loss only cares whether the next token is likely. A made up paper title in the correct format is a perfectly likely sequence of tokens.

The model must always output something. There is no "I do not know" token that wins by default. Saying so has to be taught in fine tuning, and it competes with the pretraining habit of continuing confidently.

Knowledge is compressed. Hundreds of billions of weights cannot store the internet verbatim. Rare facts are stored fuzzily or not at all, and the model fills gaps with the pattern of similar facts. It knows what a citation looks like far better than it knows any specific citation.

Raters like confidence. If human feedback penalised hedging more than it penalised errors, the model learned to hedge less.

Give the model relevant documents and access to tools when appropriate.

Test retrieval, reasoning instructions and temperature on your task, then use our AI search fact checking steps to check its sources. None of them remove it.

Treat model output as a strong first draft from a well read colleague who never says "I am not sure".

Knowledge cutoff

Pretraining data stops at some date, so the model knows nothing after it unless the product feeds it search results. This is the knowledge cutoff.

A model asked about last week's news will either say it does not know, if fine tuning taught it to, or hallucinate a plausible answer, if it did not.

The whole series in one paragraph

AI is the broad field; machine learning finds rules from examples; deep learning does it with many layered networks; a language model is a deep network trained to predict the next token. Training means adjusting weights by gradient descent to lower a loss.

Text enters as tokens, becomes embedding vectors, passes through transformer layers where attention lets each token gather context, and comes out as a probability over the next token.

Pretraining on the internet makes a base model that can continue anything; fine tuning and RLHF shape it into an assistant; a system prompt and a temperature setting govern each conversation.

The model is fluent because it was trained on fluent text, and it is sometimes wrong for the same reason.

Where to go from here

Our AI in practice guides cover coding agents, MCP, local models and structured output.

Try minifying a document in our JSON formatter, then compare the original and result with your model's tokenizer. Check warnings before using the output; minification does not guarantee fewer tokens.