From Next Word Predictor to Chatbot: Pretraining, Fine Tuning and RLHF

By Sheng Pang · Published · 8 min read

After part four we have a transformer trained to predict the next token on a huge pile of text. That model is powerful and nearly useless as a product. Ask it "What is the capital of France?" and it may reply "What is the capital of Germany? What is the capital of Spain?" because on the internet, a list of quiz questions is a common thing to follow a quiz question. This post covers the stages that turn the raw model into an assistant, and the settings you meet when you use one.

Stage 1: pretraining

This is the expensive part and the one the earlier posts described. Take trillions of tokens of text: web crawl, books, code, papers, transcripts. At each position, predict the next token. Measure the loss, backpropagate, update the weights. Run on thousands of graphics cards for months.

The result is a base model, sometimes called a foundation model. It has absorbed grammar, facts, styles and reasoning patterns as a side effect of getting good at prediction. It is a completion engine: give it the start of anything and it continues in the most likely way. It has no notion of being asked a question or being helpful. It is imitating the internet.

Base models are what researchers mean when they discuss what "the model knows". Everything after this stage is about steering that knowledge, not adding much to it.

Stage 2: supervised fine tuning

To make the model behave like an assistant, you show it what an assistant looks like. Humans write thousands of example conversations: a user message, then an ideal reply. The model is trained on these with the same next token prediction, but now the text it learns to continue is formatted as a dialogue and the continuation is always a helpful answer.

User: What is the capital of France?
Assistant: The capital of France is Paris.

This is a tiny amount of data compared to pretraining, tens of thousands of examples rather than trillions of tokens, and it takes hours or days rather than months. But it changes the behaviour completely. The model learns the format, learns to stop after answering, learns to say "I" and to address the user. This is instruction tuning or supervised fine tuning. Fine tuning in general means further training an already trained model on a smaller specialised dataset, and companies use the same technique to specialise a model on their own documents or style.

Stage 3: learning from human preferences

Writing perfect example answers is slow and does not cover everything. It is much easier for a human to look at two answers and say which is better. That comparison is the basis of reinforcement learning from human feedback, RLHF.

It works in two steps:

  • Train a reward model. Generate several answers to many prompts. Have people rank them. Train a separate network to predict, for any answer, the score a human would give it. This reward model is a learned stand in for human judgement.
  • Optimise against the reward model. Let the chat model produce answers, score them with the reward model, and use reinforcement learning, the third kind of learning from part one, to nudge the weights toward answers that score higher. The model is no longer imitating examples. It is exploring and being rewarded.

This is where models learn to be polite, to refuse harmful requests, to admit uncertainty and to format answers with headings and lists. It is also where some of their annoying habits come from. If raters slightly prefer confident, longer, agreeable answers, the model becomes more confident, longer and more agreeable than it should be. The industry is still working on this. Newer variants replace the human raters with rules or with another model for parts of the process, and use reinforcement learning to reward correct final answers on maths and code, which is how "reasoning" models that think out loud before answering are trained.

What you get: a chat model

The finished product is a chat model or instruct model. When you use it, your conversation is formatted into a single text with special tokens marking who said what, and the model continues it from the assistant's turn. Under the hood it is still the same next token predictor. The training has simply made "a helpful assistant's reply" the most likely continuation.

The system prompt

Before your first message, the product inserts a hidden block of text: the system prompt. It tells the model who it is, today's date, what tools it has, house rules about tone and topics. When a chatbot knows the date or refuses a topic, that usually comes from the system prompt plus RLHF, not from the weights. If you use a model through an API, you write the system prompt yourself, and it is the most effective single lever you have over behaviour.

Temperature and sampling

Part three ended with the model producing a probability for every token in the vocabulary. Something has to pick one. The options:

  • Greedy. Always take the most likely token. Deterministic but tends to be repetitive and dull.
  • Sampling. Pick randomly in proportion to the probabilities. "Paris" at 92 percent is chosen 92 times out of 100.
  • Temperature. A knob that reshapes the probabilities before sampling. Temperature 0 is greedy. Temperature 1 samples as is. Higher temperatures flatten the distribution so unlikely tokens get chosen more often. Low temperature for code and data extraction where you want consistency. Higher for brainstorming where you want variety.

This is why the same question gives different answers on different runs, and why setting temperature to 0 makes output far more repeatable, though not perfectly so, because the underlying arithmetic on graphics cards is not always bit for bit identical.

Why models hallucinate

Now the pieces are in place to explain the most important failure mode. A hallucination is a fluent, confident, false statement: a citation that does not exist, a function that was never in the library, a date that is wrong.

It happens for reasons built into the training:

  • Pretraining rewards plausibility. The loss only cares whether the next token is likely. A made up paper title in the correct format is a perfectly likely sequence of tokens.
  • The model must always output something. There is no "I do not know" token that wins by default. Saying so has to be taught in fine tuning, and it competes with the pretraining habit of continuing confidently.
  • Knowledge is compressed. Hundreds of billions of weights cannot store the internet verbatim. Rare facts are stored fuzzily or not at all, and the model fills gaps with the pattern of similar facts. It knows what a citation looks like far better than it knows any specific citation.
  • Raters like confidence. If human feedback penalised hedging more than it penalised errors, the model learned to hedge less.

Things that reduce it: giving the model the relevant documents in the prompt so it can read rather than recall, which is retrieval augmented generation; letting it call tools like search or a code runner; asking it to reason step by step before answering; lower temperature; and asking for sources you can check. None of them remove it. Treat model output as a strong first draft from a well read colleague who never says "I am not sure".

Knowledge cutoff

Pretraining data stops at some date, so the model knows nothing after it unless the product feeds it search results. This is the knowledge cutoff. A model asked about last week's news will either say it does not know, if fine tuning taught it to, or hallucinate a plausible answer, if it did not.

The whole series in one paragraph

AI is the broad field; machine learning finds rules from examples; deep learning does it with many layered networks; a language model is a deep network trained to predict the next token. Training means adjusting weights by gradient descent to lower a loss. Text enters as tokens, becomes embedding vectors, passes through transformer layers where attention lets each token gather context, and comes out as a probability over the next token. Pretraining on the internet makes a base model that can continue anything; fine tuning and RLHF shape it into an assistant; a system prompt and a temperature setting govern each conversation. The model is fluent because it was trained on fluent text, and it is sometimes wrong for the same reason.

Where to go from here

You now have the background to read almost any AI announcement and understand what actually changed. Our AI in practice posts pick up from here with the tools built on top of these models: coding agents, MCP, running a model locally and getting reliable structured output. And if you want to see tokens in action, paste a JSON document into our JSON formatter, minify it, and notice how much shorter it gets. That is real tokens saved on every request.

← Back to all articles