Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Learn · 6 of 7

How a model gets made

Four stages, and the fact that almost all of the behaviour comes from the cheap ones.

Four stages, and their costs are wildly unequal in a way that explains a lot about how these systems behave.

Pretraining: months, and nearly all the money

A model starts as random numbers and reads an enormous body of text, adjusting itself each time to predict the next token slightly better. This runs for months across thousands of accelerators, costs tens or hundreds of millions of dollars, and is where essentially everything the model knows comes from.

What comes out is a base model (The first and most expensive stage: learning to predict the next token across a very large body of text.): fluent, knowledgeable, and useless. It completes text rather than answering questions, and it has no notion that it ought to be helpful.

Fine-tuning: days

Training continues on a much smaller, curated set of example conversations. This is where a text completer becomes something that answers a question when asked one.

It is remarkably cheap compared to the first stage, and remarkably consequential.

Preference tuning: days

For most of what an assistant does there is no single right answer, so there is nothing to train against directly. Instead people compare pairs of answers, a second model learns to predict those preferences, and the first is trained to score well against it.

Tone comes from here. So does helpfulness, and so do refusals. And so does one of the characteristic failures: a model trained to produce answers people prefer learns, among other things, that people prefer confident answers.

Evaluation and release: ongoing

Score it, decide whether to ship it, keep scoring it. The published benchmark (A repeatable test set you score a model against, so a change can be shown to help rather than assumed to.) figures come from here, and they are worth less than they look — mostly self-reported, with test sets that leak into training data over time.

The thing worth taking away

Nearly all the compute is in stage one, and nearly all the behaviour you interact with comes from stages two and three. When a model feels different after an update, it is rarely because it learned new facts.

How a model gets made

Width shows roughly where the compute goes

The four stages of training a modelFour stages left to right — pretraining, fine-tuning, preference tuning, evaluation and release — drawn as bands whose widths show roughly where the compute goes. Pretraining dwarfs the rest. Each stage is described in the list below.1months, most of the moneyPretraining2daysFine-tuning3daysPreference tuning4ongoingEvaluation and releaseYou are never charged for any of this. What you pay for is inference — running the finished model.
  1. Pretraining (months, most of the money). Predict the next token across an enormous body of text. This is where almost everything the model knows comes from, and almost all of the cost.
  2. Fine-tuning (days). Train on curated example conversations. This is what turns a text completer into something that answers a question.
  3. Preference tuning (days). Train against human judgements of which answer is better. Tone, helpfulness and refusals come from here.
  4. Evaluation and release (ongoing). Score it, decide whether to ship it, and keep scoring it afterwards.
ShareOpen LinkedIn