Reference
The words, explained
33 terms, defined without assuming you already know the other32. Each one appears as a dotted underline wherever it is used on this site — and the Explain jargon button in the header puts the definition inline, in the sentence, so you never have to leave the page you are reading.
33 terms.
How models work
- Context windowThe total tokens a model can consider at once: your prompt, the conversation so far, and the answer it is writing.
- EmbeddingA list of numbers representing a piece of text, arranged so that similar meanings land near each other.
- InferenceRunning a trained model to get an answer. Everything you pay for by the token is inference, not training.
- Knowledge cutoffThe date after which a model saw no training data. It knows nothing later unless you tell it or it can search.
- Large language modelA model trained to predict the next token over enormous amounts of text, which turns out to be enough to do a great deal.
- MultimodalA model that accepts more than text — images, audio or video — in the same prompt.
- Reasoning modelA model trained to produce a long internal working-out before its answer. More accurate on hard problems, slower and dearer.
- TemperatureHow much randomness to allow when picking each token. Low is repeatable and dull; high is varied and unreliable.
- TokenThe unit a model reads and writes. Roughly three quarters of an English word, so 1,000 tokens is about 750 words.
- TokenizerThe fixed table that chops text into tokens. Decided before training and never changed afterwards.
What things cost
- Cache writeThe charge for putting a prefix into the prompt cache. Usually more than the input rate, not less.
- Per million tokensHow model prices are quoted: dollars per million tokens, charged separately for what you send and what comes back.
- Prompt cachingPaying a reduced rate for a prefix the vendor has already processed, instead of full price for sending it again.
- Rate limitA cap on how much you may send in a window, usually counted in requests and tokens per minute.
Getting good results
- Chain of thoughtAsking a model to work through its reasoning before answering, which measurably improves accuracy on multi-step problems.
- Context rotThe tendency of a model to attend less to material buried in the middle of a long context than to its start or end.
- Few-shot promptingShowing the model two or three worked examples of what you want, instead of describing it.
- HallucinationA fluent, confident, entirely invented answer. The failure mode that matters, because it looks exactly like a correct one.
- PromptEverything you send the model on one call: instructions, context, history and question together.
- Retrieval-augmented generationLooking up relevant documents and putting them in the prompt, so the model reads facts rather than recalling them.
- System promptStanding instructions sent ahead of every conversation, setting the role, rules and format.
Agents and tools
- AgentA model given tools and a goal, left to loop — deciding what to do next and doing it — until the job is done or it gives up.
- Agent loopPlan, call a tool, read the result, decide again. The cycle an agent repeats until it finishes or is stopped.
- Computer useA model driving a screen directly — looking at pixels, moving a cursor, typing — instead of calling an API.
- Model Context ProtocolAn open standard for connecting models to tools and data, so one integration works across many applications.
- Tool useLetting a model call functions you define — search, read a file, send a request — and read back what they return.
How models are made
- EvalsA repeatable test set you score a model against, so a change can be shown to help rather than assumed to.
- Fine-tuningContinuing training on a smaller, curated set of examples to specialise a model or teach it a format.
- PretrainingThe first and most expensive stage: learning to predict the next token across a very large body of text.
- Reinforcement learning from human feedbackTraining a model against human judgements of which answer is better, rather than against a single correct answer.
The landscape
- Frontier modelThe largest and most capable model a vendor currently offers, priced accordingly.
- Open weightsA model whose trained parameters are published, so anyone can run it on their own hardware.
- Vibe codingBuilding software by describing what you want and iterating on what the model produces, rather than writing it line by line.