Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Glossary

Token

The unit a model reads and writes. Roughly three quarters of an English word, so 1,000 tokens is about 750 words.

Also written: tokens.

A model does not see letters or words. It sees tokens: fragments of text that its tokenizer (The fixed table that chops text into tokens. Decided before training and never changed afterwards.) carved the language into before training began.

For ordinary English, one token is about four characters, so 1,000 tokens is roughly 750 words, or a page and a half. That ratio is a rule of thumb and not a rule. Code tokenises less efficiently than prose because punctuation and indentation each cost something. Rare words split into several tokens. Languages that do not use spaces, and languages written in scripts the tokenizer saw little of, can cost two or three times more tokens for the same meaning — which means the same question costs more to ask in some languages than in others.

Everything you are charged for is counted in tokens, in both directions, and almost always at a different rate each way. Output usually costs several times more than input.

ShareOpen LinkedIn