How it works
Before a model sees your text, a tokeniser splits it into tokens: common words are one token, rarer words are broken into pieces, and spaces and punctuation count too. In English a token averages about four characters, or three quarters of a word; many other languages, and code, need more tokens for the same content.
Providers bill per million tokens, and output tokens usually cost several times more than input. The context window is the total a model can take in one request (instructions, history, documents and the reply), and current models range from tens of thousands to about a million tokens. A full window is slower and dearer, and models can overlook details buried in the middle, so sending only what is relevant still pays off.
Related terms
More in AI and LLMs
Basics