Tokenization
The conversion of input text or other content into discrete model units whose integer identifiers are processed by a neural network.
Language-Model Context
Tokenization converts input into the discrete units and integer identifiers processed by a model. Subword methods reduce out-of-vocabulary problems while controlling vocabulary size, but token efficiency differs across languages and directly affects context consumption and inference cost.
Representation Boundary
A token is not the same as a word. One word can span several tokens, while one token can represent several characters, bytes, punctuation marks, or a complete short word.
Related LLM Concepts
Related technical article: Artificial Intelligence: Philosophy, Theory and Practice.