Byte Pair Encoding
A subword tokenization method that repeatedly merges frequent adjacent symbol pairs to build a compact vocabulary from raw text units.
Language-Model Context
Byte Pair Encoding builds a subword vocabulary by repeatedly merging frequent adjacent symbols. In NLP tokenizers, BPE and related schemes let a bounded vocabulary represent rare words as smaller units; byte-level variants also reduce dependence on a fixed character inventory.
Terminology Boundary
The original BPE compression algorithm and modern language-model tokenizers share a merging idea but not necessarily the same preprocessing, objective, byte handling, or implementation details.
Related LLM Concepts
- Tokenization
- Vocabulary
- Subword
- Context Window
Two Meanings Commonly Collapsed into BPE
The name BPE originates in a data-compression algorithm; its NLP use adapts the repeated-pair merging idea to open-vocabulary text. Sennrich, Haddow, and Birch applied this idea to neural machine translation by representing rare words as sequences of subword units. Modern "byte-level BPE" tokenizers are not necessarily identical to that early formulation because preprocessing, byte mapping, normalization, and special-token rules can differ.
Merge ranks and vocabulary are part of the model/tokenizer contract. The same visible text can map to a different token-ID sequence under another tokenizer version, changing context consumption and model input even when model weights are unchanged. The direct source for the influential NLP adaptation is Sennrich et al..