Top-k Sampling
A decoding strategy that restricts token sampling to the k highest-probability candidates at each generation step.
Language-Model Context
Top-k sampling restricts each decoding step to the k highest-probability token candidates before sampling. This removes the long tail of very unlikely tokens, although a fixed k behaves differently when the model distribution is sharp versus highly uncertain.
Decoding Boundary
Top-k sampling is not beam search. Beam search maintains and expands sequence hypotheses according to a search objective, whereas top-k changes the candidate distribution for a stochastic sampling step.
Related LLM Concepts
Direct source: The primary paper or official specification for Top-k Sampling is linked here for verification.