Top-p Sampling
Top-p Sampling — A nucleus-sampling strategy that draws from the smallest token set whose cumulative probability reaches a chosen threshold p.
Language-Model Context
Top-p, or nucleus, sampling forms the smallest candidate set whose cumulative probability reaches a chosen threshold and samples from that set. The number of eligible tokens therefore expands or contracts with the uncertainty of the current distribution.
Decoding Boundary
The value of p is not a quality score and does not imply the same behavior across different models. When temperature and top-p are combined, their interaction should be evaluated rather than treated as two independent knobs.
Related LLM Concepts
Variable Candidate Set
Top-p does not keep a fixed number of tokens. After token probabilities are ordered from largest to smallest, it retains the smallest prefix whose cumulative probability is at least p. The same threshold can therefore leave only a few candidates when the distribution is sharp and many candidates when it is flat. This is the central distinction from Top-k Sampling: top-k fixes candidate count, while top-p limits retained probability mass.
The value of p is not an accuracy or confidence score. Lowering it truncates the tail more aggressively but does not establish which continuation is correct. Temperature can reshape the distribution before sampling, so temperature and top-p should not be treated as independent quality controls. The dynamic nucleus formulation is described directly by Holtzman et al.: original paper.