Logit
In statistics, the log-odds transform of a probability; in machine learning, the term is also used for model scores before probability normalization.
The term logit is used in two closely related but not identical ways.
Statistical Definition
For a probability p:
logit(p) = ln(p / (1 - p))p/(1-p) is the odds. The logit maps a probability in (0,1) to the real-number line and is the quantity modeled by the linear predictor in logistic regression.
Machine-Learning Usage
In neural networks, logits commonly refers to raw output scores before softmax or sigmoid normalization. Individual logits do not need to lie in [0,1] and do not need to sum to one.
In multiclass classification, softmax transforms a vector of logits into normalized weights. In language models, token-selection methods such as temperature and top-k/top-p operate on or derive from these pre-normalization scores.
Important Distinction
Calling a neural-network output a logit does not imply that the value was explicitly computed as ln(p/(1-p)). The statistical origin of the term and its practical deep-learning usage should be interpreted in context.
Related: Softmax, Temperature, Artificial Intelligence and Neural Networks.
Inverse Relationship with the Logistic Function
The statistical logit and logistic sigmoid are inverses. If x = logit(p), then:
p = 1 / (1 + exp(-x))In binary logistic regression the linear predictor can therefore live on the real line while the sigmoid maps it into (0,1). The common neural-network use of the plural "logits" does not require each score to have been explicitly constructed by this inverse transformation; many libraries use the term for a vector of pre-normalization scores.
This distinction matters in loss APIs as well. "Cross entropy with logits" operations can combine sigmoid/softmax and logarithms in a numerically stable implementation; feeding already normalized probabilities as though they were logits changes the computation. Related: Softmax.
Relationship to Log-Odds
For a probability p, the logit is the log-odds transform log(p / (1-p)). It maps the probability interval (0,1) to the real line. Logistic regression models a linear predictor in this space, while the sigmoid function maps the logit back to a probability.