Logit

Turkish equivalent: LogitDomain: Large Language Models

In statistics, the log-odds transform of a probability; in machine learning, the term is also used for model scores before probability normalization.

The term logit is used in two closely related but not identical ways.

Statistical Definition

For a probability p:

logit(p) = ln(p / (1 - p))

p/(1-p) is the odds. The logit maps a probability in (0,1) to the real-number line and is the quantity modeled by the linear predictor in logistic regression.

Machine-Learning Usage

In neural networks, logits commonly refers to raw output scores before softmax or sigmoid normalization. Individual logits do not need to lie in [0,1] and do not need to sum to one.

In multiclass classification, softmax transforms a vector of logits into normalized weights. In language models, token-selection methods such as temperature and top-k/top-p operate on or derive from these pre-normalization scores.

Important Distinction

Calling a neural-network output a logit does not imply that the value was explicitly computed as ln(p/(1-p)). The statistical origin of the term and its practical deep-learning usage should be interpreted in context.

Related: Softmax, Temperature, Artificial Intelligence and Neural Networks.

Inverse Relationship with the Logistic Function

The statistical logit and logistic sigmoid are inverses. If x = logit(p), then:

p = 1 / (1 + exp(-x))

In binary logistic regression the linear predictor can therefore live on the real line while the sigmoid maps it into (0,1). The common neural-network use of the plural "logits" does not require each score to have been explicitly constructed by this inverse transformation; many libraries use the term for a vector of pre-normalization scores.

This distinction matters in loss APIs as well. "Cross entropy with logits" operations can combine sigmoid/softmax and logarithms in a numerically stable implementation; feeding already normalized probabilities as though they were logits changes the computation. Related: Softmax.

Relationship to Log-Odds

For a probability p, the logit is the log-odds transform log(p / (1-p)). It maps the probability interval (0,1) to the real line. Logistic regression models a linear predictor in this space, while the sigmoid function maps the logit back to a probability.