Softmax
An exponential normalization function that converts a real-valued score vector into positive weights that sum to one.
Softmax converts a real-valued score vector into positive normalized values that sum to 1.
For a vector z:
softmax(z_i) = exp(z_i) / Σ_j exp(z_j)Numerical Stability
Exponentials can overflow for large inputs. Implementations therefore commonly evaluate:
exp(z_i - max(z))Subtracting the same constant from every score leaves the normalized result unchanged while improving numerical stability.
Relationship with Logits
Scores before normalization are commonly called logits. Softmax converts their relative differences into normalized weights.
Probability Is Not Automatically Calibration
Values summing to 1 do not prove that a model's confidence is a calibrated real-world probability. Calibration must be evaluated separately.
Temperature scaling can make the distribution sharper or flatter, and attention mechanisms also use softmax to produce normalized weights.
Why a Constant Shift Leaves It Unchanged
Subtracting the maximum logit for numerical stability is not merely an implementation trick; it preserves the result algebraically. If the same constant c is subtracted from every component:
exp(z_i-c) / Σ exp(z_j-c)
= exp(z_i)exp(-c) / [exp(-c)Σ exp(z_j)]
= softmax(z_i)the common exp(-c) factor cancels. Choosing c = max(z) makes the largest exponential exp(0)=1, reducing overflow risk. Softmax still produces normalized relative weights, not automatic calibration. A model output of 0.9 does not by itself prove that comparable cases are correct exactly 90 percent of the time. Related concepts are Logit and Temperature.