Landmark 3 stops to get here · leads to 8

Softmax

A function that turns a list of scores (logits) into probabilities that are all positive and sum to 1; the standard output of classifiers and language models.

Your route here

3 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Activation Function ✓ understood

    A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.

  4. Softmax · you are here ✓ understood

Picture it

LOGITS → SOFTMAX → PROBABILITIES (SUM = 1)63.8%Alogit 2.023.5%Blogit 1.09.5%Clogit 0.13.2%Dlogit -1.0
Notice how the four logits become positive probabilities that sum to 1, and the largest logit gets a disproportionately large share.

Softmax turns a list of arbitrary scores into probabilities. Give it any real numbers, called logits, and it returns the same number of values, each between 0 and 1, adding up to 1. It’s the last step of almost every classifier, and of every language model choosing its next token.

The recipe has two moves: raise e to the power of each score, then divide each result by their total. The exponential makes every value positive; the division makes them sum to 1.

What the numbers do

In the diagram the logits are 2.0, 1.0, 0.1 and −1.0, and softmax turns them into roughly 64%, 24%, 10% and 3%. Two properties explain the shape:

  • Only the gaps matter. Add the same constant to every logit and the output doesn’t change. Implementations use this to stay numerically safe, subtracting the largest logit before exponentiating.
  • The exponential stretches gaps. Each extra unit of logit multiplies the probability by e, about 2.7. A is only one point ahead of B, yet gets 2.7 times its share. Softmax is a “soft” version of picking the maximum: the leader gets most, but never all.

Where it’s used

  • Classification. A network’s last layer produces one logit per class, softmax turns them into probabilities, and training scores them with cross-entropy. The pairing is deliberate: the log in the loss undoes the exponential, so the model keeps getting a useful learning signal even when it’s badly wrong.
  • Language models. One softmax over the entire vocabulary for every token generated.
  • Attention. Inside a transformer, softmax turns similarity scores between tokens into weights that sum to 1.

With only two classes, softmax reduces to the sigmoid.

Temperature

Dividing the logits by a number called the temperature before softmax changes how peaked the result is. Above 1 flattens the distribution; below 1 sharpens it toward the top choice. Hinton and colleagues used high temperatures to expose a large model’s “soft” preferences when distilling it into a smaller one, and the same knob, softmax temperature, controls how adventurous a language model’s sampling is.

The catch

Softmax always produces a tidy distribution, even for inputs unlike anything the model has seen. The outputs rank the options the model knows relative to each other; they aren’t a guarantee of how often the model is right. Whether 90% really means right nine times in ten is a separate question, called calibration.

Where it sits

Explore nearby

In the research

All papers →

A paper that builds on Softmax .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 6.2.2.3, Softmax Units for Multinoulli Output Distributions
  2. Hinton, Vinyals and Dean, "Distilling the Knowledge in a Neural Network" . 2015 (softmax temperature)