Reference 4 stops to get here · leads to 1
Softmax Temperature
A parameter controlling the smoothness of probability distributions in softmax - lower makes peaks sharper, higher makes it more uniform.
Your route here
4 stops · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Activation Function ✓ understood
A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.
- Softmax ✓ understood
A function that turns a list of scores (logits) into probabilities that are all positive and sum to 1; the standard output of classifiers and language models.
- Softmax Temperature · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs Temperature A sampling parameter controlling randomness in generation - lower values make output more deterministic, higher more creative. Training Distillation Temperature A hyperparameter in knowledge distillation controlling how soft the teacher's outputs are. Training Knowledge Distillation Training a smaller 'student' model to mimic a larger 'teacher' model, transferring knowledge while reducing size. Evaluation Calibration Ensuring predicted probabilities accurately reflect true likelihood of outcomes.