Standard 4 stops to get here · leads to 1
Weight Decay
A regularization technique that shrinks weights toward zero during optimization. Equivalent to L2 regularization in standard SGD, but differs when using adaptive optimizers like Adam.
Your route here
4 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training Data ✓ understood
The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.
- Overfitting ✓ understood
When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.
- Regularization ✓ understood
Techniques to prevent overfitting by adding constraints or penalties to the model (L1, L2, dropout, early stopping).
- Weight Decay · you are here ✓ understood
Where it sits
Explore nearby
Training L2 Regularization Adding the sum of squared weights to the loss function, penalizing large weights and improving generalization. Training AdamW Adam with decoupled weight decay, providing better regularization and often superior performance. Training L1 Regularization Adding the sum of absolute weights to the loss function, promoting sparsity and feature selection. Neural Networks Dropout A regularization technique that randomly deactivates neurons during training to prevent overfitting and improve generalization.