Overview
Optimizers update model parameters based on gradients to minimize the loss function.Optimizer Base Class
Methods
zero_grad
step
add_param_group
SGD - Stochastic Gradient Descent
Parameters
Iterable[Tensor]
required
Iterable of parameters to optimize.
float
default:"0.01"
Learning rate.
float
default:"0.0"
Momentum factor. Typical values: 0.9, 0.95.
float
default:"0.0"
Weight decay (L2 penalty). Typical values: 1e-4, 1e-5.
float
default:"0.0"
Dampening for momentum.
bool
default:"False"
Enables Nesterov momentum. Requires momentum > 0 and dampening == 0.
Example
When to Use
- Simple baseline
- When you have a well-tuned learning rate schedule
- Training CNNs with batch normalization
Adam - Adaptive Moment Estimation
Parameters
Iterable[Tensor]
required
Iterable of parameters to optimize.
float
default:"0.001"
Learning rate.
Tuple[float, float]
default:"(0.9, 0.999)"
Coefficients for computing running averages of gradient and its square.
float
default:"1e-8"
Term added to the denominator for numerical stability.
float
default:"0.0"
Weight decay (L2 penalty).
bool
default:"False"
Whether to use the AMSGrad variant.
Example
When to Use
- Default choice for most tasks
- Works well with sparse gradients
- Good for transformers and NLP models
- Handles noisy gradients well
AdamW
Example
RMSprop
When to Use
- RNNs and LSTMs
- Non-stationary objectives