Skip to main content

Overview

Optimizers update model parameters based on gradients to minimize the loss function.

Optimizer Base Class

Base class for all optimizers.

Methods

zero_grad

Reset the gradients of all optimized tensors. Should be called before computing gradients for a new batch.

step

Update the parameters based on the current gradients. Should be called after computing gradients.

add_param_group

Add a parameter group to the optimizer with different hyperparameters.

SGD - Stochastic Gradient Descent

Stochastic gradient descent with optional momentum.

Parameters

Iterable[Tensor]
required
Iterable of parameters to optimize.
float
default:"0.01"
Learning rate.
float
default:"0.0"
Momentum factor. Typical values: 0.9, 0.95.
float
default:"0.0"
Weight decay (L2 penalty). Typical values: 1e-4, 1e-5.
float
default:"0.0"
Dampening for momentum.
bool
default:"False"
Enables Nesterov momentum. Requires momentum > 0 and dampening == 0.

Example

When to Use

  • Simple baseline
  • When you have a well-tuned learning rate schedule
  • Training CNNs with batch normalization

Adam - Adaptive Moment Estimation

Adaptive learning rate optimization algorithm.

Parameters

Iterable[Tensor]
required
Iterable of parameters to optimize.
float
default:"0.001"
Learning rate.
Tuple[float, float]
default:"(0.9, 0.999)"
Coefficients for computing running averages of gradient and its square.
float
default:"1e-8"
Term added to the denominator for numerical stability.
float
default:"0.0"
Weight decay (L2 penalty).
bool
default:"False"
Whether to use the AMSGrad variant.

Example

When to Use

  • Default choice for most tasks
  • Works well with sparse gradients
  • Good for transformers and NLP models
  • Handles noisy gradients well

AdamW

Adam with decoupled weight decay. Often performs better than Adam with weight decay.

Example

RMSprop

Root Mean Square Propagation optimizer.

When to Use

  • RNNs and LSTMs
  • Non-stationary objectives

Advanced Usage

Different Learning Rates for Different Layers

Gradient Clipping

Learning Rate Scheduling

Mixed Precision Training

Complete Training Example

Optimizer Comparison

Tips

Start with Adam: It’s a good default choice for most tasks with lr=0.001.
Use AdamW for transformers: It often performs better than Adam for large language models.
SGD with momentum: Can achieve better final performance than Adam but requires more tuning.
Learning rate: Most important hyperparameter. Use learning rate scheduling for better results.
Weight decay: Typical values are 1e-4 or 1e-5 for regularization.