Skip to main content

Overview

Loss functions measure the discrepancy between predicted and target values, guiding the optimization process during training.

MSELoss

Mean Squared Error loss: MSE = mean((prediction - target)^2)

Parameters

str
default:"mean"
Specifies the reduction to apply: ‘none’, ‘mean’, or ‘sum’.

Example

Use Cases

  • Regression tasks
  • Image reconstruction
  • Autoencoders

L1Loss

Mean Absolute Error loss: L1 = mean(|prediction - target|)

Example

Use Cases

  • Robust regression (less sensitive to outliers than MSE)
  • Image-to-image translation

CrossEntropyLoss

Combines LogSoftmax and NLLLoss. Used for multi-class classification.

Parameters

Optional[Tensor]
Manual rescaling weight for each class. Shape: (num_classes,)
int
default:"-100"
Specifies a target value that is ignored and does not contribute to the gradient.
str
default:"mean"
Specifies the reduction to apply: ‘none’, ‘mean’, or ‘sum’.

Example

Use Cases

  • Multi-class classification
  • Image classification
  • Text classification
  • Semantic segmentation

BCELoss

Binary Cross Entropy loss. Used for binary classification with sigmoid outputs.

Parameters

Optional[Tensor]
Manual rescaling weight for the loss of each batch element.
str
default:"mean"
Specifies the reduction to apply: ‘none’, ‘mean’, or ‘sum’.

Example

Use Cases

  • Binary classification
  • Multi-label classification (independent binary decisions)

BCEWithLogitsLoss

Combines Sigmoid and BCELoss. More numerically stable than using BCELoss separately.

Parameters

Optional[Tensor]
Manual rescaling weight for the loss of each batch element.
Optional[Tensor]
Weight for positive examples. Useful for imbalanced datasets.
str
default:"mean"
Specifies the reduction to apply: ‘none’, ‘mean’, or ‘sum’.

Example

Use Cases

  • Binary classification (preferred over BCELoss for numerical stability)
  • Multi-label classification
  • Object detection

Training Example

Loss Function Selection Guide

Custom Loss Functions

Tips

Classification: Always use CrossEntropyLoss or BCEWithLogitsLoss (they include the activation). Don’t apply Softmax/Sigmoid before the loss.
Imbalanced datasets: Use class weights or pos_weight to handle class imbalance.
Regression: Start with MSELoss. If you have outliers, try L1Loss.
Numerical stability: Use the “WithLogits” versions of losses when possible for better numerical stability.