Skip to main content

Overview

Activation functions introduce non-linearity into neural networks, enabling them to learn complex patterns.

ReLU

Rectified Linear Unit: ReLU(x) = max(0, x)

Parameters

bool
default:"False"
If True, does the operation in-place.

Example

Sigmoid

Sigmoid function: Sigmoid(x) = 1 / (1 + exp(-x)) Outputs values in the range (0, 1).

Example

Tanh

Hyperbolic tangent: Tanh(x) = (exp(x) - exp(-x)) / (exp(x) + exp(-x)) Outputs values in the range (-1, 1).

Example

LeakyReLU

Leaky Rectified Linear Unit: LeakyReLU(x) = max(0, x) + negative_slope * min(0, x)

Parameters

float
default:"0.01"
Controls the angle of the negative slope.
bool
default:"False"
If True, does the operation in-place.

Example

ELU

Exponential Linear Unit: ELU(x) = max(0, x) + min(0, alpha * (exp(x) - 1))

Parameters

float
default:"1.0"
Controls the value to which an ELU saturates for negative inputs.
bool
default:"False"
If True, does the operation in-place.

Example

SELU

Scaled Exponential Linear Unit. Self-normalizing activation function. SELU(x) = scale * (max(0, x) + min(0, alpha * (exp(x) - 1))) where scale ≈ 1.0507 and alpha ≈ 1.6733.

Example

GELU

Gaussian Error Linear Unit. Used in BERT and GPT models. GELU(x) = x * Φ(x) where Φ(x) is the cumulative distribution function of the standard normal distribution.

Parameters

bool
default:"False"
If True, use an approximation of the GELU function for faster computation.

Example

Softmax

Softmax function: Softmax(x_i) = exp(x_i) / sum_j(exp(x_j)) Converts logits to probabilities.

Parameters

int
default:"-1"
Dimension along which to apply softmax.

Example

LogSoftmax

Log Softmax: LogSoftmax(x_i) = log(exp(x_i) / sum_j(exp(x_j))) More numerically stable than log(softmax(x)).

Example

Using Activations in Models

Activation Function Comparison

Tips

Default choice: Use ReLU for most cases. It’s fast and works well for CNNs.
Dead ReLU problem: If you have dead neurons (always outputting 0), try LeakyReLU or ELU.
Transformers: Use GELU for transformer-based models (BERT, GPT).
Output layer: Use Sigmoid for binary classification, Softmax for multi-class classification.