Overview
Activation functions introduce non-linearity into neural networks, enabling them to learn complex patterns.ReLU
ReLU(x) = max(0, x)
Parameters
bool
default:"False"
If True, does the operation in-place.
Example
Sigmoid
Sigmoid(x) = 1 / (1 + exp(-x))
Outputs values in the range (0, 1).
Example
Tanh
Tanh(x) = (exp(x) - exp(-x)) / (exp(x) + exp(-x))
Outputs values in the range (-1, 1).
Example
LeakyReLU
LeakyReLU(x) = max(0, x) + negative_slope * min(0, x)
Parameters
float
default:"0.01"
Controls the angle of the negative slope.
bool
default:"False"
If True, does the operation in-place.
Example
ELU
ELU(x) = max(0, x) + min(0, alpha * (exp(x) - 1))
Parameters
float
default:"1.0"
Controls the value to which an ELU saturates for negative inputs.
bool
default:"False"
If True, does the operation in-place.
Example
SELU
SELU(x) = scale * (max(0, x) + min(0, alpha * (exp(x) - 1)))
where scale ≈ 1.0507 and alpha ≈ 1.6733.
Example
GELU
GELU(x) = x * Φ(x)
where Φ(x) is the cumulative distribution function of the standard normal distribution.
Parameters
bool
default:"False"
If True, use an approximation of the GELU function for faster computation.
Example
Softmax
Softmax(x_i) = exp(x_i) / sum_j(exp(x_j))
Converts logits to probabilities.
Parameters
int
default:"-1"
Dimension along which to apply softmax.
Example
LogSoftmax
LogSoftmax(x_i) = log(exp(x_i) / sum_j(exp(x_j)))
More numerically stable than log(softmax(x)).