Skip to main content
The transformer architecture consists of stacked blocks, each containing attention and feedforward layers with layer normalization.

Block

The Block class represents a single transformer block with pre-normalization architecture.

Class definition

Location: model.py:94-106

Parameters

GPTConfig
required
Configuration object containing model hyperparameters.

Components

LayerNorm
First layer normalization applied before the attention layer.
CausalSelfAttention
Multi-head causal self-attention mechanism.
LayerNorm
Second layer normalization applied before the MLP layer.
MLP
Feedforward network applied after attention.

Architecture

The Block implements pre-normalization with residual connections:
  1. Apply LayerNorm to input
  2. Apply attention
  3. Add residual connection
  4. Apply LayerNorm
  5. Apply MLP
  6. Add residual connection
This uses pre-normalization (LayerNorm before the sublayer) rather than post-normalization, which tends to be more stable for training.

MLP

The MLP class is a two-layer feedforward network with GELU activation.

Class definition

Location: model.py:78-92

Parameters

GPTConfig
required
Configuration object containing model hyperparameters.

Components

nn.Linear
First linear layer that expands dimensionality from n_embd to 4 * n_embd.
nn.GELU
Gaussian Error Linear Unit activation function.
nn.Linear
Second linear layer that projects back down from 4 * n_embd to n_embd.
nn.Dropout
Dropout layer applied to the output.

Architecture

The MLP follows the standard transformer feedforward network design:
The hidden dimension is 4x the embedding dimension, which is standard in transformer architectures.

LayerNorm

Custom LayerNorm implementation with optional bias parameter.

Class definition

Location: model.py:18-27

Parameters

int
required
Dimensionality to normalize over (typically config.n_embd).
bool
required
Whether to include a learnable bias parameter. PyTorch’s standard LayerNorm doesn’t support bias=False.

Components

nn.Parameter
Learnable scale parameter initialized to ones with shape (ndim,).
nn.Parameter | None
Optional learnable bias parameter initialized to zeros with shape (ndim,). Set to None if bias=False.

Why custom LayerNorm?

PyTorch’s built-in nn.LayerNorm doesn’t support disabling the bias parameter. This implementation allows you to set bias=False in the config for potentially better performance.
The epsilon value is fixed at 1e-5 for numerical stability during normalization.