GPTConfig is a dataclass that defines all configuration parameters for the GPT model architecture. It controls the model size, attention mechanism, and regularization settings.
Class definition
model.py:108-116
Parameters
int
default:"1024"
Maximum sequence length that the model can handle. This determines the size of the positional embeddings and the causal attention mask.
int
default:"50304"
Size of the vocabulary. The default value of 50304 is GPT-2’s vocab_size of 50257 padded up to the nearest multiple of 64 for efficiency.
int
default:"12"
Number of transformer blocks in the model. Standard GPT-2 uses 12 layers.
int
default:"12"
Number of attention heads in each transformer block. Must evenly divide
n_embd.int
default:"768"
Dimensionality of the embeddings and hidden states throughout the model. Standard GPT-2 uses 768.
float
default:"0.0"
Dropout probability applied to attention weights, residual connections, and embeddings. Set to 0.0 to disable dropout.
bool
default:"True"
Whether to include bias terms in Linear layers and LayerNorms. Setting to
False can make the model slightly faster and may improve performance.Usage
Create a custom configuration
Use pretrained configurations
When loading pretrained models, the configuration is automatically set based on the model type:The
n_head parameter must evenly divide n_embd since each head operates on a portion of the embedding dimension (n_embd // n_head).