Prepare the dataset
First, download and tokenize the Shakespeare dataset:train.bin and val.bin files containing the character-level tokenized text.
Training configurations
- GPU training
- CPU training
- Apple Silicon
Train a baby GPT using the default configuration:
Model architecture
The configuration inconfig/train_shakespeare_char.py defines a small Transformer:This configuration trains a 6-layer Transformer with 6 attention heads and 384 feature channels.
Expected results
- Training time: ~3 minutes on A100 GPU
- Best validation loss: 1.4697
- Output directory:
out-shakespeare-char
Configuration parameters
Key parameters fromconfig/train_shakespeare_char.py:
Sample from the model
After training completes, generate text samples:Example output
After 3 minutes of training on GPU:Advanced configuration
1
Adjust model size
Modify
n_layer, n_head, and n_embd in the config file or via command line:2
Extend training
Increase
max_iters and lr_decay_iters for longer training:3
Enable logging
Track training progress with Weights & Biases:
Next steps
Reproduce GPT-2
Train a 124M parameter model on OpenWebText
Finetuning
Finetune pretrained GPT-2 models on custom data