Quick start
Finetune GPT-2 XL (1.5B parameters) on Shakespeare in just a few minutes:1
Prepare the dataset
Download and tokenize Shakespeare using GPT-2 BPE:This creates
train.bin and val.bin using the OpenAI BPE tokenizer (unlike character-level encoding).2
Run finetuning
gpt2-xl checkpoint and finetunes it on Shakespeare.3
Sample from the model
Finetuning configuration
Theconfig/finetune_shakespeare.py file shows a complete finetuning setup:
Key differences from pretraining
Available pretrained models
You can initialize from any OpenAI GPT-2 checkpoint:
Set the model in your config:
Initialization logic
Fromtrain.py:181-189, the script downloads and loads pretrained weights:
Preparing custom datasets
Using BPE tokenization
For best results, use the same GPT-2 BPE tokenizer as pretraining:Train/validation split
Create separate files:Directory structure
Finetuning hyperparameters
Learning rate
Use a much lower learning rate than pretraining:Regularization
Add dropout to prevent overfitting on small datasets:Training duration
Finetuning requires far fewer iterations:Batch size
Adjust based on dataset size and memory:Example finetuning output
After finetuning GPT-2 XL on Shakespeare:The model generates coherent Shakespearean-style dialogue with proper formatting and character names after just ~20 iterations.
Memory management
Reduce memory usage
If you run out of GPU memory:- Use a smaller model
- Decrease batch size
- Decrease context length
Multi-GPU finetuning
Scale finetuning across multiple GPUs:Resume from checkpoint
Resume finetuning from a saved checkpoint:train.py:158-180:
Domain adaptation strategies
Continue pretraining
For large domain-specific corpora (e.g., medical, legal, code):Task-specific finetuning
For specific tasks (e.g., summarization, Q&A):Few-shot finetuning
For very small datasets (less than 1000 examples):Monitoring overfitting
Watch for signs of overfitting:Best practices
Start small
Begin with
gpt2 (124M) to iterate quickly, then scale to larger modelsMonitor validation loss
Only save checkpoints when validation loss improves
Use BPE tokenization
Match the pretraining tokenizer (GPT-2 BPE) for best results
Tune learning rate
Start at 3e-5, decrease if unstable, increase if too slow
Next steps
Sampling
Generate text from your finetuned model
Distributed training
Scale finetuning across multiple GPUs