Skip to main content
Finetuning a pretrained GPT-2 model is identical to training from scratch, except you initialize from a checkpoint and use a smaller learning rate. This allows you to adapt large models to your specific domain with minimal compute.

Quick start

Finetune GPT-2 XL (1.5B parameters) on Shakespeare in just a few minutes:
1

Prepare the dataset

Download and tokenize Shakespeare using GPT-2 BPE:
This creates train.bin and val.bin using the OpenAI BPE tokenizer (unlike character-level encoding).
2

Run finetuning

This loads the pretrained gpt2-xl checkpoint and finetunes it on Shakespeare.
3

Sample from the model

Finetuning configuration

The config/finetune_shakespeare.py file shows a complete finetuning setup:

Key differences from pretraining

Available pretrained models

You can initialize from any OpenAI GPT-2 checkpoint: Set the model in your config:

Initialization logic

From train.py:181-189, the script downloads and loads pretrained weights:
This automatically downloads weights from Hugging Face on first use.

Preparing custom datasets

Using BPE tokenization

For best results, use the same GPT-2 BPE tokenizer as pretraining:

Train/validation split

Create separate files:

Directory structure

Finetuning hyperparameters

Learning rate

Use a much lower learning rate than pretraining:
Too high a learning rate can destroy pretrained knowledge. Start conservatively and increase if training is too slow.

Regularization

Add dropout to prevent overfitting on small datasets:
From the override logic:

Training duration

Finetuning requires far fewer iterations:

Batch size

Adjust based on dataset size and memory:

Example finetuning output

After finetuning GPT-2 XL on Shakespeare:
The model generates coherent Shakespearean-style dialogue with proper formatting and character names after just ~20 iterations.

Memory management

Reduce memory usage

If you run out of GPU memory:

Multi-GPU finetuning

Scale finetuning across multiple GPUs:
The same DDP logic applies as in pretraining. See Distributed Training for details.

Resume from checkpoint

Resume finetuning from a saved checkpoint:
Run:
From train.py:158-180:

Domain adaptation strategies

Continue pretraining

For large domain-specific corpora (e.g., medical, legal, code):

Task-specific finetuning

For specific tasks (e.g., summarization, Q&A):

Few-shot finetuning

For very small datasets (less than 1000 examples):

Monitoring overfitting

Watch for signs of overfitting:
If validation loss increases while training loss decreases, you’re overfitting. Reduce max_iters, increase dropout, or use a smaller model.

Best practices

Start small

Begin with gpt2 (124M) to iterate quickly, then scale to larger models

Monitor validation loss

Only save checkpoints when validation loss improves

Use BPE tokenization

Match the pretraining tokenizer (GPT-2 BPE) for best results

Tune learning rate

Start at 3e-5, decrease if unstable, increase if too slow

Next steps

Sampling

Generate text from your finetuned model

Distributed training

Scale finetuning across multiple GPUs