Skip to main content
The fastest way to get started with nanoGPT is to train a character-level model on the works of Shakespeare. This small-scale training run completes in about 3 minutes on a GPU.

Prepare the dataset

First, download and tokenize the Shakespeare dataset:
This creates train.bin and val.bin files containing the character-level tokenized text.

Training configurations

Train a baby GPT using the default configuration:

Model architecture

The configuration in config/train_shakespeare_char.py defines a small Transformer:
This configuration trains a 6-layer Transformer with 6 attention heads and 384 feature channels.

Expected results

  • Training time: ~3 minutes on A100 GPU
  • Best validation loss: 1.4697
  • Output directory: out-shakespeare-char

Configuration parameters

Key parameters from config/train_shakespeare_char.py:

Sample from the model

After training completes, generate text samples:

Example output

After 3 minutes of training on GPU:
Character-level models produce lower quality text than BPE-tokenized models. For better results, consider finetuning a pretrained GPT-2 model on this dataset.

Advanced configuration

1

Adjust model size

Modify n_layer, n_head, and n_embd in the config file or via command line:
2

Extend training

Increase max_iters and lr_decay_iters for longer training:
3

Enable logging

Track training progress with Weights & Biases:

Next steps

Reproduce GPT-2

Train a 124M parameter model on OpenWebText

Finetuning

Finetune pretrained GPT-2 models on custom data