Skip to main content
The fastest way to get started with nanoGPT is to train a character-level GPT on the works of Shakespeare. This tutorial will walk you through the complete process.

Overview

You’ll learn how to:
  • Prepare the Shakespeare dataset for training
  • Train a small GPT model from scratch
  • Generate text samples from your trained model
  • Adjust hyperparameters for different hardware
This quickstart uses character-level modeling (not BPE tokens) for simplicity. The entire dataset is just 1MB and training takes only a few minutes.

Prepare the dataset

First, download and prepare the Shakespeare dataset:
1

Run the data preparation script

The preparation script downloads the tiny Shakespeare dataset and converts it to binary format:
This script:
  • Downloads the complete works of Shakespeare (~1MB text file)
  • Creates a character-to-integer mapping
  • Splits the data into training (90%) and validation (10%) sets
  • Saves train.bin and val.bin in data/shakespeare_char/
2

Verify the output

You should see output similar to:
The dataset contains 65 unique characters including uppercase, lowercase, and punctuation.

Train the model

Now you can train a GPT model on this data. The training approach depends on your hardware.
If you have a GPU, you can train a small but capable model using the provided config file:

Model architecture

The config file config/train_shakespeare_char.py defines a “baby GPT” with:

Training progress

On an A100 GPU, this takes about 3 minutes and achieves a validation loss around 1.47:
Model checkpoints are saved to out-shakespeare-char/ directory.
On Apple Silicon Macs, add --device=mps to use the GPU for 2-3x speedup over CPU.

Generate text samples

Once training completes, generate Shakespeare-style text from your model:
If you trained on CPU, add the --device=cpu flag:

Example output

With the GPU-trained model (validation loss 1.47), you might see:
Not bad for a character-level model trained in just 3 minutes!

Customize generation

You can control the generation process with additional parameters:
string
default:"\\n"
The prompt to start generation. Can also load from a file with FILE:prompt.txt
integer
default:"10"
Number of independent samples to generate
integer
default:"500"
Maximum number of tokens to generate per sample
float
default:"0.8"
Sampling temperature. Lower values (0.6-0.8) are more conservative, higher values (1.0+) more creative
integer
default:"200"
Only sample from the top k most likely tokens at each step

Understanding the training loop

Let’s examine what happens during training. The core training loop in train.py follows this pattern:
train.py

Next steps

Finetune GPT-2

Learn how to finetune pretrained GPT-2 models on your own data for better results

Training configuration

Explore all available hyperparameters and configuration options

Character-level training

Deep dive into character-level model training and configuration

Distributed training

Scale up to multi-GPU training with PyTorch DDP

Common issues

  • Ensure you’re using a GPU with --device=cuda (or --device=mps on Mac)
  • Verify PyTorch 2.0+ is installed to enable torch.compile()
  • Check that compilation is enabled (don’t use --compile=False on GPU)
Reduce memory usage by:
  • Verify the data preparation completed successfully
  • Check that train.bin and val.bin exist in data/shakespeare_char/
  • Try increasing the learning rate or reducing dropout
  • The model may need more training iterations
  • Try lowering the temperature: --temperature=0.7
  • Check that you’re loading the correct checkpoint with --out_dir