Skip to main content
Heretic includes sophisticated hardware detection and optimization features that automatically tune performance for your system. This guide covers both automatic and manual optimization techniques.

Automatic Batch Size Detection

By default, Heretic automatically determines the optimal batch size for your hardware:
config.toml
When set to 0, Heretic will:
1

Benchmark Different Batch Sizes

Starting from batch size 1, doubles the batch size and tests performance (2, 4, 8, 16, …)
2

Measure Throughput

For each batch size, measures tokens/second after a warmup run
3

Find Optimal Size

Selects the batch size that achieves the highest throughput before OOM
4

Use Throughout Session

Applies the chosen batch size for all subsequent operations

How It Works

The automatic detection process:
Automatic detection typically adds 1-3 minutes to startup time but ensures optimal performance throughout the entire run.

Manual Batch Size Tuning

For more control, you can set the batch size manually:

When to Use Manual Tuning

Reproducibility

Ensure consistent behavior across multiple runs

Shared Resources

Control memory usage on multi-user systems

Known Configuration

Skip detection when you know the optimal value

Debugging

Isolate issues by fixing batch size

Maximum Batch Size Limit

Control the upper bound for automatic detection:
config.toml
Lower this value if automatic detection causes OOM errors or takes too long.

Multi-GPU Configuration

Heretic automatically detects and utilizes multiple GPUs:

Device Map Strategies

Control how the model is distributed across devices:

Per-Device Memory Limits

Set maximum memory allocation per device:
config.toml
This is useful for:
  • Sharing GPUs with other processes
  • Preventing a single model from consuming all VRAM
  • Forcing CPU offloading for memory-intensive layers
When using max_memory, make sure the total allocated memory is sufficient for your model. Too restrictive limits will cause loading failures.

Performance on Different Hardware

From the README, here are typical processing times:

RTX 3090 Performance

Model: Llama-3.1-8B-Instruct
Configuration: Default settings (200 trials)
Duration: ~45 minutes
This includes:
  • Model loading
  • Batch size detection
  • 200 optimization trials
  • Evaluation
Smaller models (4B-7B) typically complete in 20-40 minutes, while larger models (70B+) may take several hours even with quantization.

Duration Estimates by Model Size

These are rough estimates for 200 trials with default settings. Actual time varies based on model architecture and prompt datasets.

Advanced Memory Optimization

Expandable Segments

Heretic automatically enables PyTorch expandable segments to reduce memory fragmentation:
This is particularly beneficial for multi-GPU setups.

TorchDynamo Cache

The compilation cache is increased during batch size detection:
This prevents errors from excessive recompilation during the batch size search.

Supported Accelerators

Heretic supports a wide range of hardware:
Best supported, recommended for most users.

Optimization Best Practices

1

Start with Defaults

Let automatic batch size detection find the optimal setting
2

Enable Quantization for Large Models

Use 4-bit quantization for models >13B on consumer GPUs
3

Monitor Memory Usage

Watch VRAM during processing. If near capacity, reduce max_batch_size
4

Tune for Your Workload

If running many short sessions, fix batch size to skip detection

Configuration Examples

Single High-End GPU

config.toml

Consumer GPU with Limited VRAM

config.toml

Multi-GPU Server

config.toml

CPU Offloading

config.toml

Troubleshooting

Out of Memory (OOM)

Symptoms: RuntimeError: CUDA out of memory Solutions:
1

Enable Quantization

2

Reduce Max Batch Size

3

Set Memory Limits

4

Use CPU Offloading

Slow Batch Size Detection

Symptoms: Detection takes >5 minutes Solutions:
  • Lower max_batch_size to reduce search space
  • Set explicit batch_size based on previous runs
  • Use a smaller model for initial testing

Suboptimal Performance

Symptoms: Low tokens/second during processing Solutions:
  • Verify automatic detection chose a reasonable batch size
  • Check if CPU offloading is active (slow)
  • Ensure model fits entirely in VRAM
  • Monitor GPU utilization with nvidia-smi