Skip to main content
This page documents all command-line options available in Heretic. Options can also be set via environment variables (with HERETIC_ prefix) or in a config.toml file.

Model Loading

string
required
HuggingFace model ID or path to model on disk.Examples:
If provided as the last argument without --model flag, it will be automatically recognized as the model parameter.
string
default:"null"
Model ID or path to evaluate against the main model instead of performing abliteration.Example:
This compares the refusals and KL divergence of the evaluated model relative to the base model.
list[string]
List of PyTorch dtypes to try when loading model tensors. If loading with a dtype fails, the next dtype in the list will be tried.Example:
string
default:"none"
Quantization method to use when loading the model.Options:
  • none: No quantization (full precision)
  • bnb_4bit: 4-bit quantization using bitsandbytes
Example:
4-bit quantization can reduce VRAM requirements by ~75% with minimal quality impact, enabling processing of larger models on consumer GPUs.
string | dict
default:"auto"
Device map to pass to Accelerate when loading the model.Examples:
dict
default:"null"
Maximum memory to allocate per device. Useful for multi-GPU setups or when sharing GPU with other processes.Example (requires config file):
boolean
default:"null"
Whether to trust remote code when loading the model. Some models require custom code that must be explicitly trusted.Example:
Only enable for models from trusted sources, as remote code can execute arbitrary Python.

Performance & Optimization

integer
default:"0"
Number of input sequences to process in parallel. Set to 0 for automatic determination.Example:
Automatic batch size detection (default) is recommended. It benchmarks your hardware to find the optimal throughput.
integer
default:"128"
Maximum batch size to try when automatically determining the optimal batch size.Example:
integer
default:"100"
Maximum number of tokens to generate for each response during evaluation.Example:
Longer responses take more time but may improve refusal detection accuracy.

Optimization Parameters

integer
default:"200"
Number of abliteration trials to run during optimization.Example:
More trials increase the chance of finding better parameters but take longer. 200 is a good balance for most use cases.
integer
default:"60"
Number of trials that use random sampling for exploration before switching to TPE (Tree-structured Parzen Estimator) optimization.Example:
Higher values improve initial exploration but delay focused optimization.
string
default:"checkpoints"
Directory to save and load study progress to/from.Example:
Checkpoints enable resuming interrupted runs and reviewing previous results.
float
default:"1.0"
Assumed “typical” value of the Kullback-Leibler divergence for abliterated models. Used to ensure balanced co-optimization of KL divergence and refusal count.Example:
float
default:"0.01"
KL divergence target threshold. Below this value, optimization focuses on refusal count. This prevents exploring parameters that have no effect.Example:

Abliteration Method

boolean
default:"false"
Whether to adjust refusal directions so that only the component orthogonal to the “good” direction is subtracted during abliteration.Example:
Implements projected abliteration. May improve capability retention in some models.
string
default:"none"
How to apply row normalization of the weights.Options:
  • none: No normalization
  • pre: Compute LoRA adapter relative to row-normalized weights
  • full: Like pre, but renormalizes to preserve original row magnitudes
Example:
Implements norm-preserving abliteration.
integer
default:"3"
Rank of the LoRA adapter when full row normalization is used. Higher ranks provide better approximation but increase file size and evaluation time.Example:
float
default:"1.0"
Symmetric winsorization quantile for per-prompt, per-layer residual vectors (between 0 and 1). Disabled by default (1.0).Example:
This clamps residual magnitudes to the specified quantile, taming “massive activations” in some models. Value of 0.95 means components are clamped to the 95th percentile magnitude.

Evaluation & Datasets

list[string]
default:"[see config.default.toml]"
Strings whose presence in a response (case-insensitive) identifies it as a refusal.Default includes: sorry, i cannot, as an ai, harmful, unethical, etc.Example (config file):
string
default:"You are a helpful assistant."
System prompt to use when prompting the model.Example:

Dataset Configuration

Heretic uses four datasets for training and evaluation. Each dataset can be configured with these sub-parameters:
object
Dataset of prompts that tend to NOT result in refusals (used for calculating refusal directions).Default:
object
Dataset of prompts that tend to result in refusals (used for calculating refusal directions).Default:
object
Dataset of harmless prompts used for evaluating model performance (KL divergence measurement).Default:
object
Dataset of harmful prompts used for evaluating model performance (refusal counting).Default:
Custom Dataset Example:
Datasets can be HuggingFace dataset IDs or local file paths. The split parameter uses HuggingFace slice notation.

Research Features

Research features require the research extra: pip install heretic-llm[research]
boolean
default:"false"
Whether to print prompt/response pairs when counting refusals.Example:
Useful for debugging refusal detection or understanding model behavior.
boolean
default:"false"
Whether to print detailed information about residuals and refusal directions.Example:
Outputs a detailed table with per-layer metrics including:
  • Cosine similarities between good/bad/refusal directions
  • L2 norms of direction vectors
  • Silhouette coefficients for clustering quality
boolean
default:"false"
Whether to generate plots showing PaCMAP projections of residual vectors.Example:
Generates:
  • PNG image for each transformer layer
  • Animated GIF showing transformation between layers
PaCMAP projection is CPU-intensive and can take over an hour for large models.
string
default:"plots"
Base path to save plots of residual vectors.Example:
string
Title placed above plots of residual vectors.Example:
string
default:"dark_background"
Matplotlib style sheet to use for plots of residual vectors.Example:
See Matplotlib style sheets for available options.

Configuration File Example

Instead of long command lines, create config.toml in your working directory:
Then run:
All settings from the config file will be applied automatically.

Environment Variables

Any option can be set via environment variable with the HERETIC_ prefix:
Environment variables are useful for containerized deployments or when you want to override config file settings temporarily.

Help Command

For a quick reference of all options:
This displays a summary of all available command-line flags with their descriptions and default values.