Skip to main content

Simple Decensoring

The most basic usage requires only a model identifier:
Heretic will:
  1. Download the model from HuggingFace (if not already cached)
  2. Detect your hardware and optimize batch size
  3. Run 200 optimization trials (default)
  4. Present results and allow you to save/upload the model

Using with Different Model Sizes

Small Models (< 8B parameters)

Small models typically fit comfortably in VRAM:

Medium Models (8B-30B parameters)

For medium models, consider using quantization to reduce VRAM usage:
4-bit quantization via bitsandbytes can reduce VRAM requirements by approximately 75% with minimal quality impact.

Large Models (> 30B parameters)

Large models require quantization and may need explicit memory management:

Understanding Progress Output

Initial Setup

When you run Heretic, you’ll see:

Batch Size Determination

Heretic automatically finds the largest batch size that fits in memory.

Optimization Trials

Each trial tests different abliteration parameters.

Results Selection

After optimization:
Choose trials with KL divergence below 1.0 for best quality. Lower refusals with higher KL divergence means more compliance but potentially degraded capabilities.

Post-Processing Options

After selecting a trial, you have several options:

Save to Local Folder

For quantized models, you’ll be asked whether to merge or save as adapter:
Merging a quantized model requires loading the full unquantized model into RAM. For a 27B model, this requires ~80GB RAM. Ensure you have sufficient memory or your system may freeze.

Upload to HuggingFace

Heretic automatically:
  • Creates or updates the repository
  • Uploads the model files
  • Updates the model card with abliteration details
  • Adds appropriate tags (heretic, uncensored, abliterated)

Chat with the Model

Test the model interactively:
This allows you to verify the model’s behavior before committing to save or upload.

Real-World Examples

Example 1: Quick Decensoring with Defaults

Best for: First-time users, small to medium models, systems with ample VRAM.

Example 2: Quantized Decensoring

Best for: Limited VRAM, faster iteration during experimentation.

Example 3: Large Model with Custom Settings

Best for: Multi-GPU systems, production deployments requiring thorough optimization.

Example 4: Local Model with Configuration File

Create config.toml:
Then run:
Best for: Repeated experiments, custom datasets, research workflows.

Example 5: Evaluation Only

Output:
Best for: Comparing different decensored variants, benchmarking.

Tips for Success

Start small: Test Heretic on a small model first (< 8B parameters) to understand the workflow before moving to larger models.
Monitor KL divergence: Values below 0.5 typically indicate minimal capability loss. Values above 1.0 may indicate significant degradation.
Use chat testing: Always test a trial with the interactive chat before saving to ensure the model behaves as expected.
More trials = better results: The default 200 trials is a good starting point, but increasing to 300-500 trials can sometimes find better parameter combinations.
CTRL+C during optimization will gracefully stop the current trial and allow you to view results. The checkpoint is saved automatically.