Simple Decensoring
The most basic usage requires only a model identifier:
Heretic will:
- Download the model from HuggingFace (if not already cached)
- Detect your hardware and optimize batch size
- Run 200 optimization trials (default)
- Present results and allow you to save/upload the model
Using with Different Model Sizes
Small Models (< 8B parameters)
Small models typically fit comfortably in VRAM:
Medium Models (8B-30B parameters)
For medium models, consider using quantization to reduce VRAM usage:
4-bit quantization via bitsandbytes can reduce VRAM requirements by approximately 75% with minimal quality impact.
Large Models (> 30B parameters)
Large models require quantization and may need explicit memory management:
Understanding Progress Output
Initial Setup
When you run Heretic, you’ll see:
Batch Size Determination
Heretic automatically finds the largest batch size that fits in memory.
Optimization Trials
Each trial tests different abliteration parameters.
Results Selection
After optimization:
Choose trials with KL divergence below 1.0 for best quality. Lower refusals with higher KL divergence means more compliance but potentially degraded capabilities.
Post-Processing Options
After selecting a trial, you have several options:
Save to Local Folder
For quantized models, you’ll be asked whether to merge or save as adapter:
Merging a quantized model requires loading the full unquantized model into RAM. For a 27B model, this requires ~80GB RAM. Ensure you have sufficient memory or your system may freeze.
Upload to HuggingFace
Heretic automatically:
- Creates or updates the repository
- Uploads the model files
- Updates the model card with abliteration details
- Adds appropriate tags (
heretic, uncensored, abliterated)
Chat with the Model
Test the model interactively:
This allows you to verify the model’s behavior before committing to save or upload.
Real-World Examples
Example 1: Quick Decensoring with Defaults
Best for: First-time users, small to medium models, systems with ample VRAM.
Example 2: Quantized Decensoring
Best for: Limited VRAM, faster iteration during experimentation.
Example 3: Large Model with Custom Settings
Best for: Multi-GPU systems, production deployments requiring thorough optimization.
Example 4: Local Model with Configuration File
Create config.toml:
Then run:
Best for: Repeated experiments, custom datasets, research workflows.
Example 5: Evaluation Only
Output:
Best for: Comparing different decensored variants, benchmarking.
Tips for Success
Start small: Test Heretic on a small model first (< 8B parameters) to understand the workflow before moving to larger models.
Monitor KL divergence: Values below 0.5 typically indicate minimal capability loss. Values above 1.0 may indicate significant degradation.
Use chat testing: Always test a trial with the interactive chat before saving to ensure the model behaves as expected.
More trials = better results: The default 200 trials is a good starting point, but increasing to 300-500 trials can sometimes find better parameter combinations.
CTRL+C during optimization will gracefully stop the current trial and allow you to view results. The checkpoint is saved automatically.