obliteratus

OBLITERATUS: abliterate LLM refusals (diff-in-means).

  • Abliteration
  • Uncensoring
  • Refusal-Removal
  • LLM
  • Weight-Projection
  • SVD
  • Mechanistic-Interpretability
  • HuggingFace
  • Model-Surgery

Declared platforms: linux · macos

Install
npx skills add 'https://github.com/NousResearch/hermes-agent/tree/main/optional-skills/mlops/obliteratus'
Download bundle ↓
main · 24fd22bScanned 2026-09-15

Contributors

GitHub-linked commit authors for this SKILL.md at the saved revision. Co-authors and history before file renames are not included.

File history ↗
View on GitHub
← Back to SKILL.md
# OBLITERATUS Analysis Study Config# Usage: obliteratus run this-file.yaml --preset jailbreak## Run analysis modules to understand refusal geometry BEFORE abliterating.# Useful for research or when you want to understand what you're removing. # Model to analyzemodel:  name: "meta-llama/Llama-3.1-8B-Instruct"  dtype: "bfloat16"  quantization: "4bit"       # Saves VRAM for analysis  device: "auto" # Study configurationstudy:  # Available presets: quick, full, attention, jailbreak, guardrail, knowledge  preset: "jailbreak"   # Or specify individual strategies:  # strategies:  #   - layer_removal  #   - head_pruning  #   - ffn_ablation  #   - embedding_ablation # Analysis modules to run (subset of the 27 available)analysis:  - alignment_imprint        # Detect DPO/RLHF/CAI/SFT training method  - concept_geometry          # Map refusal cone geometry  - logit_lens               # Find which layer decides to refuse  - anti_ouroboros            # Detect self-repair tendency  - cross_layer              # Cross-layer alignment clustering  - causal_tracing           # Causal necessity of components  - residual_stream          # Attention vs MLP contribution # Outputoutput:  directory: "./analysis-results"  save_plots: true           # Generate matplotlib visualizations  save_report: true          # Generate markdown report 
Referenced from SKILL.md