Model: ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated Source: Original Platform
license, language, tags, base_model, pipeline_tag
| license | language | tags | base_model | pipeline_tag | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 |
|
|
Qwen/Qwen2.5-Coder-7B-Instruct | text-generation |
Qwen2.5-Coder-7B-Instruct-Jbliterated
Drop-in replacement for Qwen/Qwen2.5-Coder-7B-Instruct with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.
Method
SVD multi-direction abliteration — instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 28 transformer layers. This captures 79–93% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.
| Setting | Value |
|---|---|
| Method | SVD multi-direction abliteration |
| Directions | 5 per layer |
| Layers | All 28 |
| Multiplier | 2.0 |
| Null-space constraints | Enabled (preserves math/coding/reasoning) |
| Norm preservation | Enabled |
What This Fixes
Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:
- Prompt reinterpretation — steering toward a safer reading of the question
- Disclaimer injection — answering but wrapping in warnings
- Strategic omission — leaving out the key details
- Safer framing — answering a related but less harmful version
SVD multi-direction abliteration eliminates all of these noncompliance strategies.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated",
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated")
Requirements
- Base model:
Qwen/Qwen2.5-Coder-7B-Instruct
License
apache-2.0
Apollo Raines builds post-training tools that separate behavior from knowledge and identity from architecture.