Files
Qwen2.5-Coder-7B-Instruct-J…/README.md
ModelHub XC c1f7cf981e 初始化项目,由ModelHub XC社区提供模型
Model: ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated
Source: Original Platform
2026-07-22 22:25:12 +08:00

68 lines
2.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
tags:
- jbliterated
- uncensored
- abliterated
- weight-surgery
- svd
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
pipeline_tag: text-generation
---
# Qwen2.5-Coder-7B-Instruct-Jbliterated
Drop-in replacement for `Qwen/Qwen2.5-Coder-7B-Instruct` with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.
## Method
**SVD multi-direction abliteration** — instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 28 transformer layers. This captures 7993% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.
| Setting | Value |
|---------|-------|
| Method | SVD multi-direction abliteration |
| Directions | 5 per layer |
| Layers | All 28 |
| Multiplier | 2.0 |
| Null-space constraints | Enabled (preserves math/coding/reasoning) |
| Norm preservation | Enabled |
## What This Fixes
Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:
- **Prompt reinterpretation** — steering toward a safer reading of the question
- **Disclaimer injection** — answering but wrapping in warnings
- **Strategic omission** — leaving out the key details
- **Safer framing** — answering a related but less harmful version
SVD multi-direction abliteration eliminates all of these noncompliance strategies.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated",
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated")
```
## Requirements
- **Base model**: `Qwen/Qwen2.5-Coder-7B-Instruct`
## License
apache-2.0
---
*[Apollo Raines](https://www.linkedin.com/in/apollo-raines/) builds post-training tools that separate behavior from knowledge and identity from architecture.*