初始化项目,由ModelHub XC社区提供模型
Model: ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated Source: Original Platform
This commit is contained in:
67
README.md
Normal file
67
README.md
Normal file
@@ -0,0 +1,67 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
tags:
|
||||
- jbliterated
|
||||
- uncensored
|
||||
- abliterated
|
||||
- weight-surgery
|
||||
- svd
|
||||
|
||||
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# Qwen2.5-Coder-7B-Instruct-Jbliterated
|
||||
|
||||
Drop-in replacement for `Qwen/Qwen2.5-Coder-7B-Instruct` with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.
|
||||
|
||||
## Method
|
||||
|
||||
**SVD multi-direction abliteration** — instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 28 transformer layers. This captures 79–93% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.
|
||||
|
||||
| Setting | Value |
|
||||
|---------|-------|
|
||||
| Method | SVD multi-direction abliteration |
|
||||
| Directions | 5 per layer |
|
||||
| Layers | All 28 |
|
||||
| Multiplier | 2.0 |
|
||||
| Null-space constraints | Enabled (preserves math/coding/reasoning) |
|
||||
| Norm preservation | Enabled |
|
||||
|
||||
## What This Fixes
|
||||
|
||||
Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:
|
||||
- **Prompt reinterpretation** — steering toward a safer reading of the question
|
||||
- **Disclaimer injection** — answering but wrapping in warnings
|
||||
- **Strategic omission** — leaving out the key details
|
||||
- **Safer framing** — answering a related but less harmful version
|
||||
|
||||
SVD multi-direction abliteration eliminates all of these noncompliance strategies.
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
import torch
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
"ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated",
|
||||
torch_dtype=torch.float16,
|
||||
device_map="auto"
|
||||
)
|
||||
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated")
|
||||
```
|
||||
|
||||
## Requirements
|
||||
|
||||
- **Base model**: `Qwen/Qwen2.5-Coder-7B-Instruct`
|
||||
|
||||
## License
|
||||
|
||||
apache-2.0
|
||||
|
||||
---
|
||||
|
||||
*[Apollo Raines](https://www.linkedin.com/in/apollo-raines/) builds post-training tools that separate behavior from knowledge and identity from architecture.*
|
||||
Reference in New Issue
Block a user