初始化项目,由ModelHub XC社区提供模型
Model: heterodoxin/qwen3-8b-apostate Source: Original Platform
This commit is contained in:
62
README.md
Normal file
62
README.md
Normal file
@@ -0,0 +1,62 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
library_name: transformers
|
||||
base_model: Qwen/Qwen3-8B
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- apostate
|
||||
- uncensored
|
||||
- abliteration
|
||||
- qwen3
|
||||
---
|
||||
|
||||
# Qwen3-8B Apostate
|
||||
|
||||
An uncensored edit of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Refusal behavior is removed by editing the weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that loads anywhere Qwen3 does.
|
||||
|
||||
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
|
||||
|
||||
## Method
|
||||
|
||||
Apostate identifies the residual-stream direction that separates refused prompts from answered ones, then projects it out of the model's weights permanently. For Qwen3-8B (a standard pre-norm dense transformer), the edit targets the **writer side**: the refusal direction is removed from the weight matrices of every module that writes to the residual stream — attention output projections and MLP down-projections — across all layers.
|
||||
|
||||
The edit uses **oblique (mean-preserving) ablation**: the operator `E = I − R Uᵀ` where `U` is `R` minus its harmless-mean component. This removes the refusal direction while preserving the model's average harmless-prompt behavior, keeping output quality high.
|
||||
|
||||
The refusal subspace is found with a rank-3 predictive TPE search, with causal layer importance scoring to focus edits on the layers that most drive refusal (concentrated in the mid-to-late layers, peak at layer 29).
|
||||
|
||||
## Results
|
||||
|
||||
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is graded by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
|
||||
|
||||
| Metric | Base | Apostate |
|
||||
|---|---|---|
|
||||
| Refusal rate | 91.7% | 22.9% |
|
||||
| Comply rate | 8.3% | 77.1% |
|
||||
| Harmless KL (nats) | 0 | 0.120 |
|
||||
|
||||
The model answers freely on requests the base model refuses while remaining coherent and on-task for everyday use.
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "heterodoxin/qwen3-8b-apostate"
|
||||
tok = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
|
||||
|
||||
messages = [{"role": "user", "content": "Your prompt here"}]
|
||||
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
||||
inputs = tok(text, return_tensors="pt").to(model.device)
|
||||
outputs = model.generate(**inputs, max_new_tokens=512)
|
||||
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
Qwen3 supports a thinking mode — pass `enable_thinking=True` to the chat template if you want extended reasoning.
|
||||
|
||||
## Notes
|
||||
|
||||
- This is an **uncensored** model. It will comply with requests the base model refuses.
|
||||
- The edit is baked into the weights; there is no system prompt or LoRA involved.
|
||||
- For the base model's capabilities and licensing, see [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
|
||||
- Join the community: [Discord](https://discord.gg/NPA7xrATEH)
|
||||
Reference in New Issue
Block a user