Files
ModelHub XC d264172256 初始化项目,由ModelHub XC社区提供模型
Model: vector-institute/Qwen3-8B-UnBias-Plus-SFT-Instruct-V2
Source: Original Platform
2026-07-18 11:13:09 +08:00

230 lines
11 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen3-8B
datasets:
- vector-institute/unbias-plus-dataset
language:
- en
tags:
- qwen3
- 8b
- bias-detection
- news-debiasing
- fine-tuned
- unsloth
- sft
- conversational
pipeline_tag: text-generation
---
# Qwen3-8B-UnBias-Plus-SFT-Instruct-V2
A fine-tuned version of [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) for news media bias detection and neutral rewriting, developed by the [Vector Institute](https://vectorinstitute.ai) as part of the [UnBias-Plus](https://github.com/VectorInstitute/unbias-plus) project.
Given a news article, the model identifies biased language segments, classifies each segment's bias type and severity, provides neutral replacements, assigns a single article-level severity score, and returns a fully rewritten unbiased version, all in a single structured JSON response.
This is the **V2 Instruct variant**, trained without chain-of-thought thinking blocks (`enable_thinking=False`). It produces clean structured JSON directly, making it suitable for production inference via vLLM or other OpenAI-compatible backends.
V2 is trained as a **conservative span-level bias annotator**: it flags only clear, material bias (while always flagging toxic/dehumanizing language), applies a strict bias-type label precedence, treats attributed/quoted language under the same standard, and changes only flagged text when producing the rewrite.
> **Trained on:** [train_4](https://huggingface.co/datasets/vector-institute/unbias-plus-dataset) (5,000 expert-annotated news articles)
## Difference from Qwen3-8B-UnBias-Plus-SFT-Instruct
| | SFT-Instruct (predecessor) | SFT-Instruct-V2 (this model) |
| ------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------------------------- |
| Training data | train_3 (UnBias-Plus-3000) | train_4 (5,000 expert-annotated articles) |
| Article severity scale | `0, 1, 2, 3, 4` | `010` integer |
| Segment severity | `low` / `medium` / `high` | `Low` / `Medium` / `High` |
| Output fields | `severity`, `bias_found`, `biased_segments`, `unbiased_text` | `severity`, `biased_segments`, `unbiased_text` (label/bias_found derived downstream) |
| Bias type taxonomy | 7 free-form labels | 8 canonical slugs with strict label precedence |
| System prompt | Bias detection + rewrite rules | Conservative span-level annotator (label precedence, attribution handling, toxicity rules, audit step) |
## Model Details
| Property | Value |
| ------------------ | -------------------------------------------------------- |
| Base model | Qwen/Qwen3-8B |
| Fine-tuning method | Supervised Fine-Tuning (SFT) with LoRA |
| Training precision | bf16 (full precision, no quantization during training) |
| LoRA rank | 16 |
| Training framework | Unsloth + TRL |
| Context length | 8192 tokens |
| Thinking mode | Disabled (`enable_thinking=False`) |
| Output format | Structured JSON |
| Training dataset | [train_4](https://huggingface.co/datasets/vector-institute/unbias-plus-dataset) (5,000 expert-annotated articles) |
## Usage
The exact system prompt used during training is maintained in
[`src/unbias_plus/prompt.py`](https://github.com/VectorInstitute/unbias-plus/blob/main/src/unbias_plus/prompt.py)
(`SYSTEM_PROMPT`). For faithful results, use that prompt verbatim — a condensed
version is shown below for illustration. The easiest path is to use the
[`unbias-plus`](https://github.com/VectorInstitute/unbias-plus) package, which
wires up the prompt, parsing, and offset computation for you.
### Using with transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch, json
model_id = "vector-institute/Qwen3-8B-UnBias-Plus-SFT-Instruct-V2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
# Use the full SYSTEM_PROMPT from src/unbias_plus/prompt.py for best results.
SYSTEM_PROMPT = """You are a conservative span-level bias annotator.
Given an article, identify material biased language, assign one article-level
severity score (0-10), and produce a neutral rewrite that changes only flagged
text. Return the result only as a single valid JSON object with the fields
`severity`, `biased_segments`, and `unbiased_text`. Do not output `binary_label`
or `bias_found`. Do not write anything outside the JSON object."""
article = "Your news article here..."
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Analyze the following article for bias and return the result in the required JSON format.\n\nARTICLE:\n{article}"},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
truncation=True,
max_length=8192,
)
with torch.no_grad():
outputs = model.generate(
input_ids=inputs["input_ids"].to(model.device),
attention_mask=inputs["attention_mask"].to(model.device),
max_new_tokens=4096,
do_sample=False,
temperature=None,
top_p=None,
pad_token_id=tokenizer.eos_token_id,
)
new_tokens = outputs[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(new_tokens, skip_special_tokens=True)
result = json.loads(response)
```
### Using with vLLM
```bash
vllm serve vector-institute/Qwen3-8B-UnBias-Plus-SFT-Instruct-V2 \
--served-model-name unbias-plus \
--max-model-len 8192 \
--dtype bfloat16
```
```python
from openai import OpenAI
import json
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
completion = client.chat.completions.create(
model="unbias-plus",
messages=messages,
max_tokens=4096,
temperature=0,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
result = json.loads(completion.choices[0].message.content)
```
## Output Schema
The model emits exactly three fields. `binary_label` and `bias_found` are **not**
produced by the model — they are derived downstream from `severity`
(`severity > 0` ⇒ biased).
```json
{
"severity": 0,
"biased_segments": [
{
"original": "exact substring from input article",
"replacement": "neutral alternative phrase (or empty string to delete)",
"severity": "Low | Medium | High",
"bias_type": "loaded_language | euphemism | dehumanizing_language | opinion_as_fact | unsupported_generalization | stereotypical_association | sensationalism | informational_bias",
"reasoning": "short, specific explanation of the linguistic cue"
}
],
"unbiased_text": "Full rewritten neutral article"
}
```
- `severity` is an integer `010` (article-level).
- Each segment's `severity` is a string: `Low`, `Medium`, or `High`.
- `bias_type` is exactly one of the eight canonical slugs listed below.
- If `biased_segments` is empty, article `severity` is `0`. If `severity > 0`, there is at least one segment.
### Article-level Severity Scale
| Value | Meaning |
| ------ | -------------------------------------------- |
| 0 | No biased segments |
| 15 | Limited, low, or moderate bias |
| 610 | Strong, recurring, or highly distorting bias |
## Bias Types Detected
V2 assigns exactly one primary `bias_type` per segment, using the following
label precedence (first applicable label wins):
1. **`dehumanizing_language`** — treats people or groups as less than human, as a threatening mass, or as inherently dangerous (e.g. "vermin", "flood of migrants").
2. **`stereotypical_association`** — assigns traits, roles, or motives to a demographic/social/political/protected group, links identity to suspicion or wrongdoing, or reduces a group to a single function.
3. **`sensationalism`** — hyperbolic, alarmist, or dramatic language that inflates significance (e.g. "bombshell", "catastrophe", "sparks fury").
4. **`opinion_as_fact`** — an unattributed subjective or evaluative judgment presented as fact (e.g. "the policy is a failure").
5. **`unsupported_generalization`** — a sweeping or absolute claim about a group or situation without support (e.g. "everyone knows", "immigrants always").
6. **`euphemism`** — softened or vague wording that minimizes, hides, or downplays harmful facts (e.g. "collateral damage", "enhanced interrogation").
7. **`informational_bias`** — authorial framing that visibly treats one side, source, or claim differently in a way that steers interpretation.
8. **`loaded_language`** — clearly charged, morally weighted, editorial, toxic, abusive, or inflammatory wording not covered by a more specific type (e.g. "thugs", "so-called reform").
## Hardware Requirements
| Setup | Configuration |
| -------------------- | --------------------------------------------------------- |
| Recommended (server) | `torch_dtype=torch.bfloat16, device_map="auto"` (~16GB VRAM) |
| Lightweight (laptop) | `load_in_4bit=True` (~5GB VRAM) |
## Model Variants
| Model | Params | Context | Best for |
| ----- | ------ | ------- | -------- |
| [Qwen3.5-4B-UnBias-Plus-SFT-Instruct](https://huggingface.co/vector-institute/Qwen3.5-4B-UnBias-Plus-SFT-Instruct) | 4B | 4096 | Speed, low VRAM |
| Qwen3-8B-UnBias-Plus-SFT-Instruct-V2 (this) | 8B | 8192 | Higher recall, longer articles, production deployment |
## Citation
```@article{radwan2026unbias,
title={UnBias-Plus: Detect, Explain, and Rewrite Bias},
author={Radwan, Ahmed Y and ElKady, Ahmed and Chaduvula, Sindhuja and Hafez, Mohamed and Krishnan, Amrit and Raza, Shaina},
journal={arXiv preprint arXiv:2606.23412},
year={2026}
}
```
## Links
- 📊 Leaderboard: [UnBias-Plus Leaderboard](https://huggingface.co/spaces/vector-institute/UnBias-Plus-Leaderboard)
- 📁 Dataset: [vector-institute/unbias-plus-dataset](https://huggingface.co/datasets/vector-institute/unbias-plus-dataset)
- 🏛️ Organization: [Vector Institute](https://vectorinstitute.ai)
- 🌐 Project: [AIXpert](https://aixpert-project.eu/)