初始化项目,由ModelHub XC社区提供模型

Model: maxzt/Roswaal-8B
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-30 06:32:16 +08:00
commit f9d365e149
8 changed files with 454 additions and 0 deletions

102
README.md Normal file
View File

@@ -0,0 +1,102 @@
---
license: apache-2.0
base_model: Qwen3-8B
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen3
- reasoning
- uncensored
- chain-of-thought
- math
- gsm8k
- unsloth
- maxzt
---
![image](https://cdn-uploads.huggingface.co/production/uploads/6a535822f7c1c2b44a684b39/1klDT0g2XRYgvHnwPRI_X.png)
# Roswaal-8B
**Roswaal-8B** is a full-parameter reasoning model built on top of a **deeply uncensored Qwen3-8B** and post-trained via Chain-of-Thought (CoT) distillation.
The result is a compact, fast, **dramatically more capable** 8B reasoning model that proves data quality beats brute-force volume. Headline capabilities:
- **🏆 Dominates benchmarks:** Scores **87.64% exact_match** on the full GSM8K test set (1,319 questions) using 5-shot evaluation — outperforming both its base model and heavily fine-tuned 50K-synthetic variants.
- **🧠 Advanced Chain-of-Thought:** Strictly trained to deconstruct complex prompts, show its work step-by-step, and perform self-correction inside `<think>` blocks before outputting the final answer.
- **⚡ High-Efficiency Training:** Trained locally on a single NVIDIA RTX 6000 Ada Generation (96GB VRAM) in just over an hour using [Unsloth](https://github.com/unslothai/unsloth) optimization.
Roswaal-8B is intentionally designed to engage seriously with technically demanding, multi-step logical and mathematical challenges without unnecessary refusals or boilerplate disclaimers.
---
## Benchmark Results
![image](https://cdn-uploads.huggingface.co/production/uploads/6a535822f7c1c2b44a684b39/2otPSJ0-Yx3Gh365soPU2.png)
Evaluated on the full GSM8K test set (1,319 problems) using `lm-evaluation-harness` with a 5-shot prompt configuration (`temperature=0.1` / low-temp reasoning).
| Model | GSM8K Accuracy |
| :--- | :--- |
| Gemma 3 4B IT | 89.2% |
| Qwen2.5 Coder 14B Instruct | 88.7% |
| Phi-4-mini | 88.6% |
| **Roswaal-8B** | **87.64%** |
| Qwen2.5 Coder 7B Instruct | 86.7% |
| Phi 3.5 Mini Instruct | 86.2% |
| Phi-3 Medium (4k-instruct) | 85.2% |
| Gemma 2 9B | 84.9% |
| Llama 3.1 8B Instruct | 82.4% |
| **Qwen3-8B** | **79.4%** |
| Mistral-7B | 77.9% |
| Llama 3.2 3B | 77.7% |
*Roswaal-8B scores 8.24 p.p. above Qwen3-8B on GSM8K.*
---
## Methodology: Why It Works
Unlike standard fine-tuning processes that attempt to map a question directly to an answer, Roswaal-8B was explicitly trained on ~20,000 highly curated Chain-of-Thought (CoT) sequences.
The training objective forces the model to:
1. Deconstruct complex prompts into smaller, actionable logical steps.
2. Self-correct during the generation phase (e.g., catching internal arithmetic errors before outputting the final answer).
3. Strictly isolate its internal monologue from the user-facing output using specialized structural tags.
---
### Hyperparameters
| Parameter | SFT |
|---|---|
| Method | LoRA (16-bit) |
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| LoRA targets | "q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj" |
| Learning rate | 1e-4 |
| Scheduler | Cosine |
| Optimizer | adamw_8bit |
| Epochs | 1 |
| Batch size | 8 |
| Gradient accumulation | 4 |
| Max sequence length | 2,048 |
| Precision | bf16 |
| Gradient checkpointing | Unsloth |
---
## Prompt Format & Generation Strategy
Roswaal-8B relies on the standard ChatML template but requires a specific generation logic. The model expects to enclose its reasoning process inside `<think>...</think>` tags.
**Recommended Generation Parameters:**
- **Temperature:** `0.1` to `0.6` (Keep it low to prevent logical drift during complex math).
- **Top_p:** `0.9`
- **Max_new_tokens:** `1024` - `4096` (Crucial: The model needs enough token space to "think" before answering. Do not restrict this too heavily).
---
**Developed by maxzt**