Files
AL-0-SFT/README.md
ModelHub XC 6e70e4873a 初始化项目,由ModelHub XC社区提供模型
Model: zxcAsD01/AL-0-SFT
Source: Original Platform
2026-08-01 15:05:13 +08:00

117 lines
3.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
frameworks: PyTorch
tasks:
- text-generation
license: Apache License 2.0
base_model:
- zxcAsD01/AL-0
datasets:
- zxcAsD01/MiMo-V2-Flash_distill
language:
- zh
---
***
# AL-0 (Distilled on Qwen3)
<div align="center">
![visitors](https://visitor-badge.laobi.icu/badge?page_id=your-repo/AL-0)
[![GitHub Code License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)
</div>
## 📌 模型介绍 (Model Introduction)
**AL-0** 是一个由 **zxcAsD01** 创建的学习性质的小型语言模型。该模型是基于 **Qwen3-0.6B** 架构进行蒸馏(Distillation)并结合常规训练集进行微调(Fine-tuning)得到的。
尽管体积小巧,AL-0 旨在通过蒸馏(Distillation)技术,在保持高效推理的同时,尽可能继承大模型的逻辑推理与对话能力。
**AL-0** 的前置模型为同为 **zxcAsD01** 创建的 **AL-0-Base** , **AL-0-Base** 是一个纯基座模型,经过极少的训练即可从接近随机输出收敛到输出连贯文本。 **AL-0-SFT** 即为其 **Base** 模型经少量额外预训练以及较多的指令微调得到的。
### 模型架构参数 (Architecture)
**AL-0** 的关键参数如下:
* **Base Model:** Qwen3
* **Hidden Size:** 768
* **Num Hidden Layers:** 12
* **Num Attention Heads:** 12
* **Vocabulary Size:** 151,936
* **Context Length:** 40,96 tokens
## 🛠️ 使用说明 (Usage)
### 1. 环境依赖
请确保安装了最新的 `transformers` 库:
```bash
pip install transformers>=4.57.1
```
### 2. 加载模型与推理
您可以使用以下 Python 代码加载并使用 AL-0 模型:
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
model_path = "your-username/AL-0"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype="auto",
device_map="auto"
)
messages = [
{"role": "system", "content": "You are AL-0, a helpful assistant."},
{"role": "user", "content": "你好,请介绍一下你自己。"}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.7,
top_p=0.8
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
```
### 3. 推理参数建议
根据 Qwen3 的特性,建议如下参数:
* **非思考模式 (Non-Thinking Mode):** `temperature=0.7`, `top_p=0.8`, `top_k=20`。
* **思考模式 (Thinking Mode):** 如果模型支持,可使用 `temperature=0.6`, `top_p=0.95`, `top_k=20`。
## 🌐 模型来源 (Model Source)
AL-0 模型是基于 **Qwen3** 系列模型进行蒸馏和微调得到的。
**硬件**
Intel Core Ultra9-285H 32G
RTX5070laptop 8G
* **基础模型 (Base Model):** Qwen3-0.6B
* **来源团队:** Qwen Team (阿里云)
* **原始模型特点:** Qwen3 是阿里团队的开源大语言模型系列,支持思考模式与非思考模式,预训练数据覆盖 119 种语言,总计约 36 万亿 Token。
本模型的开发得益于 Qwen3 强大的基础能力以及“强到弱蒸馏”技术,使得小型模型也能具备出色的性能。
## 📄 许可证 (License)
本模型基于 **Apache 2.0 许可证** 开源发布。
Qwen3 系列模型本身也遵循 Apache 2.0 许可证。Apache 2.0 是一个商业友好的开源许可证,允许用户自由使用、修改和分发模型,包括用于商业目的和盈利,前提是保留版权声明和免责声明。
详细条款请参见 [LICENSE](LICENSE) 文件。
## 🙏 致谢 (Acknowledgements)
* 感谢 **Qwen Team** 开源了强大的 Qwen3 模型系列。
* 感谢开源社区提供的训练数据集与工具支持。
***