Files
Haidass-143M-v1/README.md
ModelHub XC abc6e7ed8a 初始化项目,由ModelHub XC社区提供模型
Model: DALabCommunity/Haidass-143M-v1
Source: Original Platform
2026-09-10 10:44:16 +08:00

136 lines
3.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
language:
- en
- zh
license: apache-2.0
datasets:
- openbmb/Ultra-FineWeb
- mlfoundations/dclm-baseline-1.0
- HuggingFaceTB/finemath
tags:
- haidass
- npu
- bilingual
- mindspeed-llm
library_name: transformers
pipeline_tag: text-generation
---
<div align="center">
<img src="logo.png" width="400"/>
</div>
# Haidass-143M
<p align="center">
English |
<a href="https://huggingface.co/DALabCommunity/Haidass-143M-v1/blob/main/README_ZH.md">中文</a>
</p>
A bilingual (English/Chinese) small language model trained entirely on **Huawei Ascend** NPU ecosystem.
## Model Overview
Haidass-143M is a 143M-parameter bilingual language model trained on approximately 100B tokens of English and Chinese data. The entire training pipeline runs on the Huawei Ascend ecosystem, using the **MindSpeed-LLM** framework on Atlas A2 servers (910B). A custom 64,000-token bilingual vocabulary (SentencePiece BPE) was trained alongside the model. This model is competitive among multilingual models under 150M parameters and ranks favorably across multiple evaluation benchmarks.
## Model Architecture
| Parameter | Value |
|------|------|
| Architecture | Qwen3 |
| Layers | 30 |
| Hidden size | 576 |
| Attention heads | 9 |
| KV heads (GQA) | 3 |
| Head dim | 64 |
| FFN intermediate size | 1,536 |
| Vocabulary size | 64,000 |
| Max sequence length | 4,096 |
| Tie word embeddings | Yes |
| Position encoding | RoPE (θ=100,000) |
| Attention bias | None |
| Precision | BF16 |
| Total parameters | ~143M |
## Training Data
The model was trained on approximately 100B tokens of mixed English and Chinese data. Primary data sources:
- [openbmb/Ultra-FineWeb](https://huggingface.co/datasets/openbmb/Ultra-FineWeb) (ultrafineweb-en + ultrafineweb-zh)
- [mlfoundations/dclm-baseline-1.0-parquet](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet) (dclm)
- [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) (finemath-4plus)
## Training Configuration
| Parameter | Value |
|------|------|
| Framework | MindSpeed-LLM (v2.3.0) |
| Hardware | 8 × Atlas A2 servers (8 NPUs per node, 256 cores) |
| NPU model | Huawei Ascend 910B |
| Total NPUs | 64 (8 nodes × 8 cards) |
| Sequence length | 4,096 |
## Optimizer
| Parameter | Value |
|------|------|
| Optimizer | AdamW |
| Peak learning rate | 3e-4 |
| Min learning rate | 3e-5 |
## Tokenizer
| Property | Value |
|------|------|
| Type | SentencePiece BPE |
| Vocabulary size | 64,000 |
| Language coverage | English + Chinese |
## Evaluation
Evaluated at checkpoint (~98B tokens) using the lighteval framework (v0.9.2).
| Benchmark | Score |
|------|------|
| ARC-Easy | 60.44 |
| ARC-Challenge | 27.13 |
| PIQA | 67.25 |
| HellaSwag |37.91 |
| OpenBookQA | 31.8 |
| Winogrande | 52.17 |
| agi_eval | 23.78 |
## Key Features
- **Fully Ascend-native**: Trained entirely on Huawei Ascend 910B NPUs using the MindSpeed-LLM framework
- **Bilingual**: Trained on a mixture of English and Chinese data
## Intended Use
This is a research model, suitable for:
- Studying training dynamics of small models on Ascend NPUs
- English/Chinese language modeling research
- Serving as a base model for fine-tuning or annealing experiments
## Limitations
- Small model scale; reasoning and generation capabilities are limited
- raw pretrained model only
## Citation
```bibtex
@misc{haidass-143m,
title={haidass-143M: A Bilingual Small Language Model Trained on Ascend 910B},
year={2026},
note={Based on Qwen3 architecture, trained from scratch on 100B tokens using MindSpeed-LLM on 64× Ascend 910B NPUs}
}
```
## License
Apache 2.0