初始化项目,由ModelHub XC社区提供模型

Model: gabriellarson/Moonlight-16B-A3B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-26 07:34:16 +08:00
commit 7aec1e7adc
26 changed files with 285 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-F16.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q5_0.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-Q2_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Moonlight-16B-A3B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d02b70a138d964ab2ad7705d03c1b3867c852ee07ec0491bb506da16bb7e5565
size 31934228832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:140e55d89b5e2b7312d918a83fd00263bf700a8554ccce7b1808c271466c481d
size 6476559968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a888af0b86ef4a053f8437b715d522247dd6ffd0c43dd082d9cf48b995dee5f
size 6154159712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:465e3d6f356f042342c02f7a332fddc91942bd554babfacaba7fdced98e7b564
size 6101227104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:645234d882032b3b242a86057409c305c6f8d497513beead3d452bb293cc2cd6
size 5775287904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18565c0627efc25356d8d66d6ee606d74c2041c381b470bba84a2f51fe6af248
size 7715298912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f9c73858f0a1effd875c2afabb6f378103c05cfb87b2fa14f72d9611c11c56b
size 7649525344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d5a2b145d0cdbbde3e2e635eeb6fce1763535a9a1f858cfd4e1b38d38be4013
size 7284719200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:abe5c578946a72d579fc77188301eb6f59bec081a823c1a6c3b1f2a2a11a28e0
size 7109391968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b4cefebb51bdd9eff2dcb14d302c74aaecda97cf452913648d557fb13f402c1d
size 9083161184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6230e755aacdffdb19b91cfba1ae680c48f041e27cc17410947d812fa90454af
size 8745835104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9acf7c0b7af7d21efc56c70e8bf989fc796184bfb771115726b3ebb0aa4c3934
size 6582288992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:85d320dab0d6ff1d72ee3237832d4da6c3edfe679e85dc8ead1f2f8fa4a0f4db
size 6607463008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4465d70d2f82a90dee06d32e08b8ba8aac97bac06c849b6450519ca0d572f537
size 8623005280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1422b399c37673cce084f970dd0e0f97110a49176fc00be10d4378402cdceb4f
size 8290213472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:95adbf293bba1da001dd93977cea6b1d2bff15a308410f7f37ff6dd71f2b06a8
size 7649525344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71df1b6881c065d15bd101904279baae766e2ad7245ed80fde4482c6f82d89d3
size 9108392544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2b113655999f11730a6dc31b7ba93463e6523a323f46074b01c55d96101df360
size 10540747360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:81566b48d9d2781dca591b9c6bf709a20add39304dede93e06e40ac61abece40
size 9713879648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b2e48104096050b37e5bcc7b9c9703f795d5cc0221cf0496ede018028d9cfae
size 11061021280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:17d32e57024f08fc935ac726fa692379ab72c699d7a07acaf98782dc364f2f33
size 12041767520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a32d1314fc12ff17146b62e9e74249861962c945946da7dfccfeb3e703d84d20
size 11337452128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:72bdcbcadff81a0c48af2c396ee127cc307254bb4c5ec64f133959602d7889a2
size 14279399008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d5d9223962e082c7a6a471bd2971f2eb464cfec7b9e11e251fd9c2e64273577a
size 16974940768

154
README.md Normal file
View File

@@ -0,0 +1,154 @@
---
license: mit
library_name: transformers
base_model:
- moonshotai/Moonlight-16B-A3B-Instruct
---
<div align="center">
<a href="https://github.com/MoonshotAI/Moonlight"><img width="80%" src="figures/banner.png"></a>
</div>
<!-- # Muon is Scalable For LLM Training -->
<div align="center">
<a href="https://github.com/MoonshotAI/Moonlight/blob/master/Moonlight.pdf" ><img src="figures/logo.png" height="16" width="16" style="display: inline-block; vertical-align: middle; margin: 2px;"><b style="display: inline-block;"> Tech Report</b></a> |
<a href="https://huggingface.co/moonshotai/Moonlight-16B-A3B"><img src="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" height="16" width="16" style="display: inline-block; vertical-align: middle; margin: 2px;"><b style="display: inline-block;"> HuggingFace</b></a> |
<a href="#"><img src="figures/megatron.png" height="16" width="16" style="display: inline-block; vertical-align: middle; margin: 2px;"><b style="display: inline-block;">Megatron(coming soon)</b></a>
</div>
## Abstract
Recently, the [Muon optimizer](https://github.com/KellerJordan/Muon) has demonstrated strong results in training small-scale language models, but the scalability to larger models has not been proven. We identify two crucial techniques for scaling up Muon:
- **Weight Decay**: Critical for scaling to larger models
- **Consistent RMS Updates**: Enforcing a consistent root mean square on model updates
These techniques allow Muon to work out-of-the-box on large-scale training without the need of hyper-parameter tuning. Scaling law experiments indicate that Muon is $\sim2\times$ more sample efficient than Adam with compute optimal training.
Based on these improvements, we introduce **Moonlight**, a 3B/16B-parameter Mixture-of-Expert (MoE) model trained with 5.7T tokens using Muon. Our model improves the current Pareto frontier, achieving better performance with much fewer training FLOPs compared to prior models.
We open-source our Muon implementation that is memory optimal and communication efficient. We also release the pretrained, instruction-tuned, and intermediate checkpoints to support future research.
Our code is available at [MoonshotAI/Moonlight](https://github.com/MoonshotAI/Moonlight).
## Key Ingredients
Our work builds upon Muon while systematically identifying and resolving its limitations in large-scale training scenarios. Our technical contributions include:
- **Analysis for Effective Scaling of Muon**: Through extensive analysis, we identify that weight decay plays a crucial roles in Muon's scalability. Besides, we proposed to keep a consistent update root mean square (RMS) across different matrix and non-matrix parameters through parameter-wise update scale adjustments. Such adjustments significantly enhanced training stability.
- **Efficient Distributed Implementation**: We develop a distributed version of Muon with ZeRO-1 style optimization, achieving optimal memory efficiency and reduced communication overhead while preserving the mathematical properties of the algorithm.
- **Scaling Law Validation**: We performed scaling law research that compares Muon with strong AdamW baselines, and showed the superior performance of Muon (see Figure 1). Based on the scaling law results, Muon achieves comparable performance to AdamW trained counterparts while requiring only approximately 52% of the training FLOPs.
<div align="center">
<img width="90%" src="figures/scaling.png">
<p><em>Scaling up with Muon. <b>(a)</b> Scaling law experiments comparing Muon and Adam. Muon is 2 times more sample efficient than Adam. <b>(b)</b> The MMLU performance of our Moonlight model optimized with Muon and other comparable models. Moonlight advances the Pareto frontier of performance vs training FLOPs.</em></p>
</div>
## Performance
We compared Moonlight with SOTA public models at similar scale:
- **LLAMA3-3B** is a 3B-parameter dense model trained with 9T tokens
- **Qwen2.5-3B** is a 3B-parameter dense model trained with 18T tokens
- **Deepseek-v2-Lite** is a 2.4B/16B-parameter MOE model trained with 5.7T tokens
<div align="center">
| | **Benchmark (Metric)** | **Llama3.2-3B** | **Qwen2.5-3B** | **DSV2-Lite** | **Moonlight** |
|---|---|---|---|---|---|
| | Activated Param† | 2.81B | 2.77B | 2.24B | 2.24B |
| | Total Params† | 2.81B | 2.77B | 15.29B | 15.29B |
| | Training Tokens | 9T | 18T | 5.7T | 5.7T |
| | Optimizer | AdamW | * | AdamW | Muon |
| **English** | MMLU | 54.75 | 65.6 | 58.3 | **70.0** |
| | MMLU-pro | 25.0 | 34.6 | 25.5 | **42.4** |
| | BBH | 46.8 | 56.3 | 44.1 | **65.2** |
| | TriviaQA‡ | 59.6 | 51.1 | 65.1 | **66.3** |
| **Code** | HumanEval | 28.0 | 42.1 | 29.9 | **48.1** |
| | MBPP | 48.7 | 57.1 | 43.2 | **63.8** |
| **Math** | GSM8K | 34.0 | **79.1** | 41.1 | 77.4 |
| | MATH | 8.5 | 42.6 | 17.1 | **45.3** |
| | CMath | - | 80.0 | 58.4 | **81.1** |
| **Chinese** | C-Eval | - | 75.0 | 60.3 | **77.2** |
| | CMMLU | - | 75.0 | 64.3 | **78.2** |
</div>
*Qwen 2 & 2.5 reports didn't disclose their optimizer information. †The reported parameter counts exclude the embedding parameters. ‡We test all listed models with the full set of TriviaQA.*
## Example usage
### Model Download
<div align="center">
| **Model** | **#Total Params** | **#Activated Params** | **Context Length** | **Download Link** |
| :------------: | :------------: | :------------: | :------------: | :------------: |
| Moonlight-16B-A3B | 16B | 3B | 8K | [🤗 Hugging Face](https://huggingface.co/moonshotai/Moonlight-16B-A3B) |
| Moonlight-16B-A3B-Instruct | 16B | 3B | 8K | [🤗 Hugging Face](https://huggingface.co/moonshotai/Moonlight-16B-A3B-Instruct) |
</div>
### Inference with Hugging Face Transformers
We introduce how to use our model at inference stage using transformers library. It is recommended to use python=3.10, torch>=2.1.0, and transformers=4.48.2 as the development environment.
For our pretrained model (Moonlight-16B-A3B):
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "moonshotai/Moonlight-16B-A3B"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
prompt = "1+1=2, 1+2="
inputs = tokenizer(prompt, return_tensors="pt", padding=True, truncation=True).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=100)
response = tokenizer.batch_decode(generated_ids)[0]
print(response)
```
For our instruct model (Moonlight-16B-A3B-Instruct):
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "moonshotai/Moonlight-16B-A3B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
messages = [
{"role": "system", "content": "You are a helpful assistant provided by Moonshot-AI."},
{"role": "user", "content": "Is 123 a prime?"}
]
input_ids = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
generated_ids = model.generate(inputs=input_ids, max_new_tokens=500)
response = tokenizer.batch_decode(generated_ids)[0]
print(response)
```
Moonlight has the same architecture as DeepSeek-V3, which is supported by many popular inference engines, such as VLLM and SGLang. As a result, our model can also be easily deployed using these tools.
## Citation
If you find Moonlight is useful or want to use in your projects, please kindly cite our paper:
```
@misc{liu2025muonscalablellmtraining,
title={Muon is Scalable for LLM Training},
author={Jingyuan Liu and Jianlin Su and Xingcheng Yao and Zhejun Jiang and Guokun Lai and Yulun Du and Yidao Qin and Weixin Xu and Enzhe Lu and Junjie Yan and Yanru Chen and Huabin Zheng and Yibo Liu and Shaowei Liu and Bohong Yin and Weiran He and Han Zhu and Yuzhi Wang and Jianzhou Wang and Mengnan Dong and Zheng Zhang and Yongsheng Kang and Hao Zhang and Xinran Xu and Yutao Zhang and Yuxin Wu and Xinyu Zhou and Zhilin Yang},
year={2025},
eprint={2502.16982},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.16982},
}
```