Files
ModelHub XC d0353dd27b 初始化项目,由ModelHub XC社区提供模型
Model: FlameF0X/TinyMoE-100m-2x8-chat-stage1
Source: Original Platform
2026-08-17 20:12:17 +08:00

179 lines
6.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
language:
- en
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- moe
- mixture-of-experts
- chat
- sft
- fine-tuned
- pytorch
- tiny
- efficient
- trl
base_model: FlameF0X/TinyMoE-100m-2x8-retrained
datasets:
- HuggingFaceTB/smol-smoltalk
- yahma/alpaca-cleaned
- databricks/databricks-dolly-15k
- HuggingFaceH4/ultrachat_200k
- Open-Orca/OpenOrca
- teknium/openhermes
- HuggingFaceH4/no_robots
model-index:
- name: TinyMoE-100m-2x8-chat-stage1
results: []
---
# TinyMoE-100m-2x8 Chat (Stage 1)
A **chat fine-tuned Mixture of Experts (MoE) language model** — small, efficient, and fully open-source.
TinyMoE is a ~100M parameter transformer using a **Mixture of Experts** architecture, fine-tuned via supervised fine-tuning (SFT) on a diverse blend of high-quality chat and instruction datasets. It's designed to be a compact but capable conversational AI that knows its own identity.
## Model Details
| Property | Value |
|---|---|
| **Architecture** | Mixture of Experts (MoE) Transformer |
| **Parameters** | ~100M total |
| **Experts** | 8 experts, top-2 routing per token |
| **Context Length** | 4,096 tokens (extended via linear RoPE scaling) |
| **Chat Template** | ChatML-style — `<\|system\|>`, `<\|user\|>`, `<\|assistant\|>` |
| **Base Model** | [TinyMoE-100m-2x8-retrained](https://huggingface.co/FlameF0X/TinyMoE-100m-2x8-retrained) |
| **Training Type** | Full-weight SFT (not LoRA) |
| **Precision** | bfloat16 |
| **Creator** | [FlameF0X](https://huggingface.co/FlameF0X) |
| **License** | Apache 2.0 |
## What is Mixture of Experts?
Unlike a standard dense transformer where every token goes through the same large feed-forward network, TinyMoE uses **8 expert sub-networks** with a learned router that selects the top-2 experts per token. This means:
- **More total knowledge capacity** without proportionally increasing compute
- **Sparse activation** — only a fraction of parameters fire per token
- **Efficient inference** — you get more model per FLOP
This is the same architectural family as Mixtral, but at a much smaller scale — proving that MoE works even at ~100M parameters.
## Training Recipe
### Stage 1 — Chat SFT
The base pretrained model was fine-tuned on a carefully curated mixture of chat datasets to teach conversational ability, instruction following, and model identity.
**Hardware:** NVIDIA L4 (24 GB) on [Modal](https://modal.com)
**Framework:** TRL (Transformer Reinforcement Learning) `SFTTrainer`
**Max examples:** 80,000 (after length filtering)
| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Learning Rate | 2e-5 |
| LR Schedule | Cosine with 5% warmup |
| Optimizer | AdamW (weight decay 0.01) |
| Batch Size | 4 per device × 8 grad accum = 32 effective |
| Max Gradient Norm | 1.0 |
| Packing | Yes |
| Gradient Checkpointing | Yes |
### Training Datasets
| Dataset | Examples | Description |
|---|---|---|
| [SmolTalk](https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk) | ~4k | Synthetic diverse chat conversations |
| [Alpaca Cleaned](https://huggingface.co/datasets/yahma/alpaca-cleaned) | ~52k | Cleaned instruction-following data |
| [Dolly 15k](https://huggingface.co/datasets/databricks/databricks-dolly-15k) | ~15k | Human-written instruction/response pairs |
| [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) | 12k subset | High-quality multi-turn chat |
| [OpenOrca](https://huggingface.co/datasets/Open-Orca/OpenOrca) | 15k subset | GPT-4 augmented FLAN instructions |
| [OpenHermes](https://huggingface.co/datasets/teknium/openhermes) | 10k subset | Diverse tasks (code, write, reason, roleplay) |
| [No Robots](https://huggingface.co/datasets/HuggingFaceH4/no_robots) | 10k | Hand-curated high-quality SFT examples |
| **TinyMoE Identity** *(custom)* | 44 | Synthetic identity Q&A + multi-turn conversations |
### Identity Training
The model was explicitly taught to know it's **TinyMoE** through two mechanisms:
1. **Dedicated identity dataset** — 44 custom examples covering name, creator, architecture, capabilities, and differentiation from other models (ChatGPT, Claude, Llama, etc.)
2. **System prompt injection** — 15% of all training examples received a system prompt: *"You are TinyMoE, a helpful Mixture of Experts AI assistant created by FlameF0X."*
This means TinyMoE knows who it is — ask it "What's your name?" or "Who created you?" and it will answer correctly.
## Usage
### Chat Format
TinyMoE uses a **ChatML-style** template with special tokens:
```
<|system|>
You are TinyMoE, a helpful Mixture of Experts AI assistant created by FlameF0X.</s>
<|user|>
What's Mixture of Experts?</s>
<|assistant|>
MoE stands for Mixture of Experts! Instead of one big neural network...</s>
```
### Quick Start
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "FlameF0X/TinyMoE-100m-2x8-chat-stage1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are TinyMoE, a helpful Mixture of Experts AI assistant created by FlameF0X."},
{"role": "user", "content": "What's your name and who made you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=256,
temperature=0.7,
do_sample=True,
top_p=0.9,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Capabilities & Limitations
### ✅ Strengths
- **Efficient** — MoE architecture means more capacity per inference FLOP
- **Conversational** — trained on diverse multi-turn chat data
- **Self-aware** — knows it's TinyMoE, not ChatGPT/Claude/Llama
- **Open-source** — weights, architecture, and training code are all public
- **Fast** — small enough to run on consumer hardware or free-tier GPUs
### ⚠️ Limitations
- **Small model** — at 100M parameters, factual knowledge is limited compared to billion-parameter models
- **Stage 1 only** — this is a direct-answer SFT model; it hasn't undergone RLHF/DPO alignment
- **No CoT** — training explicitly excluded chain-of-thought reasoning traces (saved for future stages)
- **English only** — training data was English-dominant
- **May hallucinate** — like all LLMs, it can generate incorrect information with confidence
### Future Stages (planned)
- **Stage 2:** MCP/Tool usage + RL
- **Longer context** — potential extension beyond 4K tokens