Files
qwen2.5-7b-to-1.5b-liftkd-v…/README.md
ModelHub XC 00a48532c8 初始化项目,由ModelHub XC社区提供模型
Model: huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500
Source: Original Platform
2026-07-20 15:18:11 +08:00

27 lines
778 B
Markdown

---
license: apache-2.0
base_model: Qwen/Qwen2.5-1.5B-Instruct
library_name: transformers
language:
- en
- zh
tags:
- qwen2.5
- knowledge-distillation
- gkd
- liftkd
---
# Qwen2.5 7B to 1.5B LiftKD V8 Bilingual 100K - Epoch 2
This is the cumulative epoch-2 checkpoint of a Qwen2.5-1.5B-Instruct student distilled from Qwen2.5-7B-Instruct.
- Method: LiftKD V8, fully on-policy GKD JSD with normalized gap gate
- Data: 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics
- Sequence limits: 384 prompt tokens, 512 total tokens, 128 generated tokens
- Precision: BF16 full-parameter training with DeepSpeed ZeRO-2
- Global batch size: 64
- Seed: 10
The training mixture was internally deduplicated. Evaluation-set decontamination was not performed.