27 lines
778 B
Markdown
27 lines
778 B
Markdown
|
|
---
|
||
|
|
license: apache-2.0
|
||
|
|
base_model: Qwen/Qwen2.5-1.5B-Instruct
|
||
|
|
library_name: transformers
|
||
|
|
language:
|
||
|
|
- en
|
||
|
|
- zh
|
||
|
|
tags:
|
||
|
|
- qwen2.5
|
||
|
|
- knowledge-distillation
|
||
|
|
- gkd
|
||
|
|
- liftkd
|
||
|
|
---
|
||
|
|
|
||
|
|
# Qwen2.5 7B to 1.5B LiftKD V8 Bilingual 100K - Epoch 2
|
||
|
|
|
||
|
|
This is the cumulative epoch-2 checkpoint of a Qwen2.5-1.5B-Instruct student distilled from Qwen2.5-7B-Instruct.
|
||
|
|
|
||
|
|
- Method: LiftKD V8, fully on-policy GKD JSD with normalized gap gate
|
||
|
|
- Data: 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics
|
||
|
|
- Sequence limits: 384 prompt tokens, 512 total tokens, 128 generated tokens
|
||
|
|
- Precision: BF16 full-parameter training with DeepSpeed ZeRO-2
|
||
|
|
- Global batch size: 64
|
||
|
|
- Seed: 10
|
||
|
|
|
||
|
|
The training mixture was internally deduplicated. Evaluation-set decontamination was not performed.
|