Files
ModelHub XC eaede82b41 初始化项目,由ModelHub XC社区提供模型
Model: nikitastheo/babylm-ara-ell-sequential_interleaved
Source: Original Platform
2026-07-20 04:10:13 +08:00

25 lines
580 B
Markdown

---
tags:
- causal-lm
- text-generation
library_name: transformers
---
# nikitastheo/babylm-ara-ell-sequential_interleaved
Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no `Trainer`).
## Training details
- **Base config**: `gpt_base_config.json`
- **Tokenizer**: `nikitastheo/babylm-ara-ell-tokenizer`
- **Max steps**: 21390
- **Learning rate**: 0.0001
- **LR scheduler**: linear
- **Warmup steps**: 2139
- **Batch size (per device)**: 32
- **Gradient accumulation steps**: 1
- **Total train batch size**: 32
- **Language switch epoch**: 10