Files
babylm-afr-ell-sequential_i…/README.md
ModelHub XC bfd417ce07 初始化项目,由ModelHub XC社区提供模型
Model: nikitastheo/babylm-afr-ell-sequential_interleaved
Source: Original Platform
2026-07-20 11:04:22 +08:00

580 B

tags, library_name
tags library_name
causal-lm
text-generation
transformers

nikitastheo/babylm-afr-ell-sequential_interleaved

Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).

Training details

  • Base config: gpt_base_config.json
  • Tokenizer: nikitastheo/babylm-afr-ell-tokenizer
  • Max steps: 24270
  • Learning rate: 0.0001
  • LR scheduler: linear
  • Warmup steps: 2427
  • Batch size (per device): 32
  • Gradient accumulation steps: 1
  • Total train batch size: 32
  • Language switch epoch: 10