588 B
588 B
tags, library_name
| tags | library_name | ||
|---|---|---|---|
|
transformers |
nikitastheo/babylm-lem-spa-ell-sequential_interleaved
Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).
Training details
- Base config:
gpt_base_config.json - Tokenizer:
nikitastheo/babylm-lem-spa-ell-tokenizer - Max steps: 25460
- Learning rate: 0.0001
- LR scheduler: linear
- Warmup steps: 2546
- Batch size (per device): 32
- Gradient accumulation steps: 1
- Total train batch size: 32
- Language switch epoch: 10