tags, library_name
tags library_name
causal-lm
text-generation
transformers

nikitastheo/babylm-lem-spa-ell-sequential_interleaved

Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).

Training details

  • Base config: gpt_base_config.json
  • Tokenizer: nikitastheo/babylm-lem-spa-ell-tokenizer
  • Max steps: 25460
  • Learning rate: 0.0001
  • LR scheduler: linear
  • Warmup steps: 2546
  • Batch size (per device): 32
  • Gradient accumulation steps: 1
  • Total train batch size: 32
  • Language switch epoch: 10
Description
Model synced from source: nikitastheo/babylm-lem-spa-ell-sequential_interleaved
Readme 946 KiB