54b90a2566841c388d6521ca1fa72c30133881bf
Model: nikitastheo/mixed-lem-ell-ell-sequential_interleaved Source: Original Platform
tags, library_name
| tags | library_name | ||
|---|---|---|---|
|
transformers |
nikitastheo/mixed-lem-ell-ell-sequential_interleaved
Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).
Training details
- Base config:
gpt_base_config.json - Tokenizer:
nikitastheo/mixed-lem-ell-ell-tokenizer - Max steps: 24280
- Learning rate: 0.0001
- LR scheduler: linear
- Warmup steps: 2428
- Batch size (per device): 32
- Gradient accumulation steps: 1
- Total train batch size: 32
- Language switch epoch: 10
Description