初始化项目,由ModelHub XC社区提供模型
Model: nikitastheo/babylm-lem-spa-ell-sequential_interleaved Source: Original Platform
This commit is contained in:
24
README.md
Normal file
24
README.md
Normal file
@@ -0,0 +1,24 @@
|
||||
---
|
||||
tags:
|
||||
- causal-lm
|
||||
- text-generation
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# nikitastheo/babylm-lem-spa-ell-sequential_interleaved
|
||||
|
||||
Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no `Trainer`).
|
||||
|
||||
## Training details
|
||||
|
||||
- **Base config**: `gpt_base_config.json`
|
||||
- **Tokenizer**: `nikitastheo/babylm-lem-spa-ell-tokenizer`
|
||||
- **Max steps**: 25460
|
||||
- **Learning rate**: 0.0001
|
||||
- **LR scheduler**: linear
|
||||
- **Warmup steps**: 2546
|
||||
- **Batch size (per device)**: 32
|
||||
- **Gradient accumulation steps**: 1
|
||||
- **Total train batch size**: 32
|
||||
- **Language switch epoch**: 10
|
||||
|
||||
Reference in New Issue
Block a user