Model: Raghav-Singhal/pretrain-normal-smollm-1p7b-100B-20n-2048sl-960gbsz Source: Original Platform
22 lines
495 B
Markdown
22 lines
495 B
Markdown
---
|
|
library_name: transformers
|
|
pipeline_tag: text-generation
|
|
tags:
|
|
- llama
|
|
- causal-lm
|
|
- bfloat16
|
|
---
|
|
|
|
# pretrain-normal-smollm-1p7b-100B-20n-2048sl-960gbsz
|
|
|
|
Converted Hugging Face base checkpoint from the Model Raising pretraining run.
|
|
|
|
## Details
|
|
|
|
- Architecture: `LlamaForCausalLM`
|
|
- Base model size: `1.7B`
|
|
- Precision on disk: `bfloat16`
|
|
- Tokenizer: `HuggingFaceTB/SmolLM2-1.7B-Instruct`
|
|
|
|
This repo contains the verified Hugging Face export of the final pretraining checkpoint before SFT.
|