初始化项目,由ModelHub XC社区提供模型
Model: karthiksab/slm-125m-base Source: Original Platform
This commit is contained in:
31
README.md
Normal file
31
README.md
Normal file
@@ -0,0 +1,31 @@
|
||||
---
|
||||
language:
|
||||
- en
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- causal-lm
|
||||
- legal
|
||||
- finance
|
||||
- llama
|
||||
---
|
||||
|
||||
# slm-125m-base
|
||||
|
||||
125.8M-parameter decoder-only language model trained from scratch for legal,
|
||||
financial, and educational English text.
|
||||
|
||||
## Training
|
||||
|
||||
- Architecture: 12-layer Llama-compatible decoder, hidden size 768, vocabulary 16,384
|
||||
- Context length: 1,024 tokens
|
||||
- Training data: 2.034B packed training tokens; 20.6M validation tokens
|
||||
- Data mix: cleaned and deduplicated US case law, SEC filings, and FineWeb-Edu
|
||||
- Optimization: AdamW, BF16, cosine learning-rate schedule, one epoch
|
||||
- Hardware: 8 NVIDIA H100 GPUs
|
||||
- Final validation loss: 2.3981
|
||||
- Final validation perplexity: 11.00
|
||||
- Reported GPU compute cost: approximately $10.60
|
||||
|
||||
This is a base model, not an instruction-tuned or safety-tuned assistant. Outputs
|
||||
may be inaccurate and should not be treated as legal or financial advice.
|
||||
Reference in New Issue
Block a user