初始化项目,由ModelHub XC社区提供模型
Model: ppanja/slm-125m-base Source: Original Platform
This commit is contained in:
62
README.md
Normal file
62
README.md
Normal file
@@ -0,0 +1,62 @@
|
||||
---
|
||||
language: en
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- causal-lm
|
||||
- llama
|
||||
- legal
|
||||
- financial
|
||||
- from-scratch
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# slm-125m
|
||||
|
||||
125M-parameter causal language model trained from scratch on a legal/financial corpus.
|
||||
|
||||
## Model
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Architecture | Llama-style (SwiGLU, RoPE, RMSNorm) |
|
||||
| Parameters | ~125.8M |
|
||||
| Layers / dim / heads | 12 / 768 / 12 |
|
||||
| Context | 1024 tokens |
|
||||
| Vocab | 16384 (byte-level BPE) |
|
||||
|
||||
## Training data
|
||||
|
||||
| Source | Role |
|
||||
|--------|------|
|
||||
| HFforLegal/case-law | US case law |
|
||||
| PleIAs/SEC | SEC filings |
|
||||
| HuggingFaceFW/fineweb-edu | General educational web text |
|
||||
|
||||
Packed train tokens: **unknown** | val tokens: **unknown**
|
||||
|
||||
Eval benchmarks (LexGLUE, CaseHOLD) were held out during corpus construction.
|
||||
|
||||
## Training run
|
||||
|
||||
| Hyperparameter | Value |
|
||||
|----------------|-------|
|
||||
| Global batch (tokens) | 524,288 |
|
||||
| LR / min LR | 0.0006 / 6e-05 |
|
||||
| Warmup tokens | 200M |
|
||||
| Weight decay | 0.1 |
|
||||
|
||||
Last logged train loss: **n/a** | val loss: **n/a**
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
tok = AutoTokenizer.from_pretrained("ppanja/slm-125m-base")
|
||||
model = AutoModelForCausalLM.from_pretrained("ppanja/slm-125m-base")
|
||||
inputs = tok("The plaintiff shall bear the burden of proof", return_tensors="pt")
|
||||
out = model.generate(**inputs, max_new_tokens=64)
|
||||
print(tok.decode(out[0]))
|
||||
```
|
||||
|
||||
Chat special tokens (`<|user|>`, `<|assistant|>`, `<|system|>`) are in the vocabulary for future instruction tuning.
|
||||
Reference in New Issue
Block a user