初始化项目,由ModelHub XC社区提供模型
Model: asinha08/slm-125m-sft-pilot-2k Source: Original Platform
This commit is contained in:
91
README.md
Normal file
91
README.md
Normal file
@@ -0,0 +1,91 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: thesreedath/slm-125m-base
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- llama
|
||||
- causal-lm
|
||||
- sft
|
||||
- legal
|
||||
- financial
|
||||
- small-language-model
|
||||
---
|
||||
|
||||
# SLM 125M SFT Pilot 2K
|
||||
|
||||
This is a supervised fine-tuned 125M-parameter causal language model based on
|
||||
[`thesreedath/slm-125m-base`](https://huggingface.co/thesreedath/slm-125m-base).
|
||||
|
||||
## Training Summary
|
||||
|
||||
| Field | Value |
|
||||
|---|---:|
|
||||
| SFT dataset run | `pilot-2k-v2` |
|
||||
| Train examples | 1800 |
|
||||
| Validation examples | 100 |
|
||||
| Epochs | 3 |
|
||||
| Optimizer steps | 339 |
|
||||
| Train tokens per epoch | 1262981 |
|
||||
| Assistant-label tokens per epoch | 62190 |
|
||||
| Total train-token exposures | 3788943 |
|
||||
| Total assistant-label token exposures | 186570 |
|
||||
| Initial validation loss | 3.6077 |
|
||||
| Initial validation perplexity | 36.88 |
|
||||
| Final validation loss | 1.5028 |
|
||||
| Final validation perplexity | 4.49 |
|
||||
| Training elapsed seconds | 160.1 |
|
||||
| Estimated GPU-only training cost | $0.0355 |
|
||||
|
||||
## Dataset
|
||||
|
||||
The SFT dataset contains 2,000 generated and judged instruction examples:
|
||||
|
||||
| Split | Examples |
|
||||
|---|---:|
|
||||
| Train | 1,800 |
|
||||
| Validation | 100 |
|
||||
| Test | 100 |
|
||||
|
||||
Source mix:
|
||||
|
||||
| Source | Examples |
|
||||
|---|---:|
|
||||
| Case law | 703 |
|
||||
| SEC filings | 890 |
|
||||
| FineWeb-Edu | 407 |
|
||||
|
||||
Task types include grounded question answering, unanswerable/refusal cases,
|
||||
multi-step QA, JSON extraction, summarization, plain-English rewrites, and
|
||||
comparative QA.
|
||||
|
||||
## Intended Use
|
||||
|
||||
This is a small research/experimentation model for legal and financial
|
||||
instruction-following behavior. It is not a substitute for professional legal,
|
||||
financial, or compliance advice.
|
||||
|
||||
## Prompt Format
|
||||
|
||||
The model was trained with the tokenizer/chat markers from the base model:
|
||||
|
||||
```text
|
||||
<|system|>
|
||||
You are a careful legal and financial assistant...
|
||||
<|user|>
|
||||
Context:
|
||||
...
|
||||
|
||||
Question:
|
||||
...
|
||||
<|assistant|>
|
||||
...
|
||||
<|eos|>
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- The SFT dataset is small, so behavior should be evaluated carefully.
|
||||
- Answers should be grounded in supplied context.
|
||||
- The model may hallucinate if used without retrieval/context.
|
||||
- Perplexity is measured on the generated validation split, not a broad external benchmark.
|
||||
Reference in New Issue
Block a user