Files
legal-slm-125m-sft/README.md
ModelHub XC 25bebbe832 初始化项目,由ModelHub XC社区提供模型
Model: jonam-ai/legal-slm-125m-sft
Source: Original Platform
2026-08-08 17:51:17 +08:00

104 lines
3.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: thesreedath/slm-125m-base
tags:
- llama
- legal
- financial
- small-language-model
- sft
- instruction-tuning
---
# legal-slm-125m-sft
A 125M-parameter Llama-architecture model **supervised-fine-tuned to answer
legal & financial questions**. It is the instruction-tuned sibling of a
from-scratch base model: where the base model only *continues* text, this model
*responds* to a question.
- **Base model:** [`thesreedath/slm-125m-base`](https://huggingface.co/thesreedath/slm-125m-base) (125M, pretrained 10 epochs)
- **Fine-tuned on:** 5,846 grounded legal/financial Q&A pairs (synthetic, quality-filtered)
- **Validation loss:** 2.06 (cross-entropy on held-out Q&A)
> **Base-capacity model.** At 125M parameters this is a study in end-to-end LLM
> engineering, not a production assistant. It learns the *shape* of a good answer
> and often gets the gist right, but it will **fabricate** case names, statutes,
> and figures. **Never** use its output as legal, financial, or factual advice.
## What it does
The base model, given "In a breach of contract claim, the plaintiff…", would ramble
onward. This model, asked a question, answers it:
> **Q:** In a breach of contract claim, what must the plaintiff prove?
> **A:** The plaintiff must show: (1) a contract; (2) an agreement to perform a
> specific act; (3) an obligation to perform…
## Chat format
The tokenizer has role special tokens but **no chat template**, so format prompts
manually as `<|bos|><|system|>{system}<|user|>{question}<|assistant|>` and let the
model complete the answer up to `<|eos|>`:
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("jonam-ai/legal-slm-125m-sft")
model = AutoModelForCausalLM.from_pretrained("jonam-ai/legal-slm-125m-sft", torch_dtype=torch.bfloat16)
system = "You are a knowledgeable legal and financial assistant. Answer accurately and concisely."
question = "What is the purpose of a Form 10-K filing?"
def sid(t): return tok.convert_tokens_to_ids(t)
ids = (tok("<|bos|>", add_special_tokens=False)["input_ids"]
+ [sid("<|system|>")] + tok(system, add_special_tokens=False)["input_ids"]
+ [sid("<|user|>")] + tok(question, add_special_tokens=False)["input_ids"]
+ [sid("<|assistant|>")])
out = model.generate(torch.tensor([ids]), max_new_tokens=120, do_sample=True,
temperature=0.7, top_p=0.9, eos_token_id=sid("<|eos|>"),
pad_token_id=sid("<|pad|>"))
print(tok.decode(out[0][len(ids):], skip_special_tokens=True))
```
## Training data
5,846 grounded question–answer pairs, synthesized with a **teacher-LLM
distillation** pipeline over a cleaned corpus of US case law, SEC filings, and
educational web text:
1. **Chunk** the corpus into ~800-token passages.
2. **Generate** (Gemini Flash-Lite): write self-contained Q&A answerable *only*
from the passage, balanced across task types (QA, extraction, summarization,
rewrite) and difficulty (easy → hard).
3. **Judge** (Gemini Flash as LLM-as-judge): keep only pairs that are grounded,
correct, and self-contained (score ≥ 4/5). ~78% pass.
4. **Dedup** (exact + MinHash-LSH near-duplicate removal).
5. **Format** as chat and **tokenize** with the base model's own tokenizer, with
**loss computed only on the answer tokens**.
Mix: case-law 45% · SEC 45% · web 10%. Split: 5,554 train / 292 val.
## Training
| | |
|---|---|
| Method | full fine-tune (not LoRA) |
| Hardware | 1 × NVIDIA L4 |
| Epochs | 2 |
| Tokens seen | ~1.0M |
| LR | 3e-5, cosine decay, 3% warmup |
| Precision | bf16 compute, fp32 master weights |
| Time | ~80 seconds |
## Limitations
125M parameters, English only, 1,024-token context, not aligned/RLHF'd. It imitates
the form of grounded answers learned from a synthetic dataset; factual reliability
is limited by model size. Not legal or financial advice.