Files
legal-slm-125m-sft/README.md
ModelHub XC 25bebbe832 初始化项目,由ModelHub XC社区提供模型
Model: jonam-ai/legal-slm-125m-sft
Source: Original Platform
2026-08-08 17:51:17 +08:00

3.9 KiB
Raw Blame History

license, language, library_name, pipeline_tag, base_model, tags
license language library_name pipeline_tag base_model tags
apache-2.0
en
transformers text-generation thesreedath/slm-125m-base
llama
legal
financial
small-language-model
sft
instruction-tuning

legal-slm-125m-sft

A 125M-parameter Llama-architecture model supervised-fine-tuned to answer legal & financial questions. It is the instruction-tuned sibling of a from-scratch base model: where the base model only continues text, this model responds to a question.

  • Base model: thesreedath/slm-125m-base (125M, pretrained 10 epochs)
  • Fine-tuned on: 5,846 grounded legal/financial Q&A pairs (synthetic, quality-filtered)
  • Validation loss: 2.06 (cross-entropy on held-out Q&A)

Base-capacity model. At 125M parameters this is a study in end-to-end LLM engineering, not a production assistant. It learns the shape of a good answer and often gets the gist right, but it will fabricate case names, statutes, and figures. Never use its output as legal, financial, or factual advice.

What it does

The base model, given "In a breach of contract claim, the plaintiff…", would ramble onward. This model, asked a question, answers it:

Q: In a breach of contract claim, what must the plaintiff prove? A: The plaintiff must show: (1) a contract; (2) an agreement to perform a specific act; (3) an obligation to perform…

Chat format

The tokenizer has role special tokens but no chat template, so format prompts manually as <|bos|><|system|>{system}<|user|>{question}<|assistant|> and let the model complete the answer up to <|eos|>:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("jonam-ai/legal-slm-125m-sft")
model = AutoModelForCausalLM.from_pretrained("jonam-ai/legal-slm-125m-sft", torch_dtype=torch.bfloat16)

system = "You are a knowledgeable legal and financial assistant. Answer accurately and concisely."
question = "What is the purpose of a Form 10-K filing?"

def sid(t): return tok.convert_tokens_to_ids(t)
ids = (tok("<|bos|>", add_special_tokens=False)["input_ids"]
       + [sid("<|system|>")] + tok(system, add_special_tokens=False)["input_ids"]
       + [sid("<|user|>")]   + tok(question, add_special_tokens=False)["input_ids"]
       + [sid("<|assistant|>")])
out = model.generate(torch.tensor([ids]), max_new_tokens=120, do_sample=True,
                     temperature=0.7, top_p=0.9, eos_token_id=sid("<|eos|>"),
                     pad_token_id=sid("<|pad|>"))
print(tok.decode(out[0][len(ids):], skip_special_tokens=True))

Training data

5,846 grounded question–answer pairs, synthesized with a teacher-LLM distillation pipeline over a cleaned corpus of US case law, SEC filings, and educational web text:

  1. Chunk the corpus into ~800-token passages.
  2. Generate (Gemini Flash-Lite): write self-contained Q&A answerable only from the passage, balanced across task types (QA, extraction, summarization, rewrite) and difficulty (easy → hard).
  3. Judge (Gemini Flash as LLM-as-judge): keep only pairs that are grounded, correct, and self-contained (score ≥ 4/5). ~78% pass.
  4. Dedup (exact + MinHash-LSH near-duplicate removal).
  5. Format as chat and tokenize with the base model's own tokenizer, with loss computed only on the answer tokens.

Mix: case-law 45% · SEC 45% · web 10%. Split: 5,554 train / 292 val.

Training

Method full fine-tune (not LoRA)
Hardware 1 × NVIDIA L4
Epochs 2
Tokens seen ~1.0M
LR 3e-5, cosine decay, 3% warmup
Precision bf16 compute, fp32 master weights
Time ~80 seconds

Limitations

125M parameters, English only, 1,024-token context, not aligned/RLHF'd. It imitates the form of grounded answers learned from a synthetic dataset; factual reliability is limited by model size. Not legal or financial advice.