language, license, base_model, tags
language license base_model tags
en apache-2.0 Kotichitturu/slm-125m-base
legal
financial
small-language-model
llama
sft
raft
rag

SLM-125M SFT (RAFT — grounded)

Answers legal/financial questions only from a context passage you supply, and refuses when the answer isn't in it. Trained with RAFT (Retrieval-Augmented Fine-Tuning): 20% of training examples had no answer in the context, teaching the model that "I don't know" is a correct response.

The number that matters

A model that never refuses can post a great loss and be useless — it hallucinates exactly when you most need it not to. So refusal is measured as a confusion matrix, because refusal rate alone is gameable: a model that refuses everything scores 100% recall and is worthless.

answer ABSENT  -> refused   :  88/91   (96.7%)   <- the point of RAFT
answer ABSENT  -> answered  :   3/91   ( 3.3%)   <- hallucination
answer PRESENT -> refused   :   7/394  ( 1.8%)   <- over-refusal
answer PRESENT -> answered  : 387/394
overall accuracy: 97.9%

Both halves matter. 96.7% refusal recall would be worthless from a model that refuses everything — over-refusal is only 1.8%. It refuses when it should and answers when it should.

⚠️ What the 97.9% does NOT say — read this before trusting an answer

It measures whether the model REFUSES when it should. Not whether the answer it gives is correct. Those are different claims, and only the first is in that number. A confidently wrong answer on an answer-present row scores in the correct cell of the matrix above, leaving 97.9% untouched.

Measured separately on the 394 answer-present validation rows:

Metric Result
Exact match (normalized) 1.8%
Token F1 vs gold answer 46.3%
Numeric match (all gold numbers reproduced) 33.7%
Prediction contained no number at all 24.3%
Single-number answers — the clean test 37.0%

When exactly one figure is being asked for, this model produces it 37% of the time. No truncation excuse, no metric strictness — it fails nearly two times in three. The dominant failure is omission: it hedges into fluent, grounded- sounding, figure-free prose. Asked "what specific article of the Code of Criminal Procedure?" it replied "the article of the Code of Criminal Procedure states that the trial judge is re…" — dodging the number entirely.

Why this coexists with 96.7% refusal: refusal is a coarse binary judgement (does this passage address the question at all?). Extraction demands precision (which figure, exactly?). A 125M model learns the first and not the second.

What that means for you: trust this model to tell you when an answer isn't in the passage. Do not trust the specific figures it reports without checking them against the passage yourself. It is a triage tool, not an extraction tool.

(Caveat: gold answers are teacher-generated and judge-filtered — a proxy for truth. Some numeric misses are defensible. The 37% is not.)

Val loss and the alignment tax

Val loss 1.9551 → 0.9286 (52%), best at epoch 3. That loss is computed over response tokens only — just 8.3% of RAFT's tokens, since the context is read but never trained on. It is not comparable to the base's 2.2521, and its low value partly reflects that extractive answers are easy, not that the model is four times better.

The comparison that is fair — same held-out pretraining text the base was measured on, every token, no masking:

Model Perplexity vs base
Base 9.50
This model 10.07 +6.0%

A small, real alignment tax: it gave up 6% of its raw domain modelling to become grounded and refusal-capable. (Control: the base reproduces its published 9.51 on this harness.)

10.07 is not e^0.9286. It is e^2.3097 — this model's loss on held-out pretraining text, every token counted. Note the direction: this model has the lower val loss of the two fine-tunes (0.9286) and the higher perplexity here (10.07). Nothing is contradictory. 0.9286 is scored on 8.3% of its tokens, copying answers out of a passage sitting in the prompt; 10.07 is scored on every token of unseen legal text with nothing to copy. Two different questions wearing the same units.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Kotichitturu/slm-125m-sft-raft")
model = AutoModelForCausalLM.from_pretrained("Kotichitturu/slm-125m-sft-raft")

SYSTEM = "Answer the question based only on the provided context. If the answer is not in the context, say 'The answer is not available in the provided context.'"
user = f"Context: {passage}\n\nQuestion: {question}"
prompt = f"<|bos|><|system|>{SYSTEM}<|user|>{user}<|assistant|>"

ids = tok(prompt, return_tensors="pt", add_special_tokens=False)
out = model.generate(**ids, max_new_tokens=90, do_sample=False,
                     eos_token_id=tok.convert_tokens_to_ids("<|eos|>"))
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

Use greedy decoding — the 97.9% above was measured with do_sample=False. Sampling degrades refusal behaviour.

When the answer is absent it emits:

The answer is not available in the provided context.

Limits

  • Context + question + answer must fit 1,024 tokens (training max was 964).
  • Retrieval is your job. This model is the reader, not the retriever.
  • It is not a lawyer. Verify anything that matters.

System prompt (must match exactly)

Answer the question based only on the provided context. If the answer is not in the context, say 'The answer is not available in the provided context.'

Architecture

Identical to the base — fine-tuning changes weights, never shape.

Params 125.8M
Layers / hidden / heads 12 / 768 / 12
Vocab 16,384 (byte-level BPE)
Context 1,024

Prompt format

The chat format is not a standard template — it uses custom special tokens that exist in this tokenizer. Render exactly:

<|bos|><|system|>{system}<|user|>{user}<|assistant|>

Then generate. Stop at <|eos|>. Sending a different system prompt than the one below moves the model off-distribution and quality degrades silently.

Provenance

  • Base: Kotichitturu/slm-125m-base
  • SFT data: 8,000 passages (US case law + SEC filings), teacher-generated by gpt-4o-mini and gemini-3.1-flash-lite, then filtered by an LLM judge (gpt-5.4-mini) that rejected every answer not supported by its passage. Judge coverage 100%; keep rate 88%.
  • Decontaminated against CaseHOLD/LexGLUE (13-gram) during pretraining.
  • Full SFT (not LoRA), 3 epochs, lr 2e-5 cosine, bf16, 1×A100.
  • Loss is masked to response tokens only — the model never trains on the prompt it is given.
Description
Model synced from source: Kotichitturu/slm-125m-sft-raft
Readme 258 KiB