Model: Kotichitturu/slm-125m-sft-raft Source: Original Platform
language, license, base_model, tags
| language | license | base_model | tags | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| en | apache-2.0 | Kotichitturu/slm-125m-base |
|
SLM-125M SFT (RAFT — grounded)
Answers legal/financial questions only from a context passage you supply, and refuses when the answer isn't in it. Trained with RAFT (Retrieval-Augmented Fine-Tuning): 20% of training examples had no answer in the context, teaching the model that "I don't know" is a correct response.
The number that matters
A model that never refuses can post a great loss and be useless — it hallucinates exactly when you most need it not to. So refusal is measured as a confusion matrix, because refusal rate alone is gameable: a model that refuses everything scores 100% recall and is worthless.
answer ABSENT -> refused : 88/91 (96.7%) <- the point of RAFT
answer ABSENT -> answered : 3/91 ( 3.3%) <- hallucination
answer PRESENT -> refused : 7/394 ( 1.8%) <- over-refusal
answer PRESENT -> answered : 387/394
overall accuracy: 97.9%
Both halves matter. 96.7% refusal recall would be worthless from a model that refuses everything — over-refusal is only 1.8%. It refuses when it should and answers when it should.
⚠️ What the 97.9% does NOT say — read this before trusting an answer
It measures whether the model REFUSES when it should. Not whether the answer it gives is correct. Those are different claims, and only the first is in that number. A confidently wrong answer on an answer-present row scores in the correct cell of the matrix above, leaving 97.9% untouched.
Measured separately on the 394 answer-present validation rows:
| Metric | Result |
|---|---|
| Exact match (normalized) | 1.8% |
| Token F1 vs gold answer | 46.3% |
| Numeric match (all gold numbers reproduced) | 33.7% |
| Prediction contained no number at all | 24.3% |
| Single-number answers — the clean test | 37.0% |
When exactly one figure is being asked for, this model produces it 37% of the time. No truncation excuse, no metric strictness — it fails nearly two times in three. The dominant failure is omission: it hedges into fluent, grounded- sounding, figure-free prose. Asked "what specific article of the Code of Criminal Procedure?" it replied "the article of the Code of Criminal Procedure states that the trial judge is re…" — dodging the number entirely.
Why this coexists with 96.7% refusal: refusal is a coarse binary judgement (does this passage address the question at all?). Extraction demands precision (which figure, exactly?). A 125M model learns the first and not the second.
What that means for you: trust this model to tell you when an answer isn't in the passage. Do not trust the specific figures it reports without checking them against the passage yourself. It is a triage tool, not an extraction tool.
(Caveat: gold answers are teacher-generated and judge-filtered — a proxy for truth. Some numeric misses are defensible. The 37% is not.)
Val loss and the alignment tax
Val loss 1.9551 → 0.9286 (−52%), best at epoch 3. That loss is computed over response tokens only — just 8.3% of RAFT's tokens, since the context is read but never trained on. It is not comparable to the base's 2.2521, and its low value partly reflects that extractive answers are easy, not that the model is four times better.
The comparison that is fair — same held-out pretraining text the base was measured on, every token, no masking:
| Model | Perplexity | vs base |
|---|---|---|
| Base | 9.50 | — |
| This model | 10.07 | +6.0% |
A small, real alignment tax: it gave up 6% of its raw domain modelling to become grounded and refusal-capable. (Control: the base reproduces its published 9.51 on this harness.)
10.07 is not e^0.9286. It is e^2.3097 — this model's loss on held-out
pretraining text, every token counted. Note the direction: this model has the
lower val loss of the two fine-tunes (0.9286) and the higher perplexity here
(10.07). Nothing is contradictory. 0.9286 is scored on 8.3% of its tokens, copying
answers out of a passage sitting in the prompt; 10.07 is scored on every token of
unseen legal text with nothing to copy. Two different questions wearing the same units.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Kotichitturu/slm-125m-sft-raft")
model = AutoModelForCausalLM.from_pretrained("Kotichitturu/slm-125m-sft-raft")
SYSTEM = "Answer the question based only on the provided context. If the answer is not in the context, say 'The answer is not available in the provided context.'"
user = f"Context: {passage}\n\nQuestion: {question}"
prompt = f"<|bos|><|system|>{SYSTEM}<|user|>{user}<|assistant|>"
ids = tok(prompt, return_tensors="pt", add_special_tokens=False)
out = model.generate(**ids, max_new_tokens=90, do_sample=False,
eos_token_id=tok.convert_tokens_to_ids("<|eos|>"))
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))
Use greedy decoding — the 97.9% above was measured with do_sample=False.
Sampling degrades refusal behaviour.
When the answer is absent it emits:
The answer is not available in the provided context.
Limits
- Context + question + answer must fit 1,024 tokens (training max was 964).
- Retrieval is your job. This model is the reader, not the retriever.
- It is not a lawyer. Verify anything that matters.
System prompt (must match exactly)
Answer the question based only on the provided context. If the answer is not in the context, say 'The answer is not available in the provided context.'
Architecture
Identical to the base — fine-tuning changes weights, never shape.
| Params | 125.8M |
| Layers / hidden / heads | 12 / 768 / 12 |
| Vocab | 16,384 (byte-level BPE) |
| Context | 1,024 |
Prompt format
The chat format is not a standard template — it uses custom special tokens that exist in this tokenizer. Render exactly:
<|bos|><|system|>{system}<|user|>{user}<|assistant|>
Then generate. Stop at <|eos|>. Sending a different system prompt than the one
below moves the model off-distribution and quality degrades silently.
Provenance
- Base: Kotichitturu/slm-125m-base
- SFT data: 8,000 passages (US case law + SEC filings), teacher-generated by gpt-4o-mini and gemini-3.1-flash-lite, then filtered by an LLM judge (gpt-5.4-mini) that rejected every answer not supported by its passage. Judge coverage 100%; keep rate 88%.
- Decontaminated against CaseHOLD/LexGLUE (13-gram) during pretraining.
- Full SFT (not LoRA), 3 epochs, lr 2e-5 cosine, bf16, 1×A100.
- Loss is masked to response tokens only — the model never trains on the prompt it is given.