Files
txn-parser/smollm2-360m
ModelHub XC f70fa31854 初始化项目,由ModelHub XC社区提供模型
Model: kartikey31/txn-parser
Source: Original Platform
2026-08-09 01:23:21 +08:00
..

license, base_model, tags, language, library_name
license base_model tags language library_name
apache-2.0 HuggingFaceTB/SmolLM2-360M-Instruct
text-generation
lora
qlora
gguf
transaction-parser
en
hi
peft

txn-parser / smollm2-360m

QLoRA fine-tune of HuggingFaceTB/SmolLM2-360M-Instruct for extracting structured transaction data (amount, currency, item, category, type) from free-form Indian-English / code-switched speech and text.

This model lives in subfolder smollm2-360m/ of the kartikey31/txn-parser repo, alongside sibling fine-tunes of other base models trained on the same data.

What's in here

Training data

  • 93,348 teacher-labeled examples (data/distill/train.jsonl)
  • 300 held-out eval examples (data/distill/eval.jsonl)
  • Validator-gated: every row's output passes the project's grammar + amount-parser semantic validator.

Training config

Knob Value
Base model HuggingFaceTB/SmolLM2-360M-Instruct
Method QLoRA (4-bit) via Unsloth
LoRA rank 32 (alpha 64, dropout 0.0)
Epochs 2
Batch size (train) 64
Grad accumulation 1
Eval batch size 16
Max seq length 1024
Learning rate 2e-4 (warmup 3%)
Started 2026-05-22T21:26:43.776899+00:00
Finished 2026-05-22T21:30:14.660709+00:00

System prompt (use this EXACTLY)

The model was trained with one specific system prompt and Gemma/Smol/Qwen chat template. If you paraphrase the prompt or skip the chat template, quality degrades quickly. Copy-paste this verbatim into your inference client (no leading/trailing whitespace, no edits):

You convert voice-transcribed transaction descriptions into structured JSON.

Output ONLY a JSON object with this schema, no other text:
{"transactions":[{"amount":<number>,"currency":"INR"|"USD","item":"<lowercase singular noun phrase>","category":"<enum>","type":"expense"|"income"}]}

Categories: Food, Drinks, Groceries, Transport, Shopping, Entertainment, Bills, Health, Education, Personal, Gifts, Income, Other.

Rules:
- Currency defaults to INR. Use USD only when the input explicitly says "dollars" or contains "$".
- Amounts: "k" = ×1000, "hazaar" = ×1000, "sau" = ×100, "lakh" = ×100000. Convert number-words ("five hundred") to digits.
- type is "expense" by default; "income" only for explicit salary, cashback, refund, gift received, payment received.
- For disfluencies and corrections ("500 wait no 600"), output the CORRECTED amount only.
- For ambiguous items ("that thing", "stuff"), use item "unspecified" and category "Other".
- Item field: lowercase singular noun phrase ("uber ride", "beer", "chai" — not "Beers" or "Uber").
- Multi-transaction inputs become multiple array entries in spoken order.
- Category heuristics: uber/ola/auto/petrol/bus/metro → Transport; beer/wine/chai/coffee/juice → Drinks; rent/electricity/wifi/recharge/gas → Bills; movie/netflix/concert → Entertainment; doctor/medicine/hospital → Health.

Source of truth: scripts/_lib.py constant SYSTEM_PROMPT. Don't retype it — pull from _lib.py or this README.

Download a single GGUF

huggingface-cli download kartikey31/txn-parser \
    smollm2-360m/gguf/txn-parser-smollm2-360m-Q4_K_M.gguf \
    --local-dir .

Inference (Python, llama-cpp-python)

from llama_cpp import Llama

SYSTEM_PROMPT = '''You convert voice-transcribed transaction descriptions into structured JSON.

Output ONLY a JSON object with this schema, no other text:
{"transactions":[{"amount":<number>,"currency":"INR"|"USD","item":"<lowercase singular noun phrase>","category":"<enum>","type":"expense"|"income"}]}

Categories: Food, Drinks, Groceries, Transport, Shopping, Entertainment, Bills, Health, Education, Personal, Gifts, Income, Other.

Rules:
- Currency defaults to INR. Use USD only when the input explicitly says "dollars" or contains "$".
- Amounts: "k" = ×1000, "hazaar" = ×1000, "sau" = ×100, "lakh" = ×100000. Convert number-words ("five hundred") to digits.
- type is "expense" by default; "income" only for explicit salary, cashback, refund, gift received, payment received.
- For disfluencies and corrections ("500 wait no 600"), output the CORRECTED amount only.
- For ambiguous items ("that thing", "stuff"), use item "unspecified" and category "Other".
- Item field: lowercase singular noun phrase ("uber ride", "beer", "chai" — not "Beers" or "Uber").
- Multi-transaction inputs become multiple array entries in spoken order.
- Category heuristics: uber/ola/auto/petrol/bus/metro → Transport; beer/wine/chai/coffee/juice → Drinks; rent/electricity/wifi/recharge/gas → Bills; movie/netflix/concert → Entertainment; doctor/medicine/hospital → Health.'''

llm = Llama(
    model_path="txn-parser-smollm2-360m-Q4_K_M.gguf",
    n_gpu_layers=-1, n_ctx=2048,
)
out = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user",   "content": "200 ka samosa"},
    ],
    temperature=0.0,
)
print(out["choices"][0]["message"]["content"])

Inference (CLI, llama.cpp)

./llama-cli -m txn-parser-smollm2-360m-Q4_K_M.gguf \
    --grammar-file scripts/grammar.gbnf \
    --system-prompt "$(cat system_prompt.txt)" \
    -p "200 ka samosa" -n 256

Reproduce

git clone https://github.com/kartikeychoudhary/txn-parser.git
cd txn-parser && bash setup.sh
python scripts/train_and_publish.py --only smollm2-360m

Auto-published by scripts/train_and_publish.py on 2026-05-22T21:30:14.964463+00:00.