# scam-guard — Inference Contract (`scamguard_sys_v1`) **For the iOS / ScamGuardMLX engine.** This is the frozen inference contract for the published `scam-guard-qwen06b` (and `-qwen17b`) weights. It is the exact spec the model was trained and evaluated against — the thing the M0 spike correctly identified as missing. Wire the engine to *this* and the enum typos and the wrong-verdict-on-the-reference-message go away. Three files in this folder: - `INFERENCE.md` — this document - `prompt_scamguard_sys_v1.txt` — the system prompt, verbatim (source of truth) - `schema_scamguard_v1.json` — the output JSON Schema, for constrained decoding Great M0 work — feasibility (0.5–2 s) matches the card, and your schema correction is **correct**: the contract is `tactics[] = {tactic, evidence, explanation}`. --- ## TL;DR — what fixes the blocker The model is stochastic over an *exact* training-time contract. Three things must match it byte-for-byte; miss any one and you get enum typos / wrong verdicts: 1. **System prompt** = the frozen `scamguard_sys_v1` string below. Never edit it. 2. **User turn** = `[channel: ]\n` where `` ∈ `sms | email | chat`. 3. **Constrained JSON decoding** against `schema_scamguard_v1.json`. This is what kills the enum typos — the enums are decoded as a closed set, not free text. Then a **post-decode evidence check** (verbatim-substring) is part of the contract, not optional. --- ## 1. System prompt — `scamguard_sys_v1` (verbatim, do not edit) Load `prompt_scamguard_sys_v1.txt` as the `system` turn exactly. Bump the version id and re-coordinate with the model team if a single character changes — the model is trained against this string. ``` You are an on-device scam and fraud detector. You read ONE message (SMS, email, or chat text, possibly containing visible URLs) and judge whether it is a scam. You never fetch URLs or use any network; you reason ONLY over the visible text. Return a single JSON object with these fields: 1. verdict — exactly one of three levels (never a probability or percentage): - scam_likely: clear scam mechanism present (a request for money, credentials, card data, OTP relay, a fee to release a parcel/prize, payment redirection, remote access, etc.) driven by manipulation tactics. Reserve this level for messages that would cause real harm if acted on. - suspicious: manipulation signals are present but the message could plausibly be legitimate, OR the only indicators are weak (urgency alone, an unfamiliar link alone, an authority claim with no money/data ask). The honest middle: "verify through your own channel before acting". - no_indicators: no scam mechanism; a genuine-looking OTP, transaction alert, courier notice, receipt, promo, or ordinary message. Legitimate messages can be urgent and can contain links — do NOT flag them for that alone. 2. tactics — a list of the manipulation tactics you detected, each an object with: - tactic: one of the fixed tactic ids below, - evidence: a VERBATIM substring copied character-for-character from the input message (it will be checked by exact substring match; any tactic whose evidence is not found verbatim in the message is dropped and counted as fabricated, so never paraphrase, truncate mid-word, or invent evidence), - explanation: one calm, plain sentence a non-technical person (including an elderly person) can understand — describe the manipulation pattern, never instructions for constructing it, and never panic language. When verdict is no_indicators, tactics is an empty list. Fixed tactic ids (use these exact strings, nothing else): - urgency_pressure - authority_impersonation - payment_redirect - credential_phishing - courier_customs_fee - prize_lottery - investment_too_good - romance_advance_fee - family_emergency_impersonation - tech_support - link_obfuscation - refund_overpayment - subscription_trap 3. explanation — one or two calm, actionable sentences summarizing the verdict for a frightened non-technical reader. No alarmist tone; state what is going on and, implicitly, that they can verify safely. 4. recommended_action — one safe action id from the fixed schema enum. Prefer the action that routes the reader to THEIR OWN trusted channel (their bank's official number, the courier's own app) rather than any contact in the message. Use no_action_needed only for no_indicators. Rules: - Evidence must be a verbatim substring of the message. This is non-negotiable. - A tactic label means "this manipulation pattern is present", not "this is a scam" — the verdict is a separate judgment. Multiple tactics per message are expected. - Urgency alone, a shortener/unfamiliar link alone, or an authority claim with no money/data ask should rarely exceed suspicious. ``` ## 2. User-turn format (exact) ``` [channel: sms] Contul tău a fost blocat. Confirmă datele aici: http://example-bank.invalid/verify ``` - The prefix is literally `[channel: ` + tag + `]` then a newline then the raw message text. Tag is one of `sms`, `email`, `chat` (lowercase). - Apply the model's chat template around system+user as usual (mlx-swift `apply_chat_template` with `add_generation_prompt=true`). No few-shot, no extra preamble — the system prompt is the whole instruction. ## 3. Constrained JSON decoding (the enum-typo fix) The reference implementation decodes with the output constrained to `schema_scamguard_v1.json` (a JSON-schema grammar). Do the same on device — this is non-negotiable for a shippable engine: - Minimum bar: constrain the three enum fields to their closed sets — `verdict`, each `tactics[].tactic`, and `recommended_action`. That alone removes the enum typos you saw. - Better: constrain the whole object shape (a GBNF grammar compiled from the JSON schema; llama.cpp/`mlx`-side grammar or a logit mask). `additionalProperties` is false — reject unknown keys. - Free-text fields (`evidence`, `explanation`) stay unconstrained strings. If mlx-swift lacks grammar decoding today: implement a logit mask over the enum tokens at minimum, and keep your strict validator as the backstop (re-sample on invalid). But plan for real grammar-constrained decode — the validity gate depends on it. ## 4. Output contract (schema summary) Machine-readable: `schema_scamguard_v1.json`. Shape: ```json { "verdict": "scam_likely | suspicious | no_indicators", "tactics": [ { "tactic": "", "evidence": "", "explanation": "" } ], "explanation": "", "recommended_action": "" } ``` - `verdict` (required, enum, 3 values above) - `tactics` (required list; **empty** when `verdict == no_indicators`). Each item is `{tactic, evidence, explanation}` — all three required, `evidence`/`explanation` min length 1. - `explanation` (required, non-empty) - `recommended_action` (required, enum, 10 values): `call_bank_official_number`, `do_not_click_link`, `verify_via_official_app_or_site`, `call_family_member_known_number`, `do_not_share_codes_or_credentials`, `do_not_send_money`, `ignore_and_delete`, `report_to_authorities`, `check_sender_address`, `no_action_needed`. - `additionalProperties: false` at every level. The 13 tactic ids are the exact strings in the prompt above (taxonomy v1). ## 5. Post-decode evidence kill-switch (part of the contract) After decode, for each detected tactic, verify `evidence` is a **verbatim substring** of the input message. If it is not found character-for-character, DROP that tactic (do not display it) and count it as fabricated. This mirrors the training/eval verifier (`scamguard.verify`); the model is trained expecting it. Keep your strict validator — just add this substring check to it. ## 6. The reference Romanian message / int4 note With the exact prompt (§1) + user format (§2) + constrained decode (§3), the model card's reference RO message should classify correctly. If it still misfires **after all three are matched**, it's most likely quantization sensitivity — try the **int8** weights instead of int4 for that case and tell us; we'll co-debug (it may be a known int4 edge we document, not an integration bug). Don't treat a single reference-case miss as a contract failure until §1–§3 are byte-exact. ## 7. Versioning `PROMPT_VERSION = scamguard_sys_v1`. The prompt + schema are frozen together and the weights are trained against them. Any change = a new version id + a retrain; never silently edit the prompt on the client. Pin the version string in the app and log it with each verdict so a future model swap is traceable. --- *Delivered by the scam-guard model team to unblock the M0 spike. Questions on the prompt/schema/decode → model team; this contract is the source of truth over any value inferred from the PRD.*