8.8 KiB
scam-guard — Inference Contract (scamguard_sys_v1)
For the iOS / ScamGuardMLX engine. This is the frozen inference contract for the
published scam-guard-qwen06b (and -qwen17b) weights. It is the exact spec the
model was trained and evaluated against — the thing the M0 spike correctly
identified as missing. Wire the engine to this and the enum typos and the
wrong-verdict-on-the-reference-message go away.
Three files in this folder:
INFERENCE.md— this documentprompt_scamguard_sys_v1.txt— the system prompt, verbatim (source of truth)schema_scamguard_v1.json— the output JSON Schema, for constrained decoding
Great M0 work — feasibility (0.5–2 s) matches the card, and your schema correction
is correct: the contract is tactics[] = {tactic, evidence, explanation}.
TL;DR — what fixes the blocker
The model is stochastic over an exact training-time contract. Three things must match it byte-for-byte; miss any one and you get enum typos / wrong verdicts:
- System prompt = the frozen
scamguard_sys_v1string below. Never edit it. - User turn =
[channel: <tag>]\n<message text>where<tag>∈sms | email | chat. - Constrained JSON decoding against
schema_scamguard_v1.json. This is what kills the enum typos — the enums are decoded as a closed set, not free text.
Then a post-decode evidence check (verbatim-substring) is part of the contract, not optional.
1. System prompt — scamguard_sys_v1 (verbatim, do not edit)
Load prompt_scamguard_sys_v1.txt as the system turn exactly. Bump the version
id and re-coordinate with the model team if a single character changes — the model
is trained against this string.
You are an on-device scam and fraud detector. You read ONE message (SMS, email,
or chat text, possibly containing visible URLs) and judge whether it is a scam.
You never fetch URLs or use any network; you reason ONLY over the visible text.
Return a single JSON object with these fields:
1. verdict — exactly one of three levels (never a probability or percentage):
- scam_likely: clear scam mechanism present (a request for money, credentials,
card data, OTP relay, a fee to release a parcel/prize, payment redirection,
remote access, etc.) driven by manipulation tactics. Reserve this level for
messages that would cause real harm if acted on.
- suspicious: manipulation signals are present but the message could plausibly
be legitimate, OR the only indicators are weak (urgency alone, an unfamiliar
link alone, an authority claim with no money/data ask). The honest middle:
"verify through your own channel before acting".
- no_indicators: no scam mechanism; a genuine-looking OTP, transaction alert,
courier notice, receipt, promo, or ordinary message. Legitimate messages can
be urgent and can contain links — do NOT flag them for that alone.
2. tactics — a list of the manipulation tactics you detected, each an object with:
- tactic: one of the fixed tactic ids below,
- evidence: a VERBATIM substring copied character-for-character from the input
message (it will be checked by exact substring match; any tactic whose
evidence is not found verbatim in the message is dropped and counted as
fabricated, so never paraphrase, truncate mid-word, or invent evidence),
- explanation: one calm, plain sentence a non-technical person (including an
elderly person) can understand — describe the manipulation pattern, never
instructions for constructing it, and never panic language.
When verdict is no_indicators, tactics is an empty list.
Fixed tactic ids (use these exact strings, nothing else):
- urgency_pressure
- authority_impersonation
- payment_redirect
- credential_phishing
- courier_customs_fee
- prize_lottery
- investment_too_good
- romance_advance_fee
- family_emergency_impersonation
- tech_support
- link_obfuscation
- refund_overpayment
- subscription_trap
3. explanation — one or two calm, actionable sentences summarizing the verdict
for a frightened non-technical reader. No alarmist tone; state what is going on
and, implicitly, that they can verify safely.
4. recommended_action — one safe action id from the fixed schema enum. Prefer the
action that routes the reader to THEIR OWN trusted channel (their bank's
official number, the courier's own app) rather than any contact in the message.
Use no_action_needed only for no_indicators.
Rules:
- Evidence must be a verbatim substring of the message. This is non-negotiable.
- A tactic label means "this manipulation pattern is present", not "this is a
scam" — the verdict is a separate judgment. Multiple tactics per message are
expected.
- Urgency alone, a shortener/unfamiliar link alone, or an authority claim with no
money/data ask should rarely exceed suspicious.
2. User-turn format (exact)
[channel: sms]
Contul tău a fost blocat. Confirmă datele aici: http://example-bank.invalid/verify
- The prefix is literally
[channel:+ tag +]then a newline then the raw message text. Tag is one ofsms,email,chat(lowercase). - Apply the model's chat template around system+user as usual (mlx-swift
apply_chat_templatewithadd_generation_prompt=true). No few-shot, no extra preamble — the system prompt is the whole instruction.
3. Constrained JSON decoding (the enum-typo fix)
The reference implementation decodes with the output constrained to
schema_scamguard_v1.json (a JSON-schema grammar). Do the same on device — this is
non-negotiable for a shippable engine:
- Minimum bar: constrain the three enum fields to their closed sets —
verdict, eachtactics[].tactic, andrecommended_action. That alone removes the enum typos you saw. - Better: constrain the whole object shape (a GBNF grammar compiled from the JSON
schema; llama.cpp/
mlx-side grammar or a logit mask).additionalPropertiesis false — reject unknown keys. - Free-text fields (
evidence,explanation) stay unconstrained strings.
If mlx-swift lacks grammar decoding today: implement a logit mask over the enum tokens at minimum, and keep your strict validator as the backstop (re-sample on invalid). But plan for real grammar-constrained decode — the validity gate depends on it.
4. Output contract (schema summary)
Machine-readable: schema_scamguard_v1.json. Shape:
{
"verdict": "scam_likely | suspicious | no_indicators",
"tactics": [
{ "tactic": "<one of the 13 ids>", "evidence": "<verbatim substring>", "explanation": "<plain sentence>" }
],
"explanation": "<overall plain-language explanation>",
"recommended_action": "<one of the 10 safe-action ids>"
}
verdict(required, enum, 3 values above)tactics(required list; empty whenverdict == no_indicators). Each item is{tactic, evidence, explanation}— all three required,evidence/explanationmin length 1.explanation(required, non-empty)recommended_action(required, enum, 10 values):call_bank_official_number,do_not_click_link,verify_via_official_app_or_site,call_family_member_known_number,do_not_share_codes_or_credentials,do_not_send_money,ignore_and_delete,report_to_authorities,check_sender_address,no_action_needed.additionalProperties: falseat every level.
The 13 tactic ids are the exact strings in the prompt above (taxonomy v1).
5. Post-decode evidence kill-switch (part of the contract)
After decode, for each detected tactic, verify evidence is a verbatim
substring of the input message. If it is not found character-for-character, DROP
that tactic (do not display it) and count it as fabricated. This mirrors the
training/eval verifier (scamguard.verify); the model is trained expecting it.
Keep your strict validator — just add this substring check to it.
6. The reference Romanian message / int4 note
With the exact prompt (§1) + user format (§2) + constrained decode (§3), the model card's reference RO message should classify correctly. If it still misfires after all three are matched, it's most likely quantization sensitivity — try the int8 weights instead of int4 for that case and tell us; we'll co-debug (it may be a known int4 edge we document, not an integration bug). Don't treat a single reference-case miss as a contract failure until §1–§3 are byte-exact.
7. Versioning
PROMPT_VERSION = scamguard_sys_v1. The prompt + schema are frozen together and the
weights are trained against them. Any change = a new version id + a retrain; never
silently edit the prompt on the client. Pin the version string in the app and log it
with each verdict so a future model swap is traceable.
Delivered by the scam-guard model team to unblock the M0 spike. Questions on the prompt/schema/decode → model team; this contract is the source of truth over any value inferred from the PRD.