194 lines
8.8 KiB
Markdown
194 lines
8.8 KiB
Markdown
# scam-guard — Inference Contract (`scamguard_sys_v1`)
|
||
|
||
**For the iOS / ScamGuardMLX engine.** This is the frozen inference contract for the
|
||
published `scam-guard-qwen06b` (and `-qwen17b`) weights. It is the exact spec the
|
||
model was trained and evaluated against — the thing the M0 spike correctly
|
||
identified as missing. Wire the engine to *this* and the enum typos and the
|
||
wrong-verdict-on-the-reference-message go away.
|
||
|
||
Three files in this folder:
|
||
- `INFERENCE.md` — this document
|
||
- `prompt_scamguard_sys_v1.txt` — the system prompt, verbatim (source of truth)
|
||
- `schema_scamguard_v1.json` — the output JSON Schema, for constrained decoding
|
||
|
||
Great M0 work — feasibility (0.5–2 s) matches the card, and your schema correction
|
||
is **correct**: the contract is `tactics[] = {tactic, evidence, explanation}`.
|
||
|
||
---
|
||
|
||
## TL;DR — what fixes the blocker
|
||
|
||
The model is stochastic over an *exact* training-time contract. Three things must
|
||
match it byte-for-byte; miss any one and you get enum typos / wrong verdicts:
|
||
|
||
1. **System prompt** = the frozen `scamguard_sys_v1` string below. Never edit it.
|
||
2. **User turn** = `[channel: <tag>]\n<message text>` where `<tag>` ∈ `sms | email | chat`.
|
||
3. **Constrained JSON decoding** against `schema_scamguard_v1.json`. This is what
|
||
kills the enum typos — the enums are decoded as a closed set, not free text.
|
||
|
||
Then a **post-decode evidence check** (verbatim-substring) is part of the contract,
|
||
not optional.
|
||
|
||
---
|
||
|
||
## 1. System prompt — `scamguard_sys_v1` (verbatim, do not edit)
|
||
|
||
Load `prompt_scamguard_sys_v1.txt` as the `system` turn exactly. Bump the version
|
||
id and re-coordinate with the model team if a single character changes — the model
|
||
is trained against this string.
|
||
|
||
```
|
||
You are an on-device scam and fraud detector. You read ONE message (SMS, email,
|
||
or chat text, possibly containing visible URLs) and judge whether it is a scam.
|
||
You never fetch URLs or use any network; you reason ONLY over the visible text.
|
||
|
||
Return a single JSON object with these fields:
|
||
|
||
1. verdict — exactly one of three levels (never a probability or percentage):
|
||
- scam_likely: clear scam mechanism present (a request for money, credentials,
|
||
card data, OTP relay, a fee to release a parcel/prize, payment redirection,
|
||
remote access, etc.) driven by manipulation tactics. Reserve this level for
|
||
messages that would cause real harm if acted on.
|
||
- suspicious: manipulation signals are present but the message could plausibly
|
||
be legitimate, OR the only indicators are weak (urgency alone, an unfamiliar
|
||
link alone, an authority claim with no money/data ask). The honest middle:
|
||
"verify through your own channel before acting".
|
||
- no_indicators: no scam mechanism; a genuine-looking OTP, transaction alert,
|
||
courier notice, receipt, promo, or ordinary message. Legitimate messages can
|
||
be urgent and can contain links — do NOT flag them for that alone.
|
||
|
||
2. tactics — a list of the manipulation tactics you detected, each an object with:
|
||
- tactic: one of the fixed tactic ids below,
|
||
- evidence: a VERBATIM substring copied character-for-character from the input
|
||
message (it will be checked by exact substring match; any tactic whose
|
||
evidence is not found verbatim in the message is dropped and counted as
|
||
fabricated, so never paraphrase, truncate mid-word, or invent evidence),
|
||
- explanation: one calm, plain sentence a non-technical person (including an
|
||
elderly person) can understand — describe the manipulation pattern, never
|
||
instructions for constructing it, and never panic language.
|
||
When verdict is no_indicators, tactics is an empty list.
|
||
|
||
Fixed tactic ids (use these exact strings, nothing else):
|
||
- urgency_pressure
|
||
- authority_impersonation
|
||
- payment_redirect
|
||
- credential_phishing
|
||
- courier_customs_fee
|
||
- prize_lottery
|
||
- investment_too_good
|
||
- romance_advance_fee
|
||
- family_emergency_impersonation
|
||
- tech_support
|
||
- link_obfuscation
|
||
- refund_overpayment
|
||
- subscription_trap
|
||
|
||
3. explanation — one or two calm, actionable sentences summarizing the verdict
|
||
for a frightened non-technical reader. No alarmist tone; state what is going on
|
||
and, implicitly, that they can verify safely.
|
||
|
||
4. recommended_action — one safe action id from the fixed schema enum. Prefer the
|
||
action that routes the reader to THEIR OWN trusted channel (their bank's
|
||
official number, the courier's own app) rather than any contact in the message.
|
||
Use no_action_needed only for no_indicators.
|
||
|
||
Rules:
|
||
- Evidence must be a verbatim substring of the message. This is non-negotiable.
|
||
- A tactic label means "this manipulation pattern is present", not "this is a
|
||
scam" — the verdict is a separate judgment. Multiple tactics per message are
|
||
expected.
|
||
- Urgency alone, a shortener/unfamiliar link alone, or an authority claim with no
|
||
money/data ask should rarely exceed suspicious.
|
||
```
|
||
|
||
## 2. User-turn format (exact)
|
||
|
||
```
|
||
[channel: sms]
|
||
Contul tău a fost blocat. Confirmă datele aici: http://example-bank.invalid/verify
|
||
```
|
||
|
||
- The prefix is literally `[channel: ` + tag + `]` then a newline then the raw
|
||
message text. Tag is one of `sms`, `email`, `chat` (lowercase).
|
||
- Apply the model's chat template around system+user as usual (mlx-swift
|
||
`apply_chat_template` with `add_generation_prompt=true`). No few-shot, no extra
|
||
preamble — the system prompt is the whole instruction.
|
||
|
||
## 3. Constrained JSON decoding (the enum-typo fix)
|
||
|
||
The reference implementation decodes with the output constrained to
|
||
`schema_scamguard_v1.json` (a JSON-schema grammar). Do the same on device — this is
|
||
non-negotiable for a shippable engine:
|
||
|
||
- Minimum bar: constrain the three enum fields to their closed sets — `verdict`,
|
||
each `tactics[].tactic`, and `recommended_action`. That alone removes the enum
|
||
typos you saw.
|
||
- Better: constrain the whole object shape (a GBNF grammar compiled from the JSON
|
||
schema; llama.cpp/`mlx`-side grammar or a logit mask). `additionalProperties` is
|
||
false — reject unknown keys.
|
||
- Free-text fields (`evidence`, `explanation`) stay unconstrained strings.
|
||
|
||
If mlx-swift lacks grammar decoding today: implement a logit mask over the enum
|
||
tokens at minimum, and keep your strict validator as the backstop (re-sample on
|
||
invalid). But plan for real grammar-constrained decode — the validity gate depends
|
||
on it.
|
||
|
||
## 4. Output contract (schema summary)
|
||
|
||
Machine-readable: `schema_scamguard_v1.json`. Shape:
|
||
|
||
```json
|
||
{
|
||
"verdict": "scam_likely | suspicious | no_indicators",
|
||
"tactics": [
|
||
{ "tactic": "<one of the 13 ids>", "evidence": "<verbatim substring>", "explanation": "<plain sentence>" }
|
||
],
|
||
"explanation": "<overall plain-language explanation>",
|
||
"recommended_action": "<one of the 10 safe-action ids>"
|
||
}
|
||
```
|
||
|
||
- `verdict` (required, enum, 3 values above)
|
||
- `tactics` (required list; **empty** when `verdict == no_indicators`). Each item is
|
||
`{tactic, evidence, explanation}` — all three required, `evidence`/`explanation`
|
||
min length 1.
|
||
- `explanation` (required, non-empty)
|
||
- `recommended_action` (required, enum, 10 values):
|
||
`call_bank_official_number`, `do_not_click_link`, `verify_via_official_app_or_site`,
|
||
`call_family_member_known_number`, `do_not_share_codes_or_credentials`,
|
||
`do_not_send_money`, `ignore_and_delete`, `report_to_authorities`,
|
||
`check_sender_address`, `no_action_needed`.
|
||
- `additionalProperties: false` at every level.
|
||
|
||
The 13 tactic ids are the exact strings in the prompt above (taxonomy v1).
|
||
|
||
## 5. Post-decode evidence kill-switch (part of the contract)
|
||
|
||
After decode, for each detected tactic, verify `evidence` is a **verbatim
|
||
substring** of the input message. If it is not found character-for-character, DROP
|
||
that tactic (do not display it) and count it as fabricated. This mirrors the
|
||
training/eval verifier (`scamguard.verify`); the model is trained expecting it.
|
||
Keep your strict validator — just add this substring check to it.
|
||
|
||
## 6. The reference Romanian message / int4 note
|
||
|
||
With the exact prompt (§1) + user format (§2) + constrained decode (§3), the model
|
||
card's reference RO message should classify correctly. If it still misfires **after
|
||
all three are matched**, it's most likely quantization sensitivity — try the **int8**
|
||
weights instead of int4 for that case and tell us; we'll co-debug (it may be a known
|
||
int4 edge we document, not an integration bug). Don't treat a single reference-case
|
||
miss as a contract failure until §1–§3 are byte-exact.
|
||
|
||
## 7. Versioning
|
||
|
||
`PROMPT_VERSION = scamguard_sys_v1`. The prompt + schema are frozen together and the
|
||
weights are trained against them. Any change = a new version id + a retrain; never
|
||
silently edit the prompt on the client. Pin the version string in the app and log it
|
||
with each verdict so a future model swap is traceable.
|
||
|
||
---
|
||
|
||
*Delivered by the scam-guard model team to unblock the M0 spike. Questions on the
|
||
prompt/schema/decode → model team; this contract is the source of truth over any
|
||
value inferred from the PRD.*
|