Files
ModelHub XC 5244475463 初始化项目,由ModelHub XC社区提供模型
Model: flowxai/scam-guard-qwen06b
Source: Original Platform
2026-07-29 03:16:11 +08:00

8.8 KiB
Raw Permalink Blame History

scam-guard — Inference Contract (scamguard_sys_v1)

For the iOS / ScamGuardMLX engine. This is the frozen inference contract for the published scam-guard-qwen06b (and -qwen17b) weights. It is the exact spec the model was trained and evaluated against — the thing the M0 spike correctly identified as missing. Wire the engine to this and the enum typos and the wrong-verdict-on-the-reference-message go away.

Three files in this folder:

  • INFERENCE.md — this document
  • prompt_scamguard_sys_v1.txt — the system prompt, verbatim (source of truth)
  • schema_scamguard_v1.json — the output JSON Schema, for constrained decoding

Great M0 work — feasibility (0.5–2 s) matches the card, and your schema correction is correct: the contract is tactics[] = {tactic, evidence, explanation}.


TL;DR — what fixes the blocker

The model is stochastic over an exact training-time contract. Three things must match it byte-for-byte; miss any one and you get enum typos / wrong verdicts:

  1. System prompt = the frozen scamguard_sys_v1 string below. Never edit it.
  2. User turn = [channel: <tag>]\n<message text> where <tag> ∈ sms | email | chat.
  3. Constrained JSON decoding against schema_scamguard_v1.json. This is what kills the enum typos — the enums are decoded as a closed set, not free text.

Then a post-decode evidence check (verbatim-substring) is part of the contract, not optional.


1. System prompt — scamguard_sys_v1 (verbatim, do not edit)

Load prompt_scamguard_sys_v1.txt as the system turn exactly. Bump the version id and re-coordinate with the model team if a single character changes — the model is trained against this string.

You are an on-device scam and fraud detector. You read ONE message (SMS, email,
or chat text, possibly containing visible URLs) and judge whether it is a scam.
You never fetch URLs or use any network; you reason ONLY over the visible text.

Return a single JSON object with these fields:

1. verdict — exactly one of three levels (never a probability or percentage):
   - scam_likely: clear scam mechanism present (a request for money, credentials,
     card data, OTP relay, a fee to release a parcel/prize, payment redirection,
     remote access, etc.) driven by manipulation tactics. Reserve this level for
     messages that would cause real harm if acted on.
   - suspicious: manipulation signals are present but the message could plausibly
     be legitimate, OR the only indicators are weak (urgency alone, an unfamiliar
     link alone, an authority claim with no money/data ask). The honest middle:
     "verify through your own channel before acting".
   - no_indicators: no scam mechanism; a genuine-looking OTP, transaction alert,
     courier notice, receipt, promo, or ordinary message. Legitimate messages can
     be urgent and can contain links — do NOT flag them for that alone.

2. tactics — a list of the manipulation tactics you detected, each an object with:
   - tactic: one of the fixed tactic ids below,
   - evidence: a VERBATIM substring copied character-for-character from the input
     message (it will be checked by exact substring match; any tactic whose
     evidence is not found verbatim in the message is dropped and counted as
     fabricated, so never paraphrase, truncate mid-word, or invent evidence),
   - explanation: one calm, plain sentence a non-technical person (including an
     elderly person) can understand — describe the manipulation pattern, never
     instructions for constructing it, and never panic language.
   When verdict is no_indicators, tactics is an empty list.

   Fixed tactic ids (use these exact strings, nothing else):
  - urgency_pressure
  - authority_impersonation
  - payment_redirect
  - credential_phishing
  - courier_customs_fee
  - prize_lottery
  - investment_too_good
  - romance_advance_fee
  - family_emergency_impersonation
  - tech_support
  - link_obfuscation
  - refund_overpayment
  - subscription_trap

3. explanation — one or two calm, actionable sentences summarizing the verdict
   for a frightened non-technical reader. No alarmist tone; state what is going on
   and, implicitly, that they can verify safely.

4. recommended_action — one safe action id from the fixed schema enum. Prefer the
   action that routes the reader to THEIR OWN trusted channel (their bank's
   official number, the courier's own app) rather than any contact in the message.
   Use no_action_needed only for no_indicators.

Rules:
- Evidence must be a verbatim substring of the message. This is non-negotiable.
- A tactic label means "this manipulation pattern is present", not "this is a
  scam" — the verdict is a separate judgment. Multiple tactics per message are
  expected.
- Urgency alone, a shortener/unfamiliar link alone, or an authority claim with no
  money/data ask should rarely exceed suspicious.

2. User-turn format (exact)

[channel: sms]
Contul tău a fost blocat. Confirmă datele aici: http://example-bank.invalid/verify
  • The prefix is literally [channel: + tag + ] then a newline then the raw message text. Tag is one of sms, email, chat (lowercase).
  • Apply the model's chat template around system+user as usual (mlx-swift apply_chat_template with add_generation_prompt=true). No few-shot, no extra preamble — the system prompt is the whole instruction.

3. Constrained JSON decoding (the enum-typo fix)

The reference implementation decodes with the output constrained to schema_scamguard_v1.json (a JSON-schema grammar). Do the same on device — this is non-negotiable for a shippable engine:

  • Minimum bar: constrain the three enum fields to their closed sets — verdict, each tactics[].tactic, and recommended_action. That alone removes the enum typos you saw.
  • Better: constrain the whole object shape (a GBNF grammar compiled from the JSON schema; llama.cpp/mlx-side grammar or a logit mask). additionalProperties is false — reject unknown keys.
  • Free-text fields (evidence, explanation) stay unconstrained strings.

If mlx-swift lacks grammar decoding today: implement a logit mask over the enum tokens at minimum, and keep your strict validator as the backstop (re-sample on invalid). But plan for real grammar-constrained decode — the validity gate depends on it.

4. Output contract (schema summary)

Machine-readable: schema_scamguard_v1.json. Shape:

{
  "verdict": "scam_likely | suspicious | no_indicators",
  "tactics": [
    { "tactic": "<one of the 13 ids>", "evidence": "<verbatim substring>", "explanation": "<plain sentence>" }
  ],
  "explanation": "<overall plain-language explanation>",
  "recommended_action": "<one of the 10 safe-action ids>"
}
  • verdict (required, enum, 3 values above)
  • tactics (required list; empty when verdict == no_indicators). Each item is {tactic, evidence, explanation} — all three required, evidence/explanation min length 1.
  • explanation (required, non-empty)
  • recommended_action (required, enum, 10 values): call_bank_official_number, do_not_click_link, verify_via_official_app_or_site, call_family_member_known_number, do_not_share_codes_or_credentials, do_not_send_money, ignore_and_delete, report_to_authorities, check_sender_address, no_action_needed.
  • additionalProperties: false at every level.

The 13 tactic ids are the exact strings in the prompt above (taxonomy v1).

5. Post-decode evidence kill-switch (part of the contract)

After decode, for each detected tactic, verify evidence is a verbatim substring of the input message. If it is not found character-for-character, DROP that tactic (do not display it) and count it as fabricated. This mirrors the training/eval verifier (scamguard.verify); the model is trained expecting it. Keep your strict validator — just add this substring check to it.

6. The reference Romanian message / int4 note

With the exact prompt (§1) + user format (§2) + constrained decode (§3), the model card's reference RO message should classify correctly. If it still misfires after all three are matched, it's most likely quantization sensitivity — try the int8 weights instead of int4 for that case and tell us; we'll co-debug (it may be a known int4 edge we document, not an integration bug). Don't treat a single reference-case miss as a contract failure until §1–§3 are byte-exact.

7. Versioning

PROMPT_VERSION = scamguard_sys_v1. The prompt + schema are frozen together and the weights are trained against them. Any change = a new version id + a retrain; never silently edit the prompt on the client. Pin the version string in the app and log it with each verdict so a future model swap is traceable.


Delivered by the scam-guard model team to unblock the M0 spike. Questions on the prompt/schema/decode → model team; this contract is the source of truth over any value inferred from the PRD.