Files
wardensight-1b/README.md
ModelHub XC 2d9d0c8c0a 初始化项目,由ModelHub XC社区提供模型
Model: emiluzelac/wardensight-1b
Source: Original Platform
2026-09-21 21:01:39 +08:00

290 lines
13 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
base_model:
- fdtn-ai/antares-1b
language:
- en
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
- security
- vulnerability-localization
- agentic
- terminal-agent
- wordpress
- php
- granite
model-index:
- name: WardenSight-1B
results:
- task:
type: vulnerability-localization
name: Vulnerability file localization
dataset:
type: wp-vloc
name: WP-VLoc (1,646 held-out WordPress plugin tasks)
split: test
metrics:
- type: file-f1
value: 0.3263
name: File-F1
---
# **WardenSight-1B**
WardenSight-1B is a 1-billion parameter terminal agent specialized for **vulnerability
file localization in WordPress plugins**. It is built on
[Antares-1B](https://huggingface.co/fdtn-ai/antares-1b) (Cisco Foundation AI), itself
built on IBM Granite 4.0 1B, and adapts that general-purpose localizer to a single
ecosystem through supervised fine-tuning on vulnerability-localization trajectories
from security fixes in WordPress plugins.
Given a plugin repository and a CWE class, WardenSight explores the code through a
read-only terminal interface — issuing `grep`, `find`, `cat` and similar commands — and
submits a ranked list of files it believes contain that weakness. It runs entirely
on-premises at 1B parameters, so client code never leaves the machine.
On 1,646 held-out WordPress plugin tasks, WardenSight-1B achieves a File-F1 of **0.3263**
against the base model's **0.2671** — a **+22.2% relative improvement**, paired
bootstrap 95% CI **[+0.0378, +0.0797]**, P(better) = 1.00. It reaches that by being
more selective: it names **1.31 files per task** where the base model names 2.55.
WardenSight is a **triage instrument**. It tells an analyst which files to read first.
It does not produce line numbers, exploits, patches, or proof that code is vulnerable.
![WP-VLoc File-F1 comparison](assets/wp-vloc-file-f1.svg)
To our knowledge, WardenSight-1B is the first publicly released open-weight model
specialized and evaluated for vulnerability file localization in WordPress plugins.
## **Model Details**
| Attribute | Value |
|---|---|
| Parameters | 1B-class (`GraniteMoeHybridForCausalLM`) |
| Base | `fdtn-ai/antares-1b` |
| Task | CWE-conditioned, repository-level vulnerable-file ranking |
| Ecosystem | WordPress plugins (PHP) |
| Interface | Sandboxed terminal agent, raw completions API |
| Context | 16,384 tokens |
| License | Apache-2.0 |
| Release | July 2026 |
## **Intended Use**
### **Intended Use Cases**
WordPress plugin security review by authorized personnel, where a
vulnerability class is already known or suspected and the task is locating it. Typical
uses are triaging a disclosure, narrowing a code review, or prioritizing files for
manual audit.
### **Downstream Use**
The model emits standard tool calls and integrates into any agentic loop providing a
sandboxed read-only terminal. Reports are available in human-readable, JSON and SARIF
formats for ingestion by code-scanning platforms.
### **Out-of-Scope Use**
- **Determining that code is safe.** The model has no reliable ability to report the
absence of a vulnerability — see Limitations. Point it at clean code and it will
name a file anyway.
- **Autonomous remediation** or any workflow acting on output without human review.
- **Compliance evidence** or attestation of code security.
- **WordPress themes.** Every training and evaluation task is a **plugin** — no theme
appears anywhere in the 1,807 tasks used for this work. Themes are structurally
similar PHP and the model will run against one, but its accuracy there is entirely
unmeasured. Treat theme results as unvalidated.
- **Non-WordPress codebases.** Specialization was measured on WordPress plugin PHP
only; on other ecosystems use the general-purpose base model instead.
- **Offensive use.** The model localizes weaknesses for defensive review; it does not
generate exploits and must not be used to attack systems without authorization.
## **How to Get Started with the Model**
WardenSight is an agent model, not a one-shot text classifier. It needs three pieces:
the weights, a raw-completions model server, and the Antares CLI agent loop with the
WardenSight prompt adapter.
### 1. Download the weights
```bash
python -m pip install -U "huggingface_hub[cli]"
hf download emiluzelac/wardensight-1b --local-dir ./wardensight-1b
```
### 2. Serve with vLLM
Use the raw completions API — **not** chat completions, whose server-side templating
rewrites the tool prompt the model was trained on:
```bash
python -m pip install "vllm>=0.19,<0.20"
vllm serve ./wardensight-1b \
--served-model-name wardensight-1b \
--host 0.0.0.0 --port 8000 \
--dtype bfloat16 --max-model-len 16384 \
--trust-remote-code --generation-config vllm
```
### 3. Install the agent loop and WardenSight adapter
Install the **Antares CLI** distributed with
[`fdtn-ai/antares-1b`](https://huggingface.co/fdtn-ai/antares-1b/tree/main), then
apply the release's prompt adapter and profile:
```bash
hf download fdtn-ai/antares-1b assets/antares-cli.zip \
--local-dir ./antares-cli-source
unzip ./antares-cli-source/assets/antares-cli.zip -d ./antares-cli-source
python -m venv ~/antares-venv
~/antares-venv/bin/pip install ./antares-cli-source
git clone https://github.com/emiluzelac/wardensight-1b
cd wardensight-1b
bash tooling/install.sh ~/antares-venv
export ANTARES_ENDPOINT=http://127.0.0.1:8000/v1
~/antares-venv/bin/antares query ./some-wordpress-plugin \
--cwe CWE-79 --profile wardensight
```
For a version-pinned WordPress.org plugin, install `tooling/ws-scan` on your `PATH`
and run `wardensight plugin-slug@1.2.3 CWE-79`. It downloads and extracts the requested
plugin release before invoking the same agent loop. See
[`tooling/README.md`](https://github.com/emiluzelac/wardensight-1b/blob/main/tooling/README.md)
for installation details and the exact division between Cisco's CLI and the
WardenSight wrapper.
The adapter matters. WardenSight is fine-tuned on a three-tool contract
(`terminal`, `submit_vulnerable_files`, `submit_no_vulnerability_found`) with a
specific system message; the CLI's default adapter presents a fourth tool the model
has never seen, and under it the model degrades into repeated commands. The adapter,
profile and a convenience wrapper are in
[`wardensight-1b/tooling`](https://github.com/emiluzelac/wardensight-1b/tree/main/tooling).
Sampling settings used for every published number: temperature 0.3, frequency
penalty 0.3, 15 terminal calls, 2,048 max tokens, stop tokens
`["<|end_of_text|>", "<|start_of_role|>"]`.
## **Training and Evaluation**
### **Training Data**
584 agent trajectories over WordPress plugin repositories with known, fixed
vulnerabilities, drawn from [Wordfence Intelligence](https://www.wordfence.com/threat-intel/)
advisories (CC BY 4.0). Each trajectory explores a plugin snapshot taken at the last
vulnerable release and terminates in a file submission.
Ground truth is derived from **security-relevant diff hunks**, not whole release
diffs. This distinction is load-bearing: 58% of files in a typical WordPress
vulnerable→fix release diff contain no security change, and under whole-diff labels a
constant guess of `<slug>.php` scores 0.3428 and outranks every trained model. Labels
were rebuilt around CWE-aware hunks, which drops that control to 0.1041 and makes the
benchmark able to rank localizers at all.
Training plugins are disjoint at the plugin level from every evaluation set.
### **Training Setup**
Full fine-tune of `fdtn-ai/antares-1b`, no adapters or quantization.
| Training setting | Value |
|---|---|
| Method | Supervised fine-tuning on 584 localization trajectories |
| Training length | 2 epochs |
| Learning rate | 8e-6 |
| Maximum sequence length | 5,120 tokens |
| Numeric precision | bfloat16 |
| Training hardware | 1× ASUS Ascent GX10 (NVIDIA GB10, 128 GB unified memory) |
### **Evaluation**
Evaluated on **WP-VLoc**, a WordPress-specific benchmark built for this work: plugin
snapshots reconstructed at the pre-fix release, paired with security-hunk ground truth
files. Held-out plugins never appear in training. Scored by exact path-set File-F1
with a paired bootstrap (20,000 resamples) against the base model on identical tasks.
WardenSight-1B has not been evaluated on Cisco's separate 500-task, six-ecosystem
VLoc Bench, so its WP-VLoc score should not be inserted into that benchmark's
cross-model table.
| model | File-F1 | precision | recall | files named | n |
|---|---|---|---|---|---|
| **WardenSight-1B** | **0.3263** | **0.397** | 0.315 | **1.31** | 1,646 |
| Antares-1B (base) | 0.2671 | 0.289 | 0.330 | 2.55 | 1,646 |
| constant guess `<slug>.php` | 0.1041 | 0.132 | 0.094 | 1.00 | 1,646 |
| constant guess slug + `index.php` | 0.0816 | 0.074 | 0.104 | 2.00 | 1,646 |
Paired improvement **+0.0592 File-F1 (+22.2%)**, 95% CI [+0.0378, +0.0797],
P(better) = 1.00. Task level: 453 improved, 340 regressed, 853 unchanged; 5% trimmed
mean +0.0621, so the gain is not driven by outliers.
## **Safety Alignment**
This model operates as a component in a sandboxed agentic loop, not as a
conversational assistant, and no standalone safety alignment was performed on it.
Safety should be addressed at the system level:
- Run the agent inside an isolated sandbox (Docker with `network=none`)
- Restrict deployment to authorized security personnel
- Require human review of output before any remediation action
- Audit agent trajectories in high-risk workflows
Do not expose this model as a general-purpose assistant endpoint.
## **Limitations**
1. **It cannot report a clean repository.** On 500 already-patched snapshots it
declared a vulnerability present in every one (TNR 0.000), and its abstention rate
across all held-out evaluation is 0.000 — it has never declined to name a file.
This is a property of the model family rather than this checkpoint alone; the base
model scored TNR 0.006 on the same test. Use it only when a vulnerability is known
or strongly suspected to exist.
2. **Coverage is partial.** It places a correct file in its ranked output on ~41% of
held-out tasks. It is a filter that reorders analyst attention, not a detector that
finds everything — and it solves 76 tasks the base model misses while missing 90 the
base model solves.
3. **Terminal budget.** Performance degrades on large plugins where 15 commands cannot
cover the codebase, and on vulnerabilities spanning many files.
4. **Pattern dependence.** Strongest on weaknesses with grep-able signatures (XSS,
missing authorization); weaker where the flaw requires semantic reasoning about
behavior (unsafe deserialization, injection through indirect data flow).
5. **Ecosystem scope.** WordPress **plugins** only — 1,807 plugin tasks, no themes and
no non-WordPress code. Specialization was not measured beyond that and should not be
assumed to transfer.
6. **Knowledge cutoff.** Inherits the base model's April 2025 cutoff; training data
covers WordPress advisories through July 2026.
### **Recommendations**
Give it a plugin you already suspect and the correct CWE class — supplying the wrong
class produces unhelpful output rather than an admission that nothing was found. Pin
the plugin version when investigating a known advisory; scanning the current release
often re-finds files whose vulnerability shipped a fix. Read the ranked output as a
reading order, and confirm every finding in the source before acting on it.
## **Provenance**
The 584 training trajectories for this checkpoint were generated with **Claude Code**.
WardenSight-1B is a full fine-tune of `fdtn-ai/antares-1b`; Claude was the teacher
used to produce the WordPress localization demonstrations, not the student or base
model.
WardenSight-1B is not affiliated with or endorsed by Anthropic. The raw Claude
trajectories are not distributed with this release. Users are responsible for
evaluating any provider terms that may apply to downstream use of model-generated
training data.
## **Attribution and License**
| Component | Source and license |
|---|---|
| Base model | [`fdtn-ai/antares-1b`](https://huggingface.co/fdtn-ai/antares-1b) — Cisco Foundation AI, Apache-2.0 |
| Base architecture | IBM Granite 4.0 1B — `GraniteMoeHybridForCausalLM` |
| Evaluation harness | Cisco `vulnerability-localization-benchmark`, Apache-2.0, with local corrections to the command validator, path normalization and the true-negative scorer |
| Advisory data | [Wordfence Intelligence](https://www.wordfence.com/threat-intel/), CC BY 4.0 |
| Plugin code | each plugin's own license (GPL) |
Code and weights are released under Apache-2.0. See the Provenance section for the
training-data disclosure.