290 lines
13 KiB
Markdown
290 lines
13 KiB
Markdown
---
|
||
base_model:
|
||
- fdtn-ai/antares-1b
|
||
language:
|
||
- en
|
||
library_name: transformers
|
||
license: apache-2.0
|
||
pipeline_tag: text-generation
|
||
tags:
|
||
- security
|
||
- vulnerability-localization
|
||
- agentic
|
||
- terminal-agent
|
||
- wordpress
|
||
- php
|
||
- granite
|
||
model-index:
|
||
- name: WardenSight-1B
|
||
results:
|
||
- task:
|
||
type: vulnerability-localization
|
||
name: Vulnerability file localization
|
||
dataset:
|
||
type: wp-vloc
|
||
name: WP-VLoc (1,646 held-out WordPress plugin tasks)
|
||
split: test
|
||
metrics:
|
||
- type: file-f1
|
||
value: 0.3263
|
||
name: File-F1
|
||
---
|
||
|
||
# **WardenSight-1B**
|
||
|
||
WardenSight-1B is a 1-billion parameter terminal agent specialized for **vulnerability
|
||
file localization in WordPress plugins**. It is built on
|
||
[Antares-1B](https://huggingface.co/fdtn-ai/antares-1b) (Cisco Foundation AI), itself
|
||
built on IBM Granite 4.0 1B, and adapts that general-purpose localizer to a single
|
||
ecosystem through supervised fine-tuning on vulnerability-localization trajectories
|
||
from security fixes in WordPress plugins.
|
||
|
||
Given a plugin repository and a CWE class, WardenSight explores the code through a
|
||
read-only terminal interface — issuing `grep`, `find`, `cat` and similar commands — and
|
||
submits a ranked list of files it believes contain that weakness. It runs entirely
|
||
on-premises at 1B parameters, so client code never leaves the machine.
|
||
|
||
On 1,646 held-out WordPress plugin tasks, WardenSight-1B achieves a File-F1 of **0.3263**
|
||
against the base model's **0.2671** — a **+22.2% relative improvement**, paired
|
||
bootstrap 95% CI **[+0.0378, +0.0797]**, P(better) = 1.00. It reaches that by being
|
||
more selective: it names **1.31 files per task** where the base model names 2.55.
|
||
|
||
WardenSight is a **triage instrument**. It tells an analyst which files to read first.
|
||
It does not produce line numbers, exploits, patches, or proof that code is vulnerable.
|
||
|
||

|
||
|
||
To our knowledge, WardenSight-1B is the first publicly released open-weight model
|
||
specialized and evaluated for vulnerability file localization in WordPress plugins.
|
||
|
||
## **Model Details**
|
||
|
||
| Attribute | Value |
|
||
|---|---|
|
||
| Parameters | 1B-class (`GraniteMoeHybridForCausalLM`) |
|
||
| Base | `fdtn-ai/antares-1b` |
|
||
| Task | CWE-conditioned, repository-level vulnerable-file ranking |
|
||
| Ecosystem | WordPress plugins (PHP) |
|
||
| Interface | Sandboxed terminal agent, raw completions API |
|
||
| Context | 16,384 tokens |
|
||
| License | Apache-2.0 |
|
||
| Release | July 2026 |
|
||
|
||
## **Intended Use**
|
||
|
||
### **Intended Use Cases**
|
||
|
||
WordPress plugin security review by authorized personnel, where a
|
||
vulnerability class is already known or suspected and the task is locating it. Typical
|
||
uses are triaging a disclosure, narrowing a code review, or prioritizing files for
|
||
manual audit.
|
||
|
||
### **Downstream Use**
|
||
|
||
The model emits standard tool calls and integrates into any agentic loop providing a
|
||
sandboxed read-only terminal. Reports are available in human-readable, JSON and SARIF
|
||
formats for ingestion by code-scanning platforms.
|
||
|
||
### **Out-of-Scope Use**
|
||
|
||
- **Determining that code is safe.** The model has no reliable ability to report the
|
||
absence of a vulnerability — see Limitations. Point it at clean code and it will
|
||
name a file anyway.
|
||
- **Autonomous remediation** or any workflow acting on output without human review.
|
||
- **Compliance evidence** or attestation of code security.
|
||
- **WordPress themes.** Every training and evaluation task is a **plugin** — no theme
|
||
appears anywhere in the 1,807 tasks used for this work. Themes are structurally
|
||
similar PHP and the model will run against one, but its accuracy there is entirely
|
||
unmeasured. Treat theme results as unvalidated.
|
||
- **Non-WordPress codebases.** Specialization was measured on WordPress plugin PHP
|
||
only; on other ecosystems use the general-purpose base model instead.
|
||
- **Offensive use.** The model localizes weaknesses for defensive review; it does not
|
||
generate exploits and must not be used to attack systems without authorization.
|
||
|
||
## **How to Get Started with the Model**
|
||
|
||
WardenSight is an agent model, not a one-shot text classifier. It needs three pieces:
|
||
the weights, a raw-completions model server, and the Antares CLI agent loop with the
|
||
WardenSight prompt adapter.
|
||
|
||
### 1. Download the weights
|
||
|
||
```bash
|
||
python -m pip install -U "huggingface_hub[cli]"
|
||
hf download emiluzelac/wardensight-1b --local-dir ./wardensight-1b
|
||
```
|
||
|
||
### 2. Serve with vLLM
|
||
|
||
Use the raw completions API — **not** chat completions, whose server-side templating
|
||
rewrites the tool prompt the model was trained on:
|
||
|
||
```bash
|
||
python -m pip install "vllm>=0.19,<0.20"
|
||
vllm serve ./wardensight-1b \
|
||
--served-model-name wardensight-1b \
|
||
--host 0.0.0.0 --port 8000 \
|
||
--dtype bfloat16 --max-model-len 16384 \
|
||
--trust-remote-code --generation-config vllm
|
||
```
|
||
|
||
### 3. Install the agent loop and WardenSight adapter
|
||
|
||
Install the **Antares CLI** distributed with
|
||
[`fdtn-ai/antares-1b`](https://huggingface.co/fdtn-ai/antares-1b/tree/main), then
|
||
apply the release's prompt adapter and profile:
|
||
|
||
```bash
|
||
hf download fdtn-ai/antares-1b assets/antares-cli.zip \
|
||
--local-dir ./antares-cli-source
|
||
unzip ./antares-cli-source/assets/antares-cli.zip -d ./antares-cli-source
|
||
python -m venv ~/antares-venv
|
||
~/antares-venv/bin/pip install ./antares-cli-source
|
||
|
||
git clone https://github.com/emiluzelac/wardensight-1b
|
||
cd wardensight-1b
|
||
bash tooling/install.sh ~/antares-venv
|
||
export ANTARES_ENDPOINT=http://127.0.0.1:8000/v1
|
||
~/antares-venv/bin/antares query ./some-wordpress-plugin \
|
||
--cwe CWE-79 --profile wardensight
|
||
```
|
||
|
||
For a version-pinned WordPress.org plugin, install `tooling/ws-scan` on your `PATH`
|
||
and run `wardensight plugin-slug@1.2.3 CWE-79`. It downloads and extracts the requested
|
||
plugin release before invoking the same agent loop. See
|
||
[`tooling/README.md`](https://github.com/emiluzelac/wardensight-1b/blob/main/tooling/README.md)
|
||
for installation details and the exact division between Cisco's CLI and the
|
||
WardenSight wrapper.
|
||
|
||
The adapter matters. WardenSight is fine-tuned on a three-tool contract
|
||
(`terminal`, `submit_vulnerable_files`, `submit_no_vulnerability_found`) with a
|
||
specific system message; the CLI's default adapter presents a fourth tool the model
|
||
has never seen, and under it the model degrades into repeated commands. The adapter,
|
||
profile and a convenience wrapper are in
|
||
[`wardensight-1b/tooling`](https://github.com/emiluzelac/wardensight-1b/tree/main/tooling).
|
||
|
||
Sampling settings used for every published number: temperature 0.3, frequency
|
||
penalty 0.3, 15 terminal calls, 2,048 max tokens, stop tokens
|
||
`["<|end_of_text|>", "<|start_of_role|>"]`.
|
||
|
||
## **Training and Evaluation**
|
||
|
||
### **Training Data**
|
||
|
||
584 agent trajectories over WordPress plugin repositories with known, fixed
|
||
vulnerabilities, drawn from [Wordfence Intelligence](https://www.wordfence.com/threat-intel/)
|
||
advisories (CC BY 4.0). Each trajectory explores a plugin snapshot taken at the last
|
||
vulnerable release and terminates in a file submission.
|
||
|
||
Ground truth is derived from **security-relevant diff hunks**, not whole release
|
||
diffs. This distinction is load-bearing: 58% of files in a typical WordPress
|
||
vulnerable→fix release diff contain no security change, and under whole-diff labels a
|
||
constant guess of `<slug>.php` scores 0.3428 and outranks every trained model. Labels
|
||
were rebuilt around CWE-aware hunks, which drops that control to 0.1041 and makes the
|
||
benchmark able to rank localizers at all.
|
||
|
||
Training plugins are disjoint at the plugin level from every evaluation set.
|
||
|
||
### **Training Setup**
|
||
|
||
Full fine-tune of `fdtn-ai/antares-1b`, no adapters or quantization.
|
||
|
||
| Training setting | Value |
|
||
|---|---|
|
||
| Method | Supervised fine-tuning on 584 localization trajectories |
|
||
| Training length | 2 epochs |
|
||
| Learning rate | 8e-6 |
|
||
| Maximum sequence length | 5,120 tokens |
|
||
| Numeric precision | bfloat16 |
|
||
| Training hardware | 1× ASUS Ascent GX10 (NVIDIA GB10, 128 GB unified memory) |
|
||
|
||
### **Evaluation**
|
||
|
||
Evaluated on **WP-VLoc**, a WordPress-specific benchmark built for this work: plugin
|
||
snapshots reconstructed at the pre-fix release, paired with security-hunk ground truth
|
||
files. Held-out plugins never appear in training. Scored by exact path-set File-F1
|
||
with a paired bootstrap (20,000 resamples) against the base model on identical tasks.
|
||
WardenSight-1B has not been evaluated on Cisco's separate 500-task, six-ecosystem
|
||
VLoc Bench, so its WP-VLoc score should not be inserted into that benchmark's
|
||
cross-model table.
|
||
|
||
| model | File-F1 | precision | recall | files named | n |
|
||
|---|---|---|---|---|---|
|
||
| **WardenSight-1B** | **0.3263** | **0.397** | 0.315 | **1.31** | 1,646 |
|
||
| Antares-1B (base) | 0.2671 | 0.289 | 0.330 | 2.55 | 1,646 |
|
||
| constant guess `<slug>.php` | 0.1041 | 0.132 | 0.094 | 1.00 | 1,646 |
|
||
| constant guess slug + `index.php` | 0.0816 | 0.074 | 0.104 | 2.00 | 1,646 |
|
||
|
||
Paired improvement **+0.0592 File-F1 (+22.2%)**, 95% CI [+0.0378, +0.0797],
|
||
P(better) = 1.00. Task level: 453 improved, 340 regressed, 853 unchanged; 5% trimmed
|
||
mean +0.0621, so the gain is not driven by outliers.
|
||
|
||
## **Safety Alignment**
|
||
|
||
This model operates as a component in a sandboxed agentic loop, not as a
|
||
conversational assistant, and no standalone safety alignment was performed on it.
|
||
Safety should be addressed at the system level:
|
||
|
||
- Run the agent inside an isolated sandbox (Docker with `network=none`)
|
||
- Restrict deployment to authorized security personnel
|
||
- Require human review of output before any remediation action
|
||
- Audit agent trajectories in high-risk workflows
|
||
|
||
Do not expose this model as a general-purpose assistant endpoint.
|
||
|
||
## **Limitations**
|
||
|
||
1. **It cannot report a clean repository.** On 500 already-patched snapshots it
|
||
declared a vulnerability present in every one (TNR 0.000), and its abstention rate
|
||
across all held-out evaluation is 0.000 — it has never declined to name a file.
|
||
This is a property of the model family rather than this checkpoint alone; the base
|
||
model scored TNR 0.006 on the same test. Use it only when a vulnerability is known
|
||
or strongly suspected to exist.
|
||
2. **Coverage is partial.** It places a correct file in its ranked output on ~41% of
|
||
held-out tasks. It is a filter that reorders analyst attention, not a detector that
|
||
finds everything — and it solves 76 tasks the base model misses while missing 90 the
|
||
base model solves.
|
||
3. **Terminal budget.** Performance degrades on large plugins where 15 commands cannot
|
||
cover the codebase, and on vulnerabilities spanning many files.
|
||
4. **Pattern dependence.** Strongest on weaknesses with grep-able signatures (XSS,
|
||
missing authorization); weaker where the flaw requires semantic reasoning about
|
||
behavior (unsafe deserialization, injection through indirect data flow).
|
||
5. **Ecosystem scope.** WordPress **plugins** only — 1,807 plugin tasks, no themes and
|
||
no non-WordPress code. Specialization was not measured beyond that and should not be
|
||
assumed to transfer.
|
||
6. **Knowledge cutoff.** Inherits the base model's April 2025 cutoff; training data
|
||
covers WordPress advisories through July 2026.
|
||
|
||
### **Recommendations**
|
||
|
||
Give it a plugin you already suspect and the correct CWE class — supplying the wrong
|
||
class produces unhelpful output rather than an admission that nothing was found. Pin
|
||
the plugin version when investigating a known advisory; scanning the current release
|
||
often re-finds files whose vulnerability shipped a fix. Read the ranked output as a
|
||
reading order, and confirm every finding in the source before acting on it.
|
||
|
||
## **Provenance**
|
||
|
||
The 584 training trajectories for this checkpoint were generated with **Claude Code**.
|
||
WardenSight-1B is a full fine-tune of `fdtn-ai/antares-1b`; Claude was the teacher
|
||
used to produce the WordPress localization demonstrations, not the student or base
|
||
model.
|
||
|
||
WardenSight-1B is not affiliated with or endorsed by Anthropic. The raw Claude
|
||
trajectories are not distributed with this release. Users are responsible for
|
||
evaluating any provider terms that may apply to downstream use of model-generated
|
||
training data.
|
||
|
||
## **Attribution and License**
|
||
|
||
| Component | Source and license |
|
||
|---|---|
|
||
| Base model | [`fdtn-ai/antares-1b`](https://huggingface.co/fdtn-ai/antares-1b) — Cisco Foundation AI, Apache-2.0 |
|
||
| Base architecture | IBM Granite 4.0 1B — `GraniteMoeHybridForCausalLM` |
|
||
| Evaluation harness | Cisco `vulnerability-localization-benchmark`, Apache-2.0, with local corrections to the command validator, path normalization and the true-negative scorer |
|
||
| Advisory data | [Wordfence Intelligence](https://www.wordfence.com/threat-intel/), CC BY 4.0 |
|
||
| Plugin code | each plugin's own license (GPL) |
|
||
|
||
Code and weights are released under Apache-2.0. See the Provenance section for the
|
||
training-data disclosure.
|