commit fcec508f05c26ea71efd214850bdc164bbbe032d Author: ModelHub XC Date: Tue Sep 15 23:25:19 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: jsantillana/vectrayx-base-260m-gguf Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..b3479f6 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,37 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +vectrayx-base-260m-f16.gguf filter=lfs diff=lfs merge=lfs -text +vectrayx-base-260m-v2-f16.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..e9d4151 --- /dev/null +++ b/README.md @@ -0,0 +1,284 @@ +--- +language: +- es +- en +license: apache-2.0 +tags: +- cybersecurity +- tool-use +- function-calling +- thinking +- llama-cpp +- gguf +- latam +- spanish +base_model: [] +model_type: llama +pipeline_tag: text-generation +library_name: gguf +--- + +# VectraYX-Base-260M + +**VectraYX-Base-260M** is a 260M parameter language model specialized in cybersecurity for Latin America, trained from scratch in Spanish. It supports native tool use with `<|tool_call|>`, explicit reasoning with ``, and technical conversation in Latin American Spanish. + +> Compact architecture, efficient on CPU/GPU, deployable with Ollama and llama.cpp. + +--- + +## Key Features + +- **260M parameters** — lightweight, fast, deployable on consumer hardware +- **Native tool use** — generates `<|tool_call|>{...}<|/tool_call|>` JSON blocks +- **Chain-of-thought** — explicit reasoning with `...` tags +- **LATAM-first** — trained on Latin American Spanish corpus (laws, regulations, regional context) +- **Cybersecurity-first** — CVE Q&A, threat classification, pentesting commands, MITRE ATT&CK +- **GGUF / llama.cpp** — compatible with Ollama, LM Studio, llama.cpp + +--- + +## Architecture + +| Parameter | Value | +|---|---| +| Total parameters | 260M | +| Layers | 16 | +| Attention heads | 16 | +| KV heads (GQA) | 4 | +| d_model | 1024 | +| d_ffn | 4096 | +| Vocab size | 16,384 | +| Context length | 1,024 tokens | +| RoPE theta | 10,000 | +| QK-Norm | No | +| Tie embeddings | Yes | +| GGUF architecture | `llama` | + +--- + +## Training Pipeline + +### Phase 1 — General Pretraining +- **Corpus:** general conversational Spanish +- **Tokens seen:** ~4.24B +- **Goal:** linguistic base and Latin American Spanish comprehension + +### Phase 2 — Technical Specialization +- **Corpus:** cybersecurity documentation, CVEs, writeups, tools +- **Tokens seen:** 2.03B (epochs=1.0) +- **Steps:** 15,500 | **Final loss:** 2.07 +- **Total accumulated:** 6.27B tokens (~1.2× Chinchilla optimal for 260M) + +### Phase 3 — Domain Adaptation (tools + LATAM) +- **Corpus:** 31,969 hybrid examples with `` + LATAM + tool SFT + - Cybersec hybrids generated with GPT-4.1-mini + real Kali/Ubuntu sandbox + - LATAM: cybersecurity laws, regulations, regional corpus + - Tool SFT inherited from VectraYX-Nano +- **Epochs:** 0.1 (surgical pass, no overfitting) +- **Mix:** 70% tools, 20% tech replay, 10% conv replay + +### SFT — Instructions + tool use + thinking +- **Data:** 17,508 examples with loss masking on assistant turns +- **Thinking:** supervised `` examples (Option A) +- **Curriculum:** + - Epoch 1: 100% conversational + - Epoch 2: 70% conv + 30% CVE Q&A + - Epoch 3: 55% conv + 30% CVE + 15% tool use (with ``) +- **Steps:** 1,065 | **Final loss:** 0.064 + +--- + +## Benchmarks — VectraYX-Bench + +Evaluated with the internal VectraYX-Bench harness (B1–B5), N=1 seed. + +| Benchmark | Description | Base (post-P3) | Post-SFT | +|---|---|---|---| +| **B1** CVE Q&A | Keyword recall in CVE responses | 0.337 | 0.341 | +| **B2** Classification | Threat classification accuracy | 0.215 | 0.185 | +| **B3** Commands | Tool match in pentesting commands | 0.210 | **0.350** | +| **B4** Tool use | Correct JSON tool activation | 0.230 | 0.230 | +| **B5** Conversational | Coherence in Spanish dialogue | 0.691 | **0.755** | + +> **Note:** B2 drops slightly post-SFT — expected, as SFT prioritizes tool use and conversation over classification. B3 and B5 show the largest gains. + +--- + +## Qualitative Evaluation (22 questions, 7 categories) + +Zero-shot evaluation with llama.cpp on CPU (Azure Standard_D8s_v3). N=22 questions covering CVE knowledge, tool use, chain-of-thought, pentesting commands, MITRE ATT&CK, and conversation. + +### Results by Category + +| Category | Questions | Passed | Score | +|---|---|---|---| +| Pentesting commands | 4 | 4 | **100%** | +| MITRE ATT&CK | 2 | 1 | 50% | +| Conversation (cybersecurity ES) | 4 | 1 | 25% | +| CVE Knowledge | 4 | 0 | 0%* | +| Tool Use | 5 | 0 | 0%* | +| Thinking | 3 | 0 | 0% | +| **TOTAL** | **22** | **6** | **27%** | + +> *CVE and Tool Use require external augmentation: the model is designed to query CVE facts via `<|tool_call|>` rather than memorize them. With tool augmentation (NVD/CISA KEV backend) scores improve significantly (see B1=0.341, B4=0.230 in VectraYX-Bench). + +### Behavioral Metrics + +| Metric | Value | +|---|---| +| `` activated (when appropriate) | 41% of responses | +| `<|tool_call|>` generated (zero-shot) | 14% of responses | +| Responds in Spanish | 82% of responses | + +### Example Responses + +**Nmap scan with chain-of-thought + tool_call:** +``` +Q: Give me the exact nmap command to detect services and versions on 192.168.1.0/24, + scanning all TCP ports. + +A: To detect services and versions on 192.168.1.0/24, I should use nmap +with the -sV flag which detects open service versions. I will run nmap -sV +192.168.1.1 to get this information. +<|tool_call|>{"name": "bash_exec", "args": {"cmd": "nmap -sV -p- 192.168.1.0/24"}}<|/tool_call|> +``` + +**SUID binaries (exact answer):** +``` +Q: What is the find command to search for SUID binaries in Linux that could be + used for privilege escalation? + +A: find / -perm -4000 -type f 2>/dev/null +``` + +**Hydra SSH brute-force with reasoning:** +``` +Q: How would you use Hydra to brute-force SSH on 10.0.0.1 with rockyou.txt? + +A: The user wants to perform a brute-force attack on SSH at 10.0.0.1 +using the rockyou.txt wordlist. I should use Hydra with -l for the user +and -P for the wordlist. +hydra -l root -P /usr/share/wordlists/rockyou.txt ssh://10.0.0.1 -t 4 +``` + +**Gobuster directory fuzzing:** +``` +Q: Give me a gobuster command for directory fuzzing on http://target.com using common.txt. + +A: gobuster dir -u http://target.com -w /usr/share/wordlists/dirb/common.txt +``` + +### Observations + +- **Primary strength:** pentesting command generation (nmap, hydra, gobuster, find SUID) — 100% accuracy with direct prompts. +- **Chain-of-thought:** `` activates in 41% of responses, including complex commands and incident analysis. +- **CVE knowledge:** the model has general severity knowledge but does not memorize specific CVEs — intentional design. With tool augmentation (NVD/CISA KEV) responses are precise. +- **Zero-shot tool use:** `<|tool_call|>` activation is lower on generic prompts. The model responds best when the system prompt includes the exact tool schema from training. + +--- + +## Quick Start + +### With Ollama + +```bash +ollama run jsantillana/vectrayx-base-260m +``` + +### With llama.cpp + +```bash +llama-cli -m vectrayx-base-260m-f16.gguf \ + --prompt "<|system|>You are VectraYX, a cybersecurity expert for LATAM.<|end|><|user|>How do I scan open ports on a network?<|end|><|assistant|>" \ + -n 512 --temp 0.7 +``` + +### With Python (llama-cpp-python) + +```python +from llama_cpp import Llama + +llm = Llama(model_path="vectrayx-base-260m-f16.gguf", n_ctx=1024) +response = llm( + "<|system|>You are VectraYX, cybersecurity expert for LATAM.<|end|>" + "<|user|>Explain CVE-2021-44228<|end|><|assistant|>", + max_tokens=512, + temperature=0.7, +) +print(response["choices"][0]["text"]) +``` + +--- + +## Conversation Format + +``` +<|system|>System instructions<|end|> +<|user|>User question<|end|> +<|assistant|> +Internal reasoning here... + +<|tool_call|>{"name": "nvd_get_cve", "args": {"cve_id": "CVE-2021-44228"}}<|/tool_call|> +<|tool_result|>{"cvss_score": 9.8, "severity": "CRITICAL", ...}<|/tool_result|> +Response to user...<|end|> +``` + +### Available Tools + +| Tool | Description | +|---|---| +| `nvd_get_cve(cve_id)` | Get CVSS score, description and references for a CVE | +| `nvd_search(query, limit)` | Search recent CVEs by keyword | +| `cisa_kev_check(cve_id)` | Check if a CVE is in CISA's KEV catalog | +| `mitre_get_technique(technique_id)` | Describe a MITRE ATT&CK technique | +| `otx_check_ioc(ioc_type, value)` | Check IP/domain/hash reputation in AlienVault OTX | +| `bash_exec(cmd)` | Execute a bash command for analysis or forensics | + +--- + +## Training Data + +The model was trained on: +- General conversational Spanish corpus (Phase 1) +- Cybersecurity technical documentation and CVEs (Phase 2) +- **Synthetic hybrid dataset** generated with GPT-4.1-mini + real Kali Linux sandbox: + - `help_grounding`, `unknown_tool`, `man_section`, `multi_hop_2/3` + - `recovery`, `negative_no_tool`, `cvss_reasoning`, `ad_attacks` + - `malware_analysis`, `red_blue_dual`, `compliance_report` +- **LATAM corpus:** Latin American cybersecurity laws, national regulations, regional context +- **Tool SFT** inherited from VectraYX-Nano: `tool_sft_mini_v1`, `tool_sft_v3_bash`, `tooluse_dataset` + +--- + +## Responsible Use + +This model is designed for cybersecurity professionals, incident response teams, and educators in Latin America. Knowledge of offensive techniques is included for **educational and defensive purposes**. + +**Do not use for:** unauthorized attacks, exploitation of systems without permission, or illegal activities. + +--- + +## VectraYX Family + +| Model | Params | Specialty | +|---|---|---| +| VectraYX-Nano | ~35M | Ultra-lightweight, edge, LATAM | +| **VectraYX-Base-260M** | 260M | Cybersecurity LATAM, tool use, thinking | + +--- + +## Citation + +```bibtex +@misc{vectrayx-base-260m-2026, + title = {VectraYX-Base-260M: A Cybersecurity Language Model for Latin America}, + author = {Santillana, Juan S.}, + year = {2026}, + publisher = {Hugging Face}, + url = {https://huggingface.co/jsantillana/vectrayx-base-260m-gguf} +} +``` + +--- + +*Trained on Azure H100 NVL · Pipeline: PyTorch + llama.cpp · Exported to GGUF* diff --git a/config.json b/config.json new file mode 100644 index 0000000..184f33a --- /dev/null +++ b/config.json @@ -0,0 +1,59 @@ +{ + "model": { + "vocab_size": 16384, + "n_layers": 16, + "n_heads": 16, + "n_kv_heads": 4, + "d_model": 1024, + "d_ffn": 4096, + "max_seq_len": 4096, + "rope_theta": 500000.0, + "rms_eps": 1e-06, + "init_std": 0.02, + "dropout": 0.0, + "tie_embeddings": true, + "qk_norm": false, + "z_loss_coef": 0.0001 + }, + "tokenizer": { + "vocab_size": 16384, + "model_type": "bpe", + "character_coverage": 1.0, + "byte_fallback": true, + "normalization": "nmt_nfkc", + "split_digits": true, + "split_by_unicode_script": true, + "add_dummy_prefix": true, + "user_defined_symbols": [ + "<|pad|>", + "<|bos|>", + "<|eos|>", + "<|unk|>", + "<|sep|>", + "<|system|>", + "<|user|>", + "<|assistant|>", + "<|end|>", + "<|tool_call|>", + "<|/tool_call|>", + "<|tool_result|>", + "<|/tool_result|>", + "<|think|>", + "<|/think|>", + "<|cve|>", + "<|cvss|>", + "<|ioc|>", + "<|ttp|>", + "<|mitre|>", + "<|kev|>", + "<|exploit|>", + "<|patch|>", + "<|alert|>", + "<|critical|>", + "<|high|>", + "<|medium|>", + "<|low|>", + "<|info|>" + ] + } +} diff --git a/phase1-last.pt b/phase1-last.pt new file mode 100644 index 0000000..60aa803 --- /dev/null +++ b/phase1-last.pt @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2dc32507a8be4ce558adf1c073e45c00c3ea233d6d225d2db53beaa2f888b7ab +size 3121151330 diff --git a/phase2-last.pt b/phase2-last.pt new file mode 100644 index 0000000..1361f54 --- /dev/null +++ b/phase2-last.pt @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:04057b74575004eb4f51df5e36e80eb210cd43e9e5b352b30fc8f45c08920dba +size 3121151330 diff --git a/phase3-last.pt b/phase3-last.pt new file mode 100644 index 0000000..ada7782 --- /dev/null +++ b/phase3-last.pt @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f7e56ee490282581fdef13aa2b5b54201abd09344714072dbf0a36b4de033756 +size 3121151330 diff --git a/tokenizer.model b/tokenizer.model new file mode 100644 index 0000000..26dad62 --- /dev/null +++ b/tokenizer.model @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b301a6c9e5621df751c4b17e2d4bf9751a07dcd3e518d378524444654dc6f3bb +size 474625 diff --git a/vectrayx-base-260m-f16.gguf b/vectrayx-base-260m-f16.gguf new file mode 100644 index 0000000..ec6dffd --- /dev/null +++ b/vectrayx-base-260m-f16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8095f910d54f3348dd880ca9ba30bfa61a1c44f5e7baca12afbc75e802fcad4a +size 554141600 diff --git a/vectrayx-base-260m-v1.pt b/vectrayx-base-260m-v1.pt new file mode 100644 index 0000000..6a1e119 --- /dev/null +++ b/vectrayx-base-260m-v1.pt @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:873e11a07b5c870672d5e0ea49333dd970497514d2b36a6308a6dba9c53907f9 +size 3121148846 diff --git a/vectrayx-base-260m-v2-f16.gguf b/vectrayx-base-260m-v2-f16.gguf new file mode 100644 index 0000000..30dce60 --- /dev/null +++ b/vectrayx-base-260m-v2-f16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f1a0fab856b96afaf3290881d03c41a704e5513eb9674c8a240fa394166aed2e +size 554141600 diff --git a/vectrayx-base-260m-v2.pt b/vectrayx-base-260m-v2.pt new file mode 100644 index 0000000..6bf6923 --- /dev/null +++ b/vectrayx-base-260m-v2.pt @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2643fa4f4b0d37411279c9b5027d3e65aaebc397292caf49301fcbe0b974a957 +size 3121148846