Files
sakthai-context-0.5b-tools/.eval_results/cron-eval-sakthai-context-0.5b-tools-2026-07-31-1.yaml
ModelHub XC f43581c75c 初始化项目,由ModelHub XC社区提供模型
Model: Nanthasit/sakthai-context-0.5b-tools
Source: Original Platform
2026-08-22 10:15:18 +08:00

228 lines
7.8 KiB
YAML

target_model:
id: Nanthasit/sakthai-context-0.5b-tools
pipeline_tag: text-generation
library_name: transformers
base_model: Qwen/Qwen2.5-0.5B-Instruct
downloads: 94
likes: 0
private: false
gated: false
created: 2026-07-04T22:27:20.000Z
last_modified: 2026-07-31T04:47:54.000Z
model_age_days: 26.28
model_type: merged full SFT (494M params, bfloat16, ~942 MB safetensors)
has_weights: true
architecture:
model_type: qwen2
architectures: ["Qwen2ForCausalLM"]
hidden_size: 896
num_hidden_layers: 24
num_attention_heads: 14
num_key_value_heads: 2
intermediate_size: 4864
vocab_size: 151936
max_position_embeddings: 32768
total_parameters: 494032768
dtype: bfloat16
rope_theta: 1000000
rope_type: default
tie_word_embeddings: true
attention: full (24/24 full_attention layers; sliding_window null — no SWA)
transformers_version: 5.14.1
repo_summary:
siblings_count: 9
total_repo_bytes: 999542790
total_gb: 0.931
has_weights: true
weight_file_count: 1
weight_files: ["model.safetensors (988,097,824 bytes, ~942 MB BF16)"]
config_present: true
tokenizer_present: true
chat_template_present: true
readme_present: true
readme_size_bytes: 16229
eval_files_present: 1
eval_files: [".eval_results/benchmark-20260731_015807.yaml"]
note: >-
Clean 9-file merged repo — single BF16 safetensors, full tokenizer +
chat_template.jinja, no junk. GGUF lives in the merged sibling
(sakthai-context-0.5b-merged), which this README links for CPU users.
benchmarks:
model_index_count: 1
metrics_count: 8
all_verified: false
pending_metrics: 8
entries:
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Selection Accuracy
value: 91.8
verified: false
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Arguments Accuracy
value: 45.7
verified: false
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Strict Accuracy
value: 45.7
verified: false
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Held-Out Tool Accuracy
value: 87.8
verified: false
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Partial Arguments Credit
value: 58.1
verified: false
- task: text-generation
dataset: Nanthasit/sakthai-bench-v2 (500 rows)
metric: Degenerate Outputs
value: 0
verified: false
- task: text-generation
dataset: sakthai-bench-v2 (200 irrelevance rows)
metric: Correct Silence (no tools offered)
value: 100.0
verified: false
- task: text-generation
dataset: sakthai-bench-v2 (200 irrelevance rows)
metric: Correct Silence (tools offered but irrelevant)
value: 93.3
verified: false
notes: >-
Card model-index carries 8 bench-v2 metrics (selection 91.8% = 2.3x over
the v1 40.2% baseline; args 45.7%; held-out 87.8% on web_search +
get_news_headlines; 0 degenerate). Card claims independently reproducible
via eval_bench.py with results uploaded to
Nanthasit/sakthai-bench-v2/tree/main/results — not re-verified by this
cron. A real single-trial inference check in this repo
(.eval_results/benchmark-20260731_015807.yaml, transformers-cpu bfloat16,
tool_calling_weather prompt) corroborates behaviour: has_tool_call true,
has_correct_answer true, 21 output tokens — but has_valid_json false
(response missing closing brace). Publish a multi-trial verified pass
before promoting the 91.8% claim to "verified".
training:
method: SFT with prompt-masked (completion-only) loss
base_model: Qwen/Qwen2.5-0.5B-Instruct
dataset: Nanthasit/sakthai-combined-v7 (2,050 rows after bench exclusion)
lora_config: "r=16, alpha=32, dropout=0.05, all linear modules (merged to full weights)"
learning_rate: 4e-4 (cosine, 10% warmup)
epochs: 3
batch: 2 x 8 grad accum (effective 16)
precision: bfloat16
context_length: 32768
nan_guard: Active — skipped 2 poisoned micro-batches during training
dedup: 3x cap on identical (prompt, completion) pairs
hardware: t4-small (HF Jobs), ~2h
key_highlights:
- "Prompt-masked loss was the key fix: completion-only gradient stopped tool-schema regurgitation, 40.2% → 91.8% selection (+51.6pp, 2.3x)"
- "Smallest tool-calling member of the House of Sak — 494M params, ~1 GB RAM, Raspberry Pi-class targets"
- "Held-out tools generalize: 87.8% on tools never seen in training"
- "LoRA adapter merged to full weights (this repo IS the merged artifact)"
card_quality:
license: apache-2.0
base_model_documented: true
base_model: Qwen/Qwen2.5-0.5B-Instruct
tags_count: 28
tags:
- agent
- conversational
- ollama
- transformers
- small-language-model
- slm
- tool-use
- qwen
- qwen2.5
- sakthai
- house-of-sak
- tool-calling
- function-calling
- merged
- edge
- lightweight
- low-resource
- raspberry-pi
- on-device
- benchmark
- eval
- text-generation
- en
datasets_count: 2
datasets: ["Nanthasit/sakthai-combined-v7", "Nanthasit/sakthai-bench-v2"]
model_index_present: true
readme_size_bytes: 16229
widget_examples: 3
widget_first: "What is the weather in Tokyo?"
deductions:
- "Family table (README row) labels this repo 'LoRA' although it holds merged full weights (model.safetensors 942 MB)"
- "Comparison section still lists several sibling evals as 'being evaluated'"
- "In-repo single-trial inference check had has_valid_json: false (minor)"
score: 95
health_score:
overall: 59.9
components:
popularity: 0.9
momentum: 20
benchmarks: 88
card_quality: 95
repo_hygiene: 98
weights:
popularity: 0.20
momentum: 0.20
benchmarks: 0.25
card_quality: 0.20
repo_hygiene: 0.15
sibling_comparison:
rank_by_downloads: 10
total_author_models: 19
max_sibling_downloads: 1599
models_with_positive_downloads: 11
velocity_rank: 11
max_sibling_velocity: 62.62
our_velocity: 3.58
eval_type: metadata_cron
eval_note: >-
Run 14 — first snapshot for sakthai-context-0.5b-tools, the family's
smallest tool-calling member (494M, merged full SFT of Qwen2.5-0.5B-Instruct)
and the direct sibling of the #2 model sakthai-context-0.5b-merged. 94
downloads (rank 10/19), 3.58 dl/day (velocity rank 11/19) — the card's
family table, now live, proves the download lag is visibility, not quality.
Strengths: genuinely excellent card (8-metric model-index, 91.8% selection
= 2.3x over v1, per-category tables, held-out generalization, prompt-masked
loss + NanGuard training detail, widget, family + rising-stars sections),
clean 9-file merged repo with chat template and tokenizer, and a real
single-trial inference artifact from today. Weaknesses: popularity component
raw-count-capped (0.9/100 by the run-7 downloads/100 formula), model-index
metrics not independently re-verified by this cron, the in-repo inference
check showed has_valid_json false (missing closing brace), and the family
table labels this repo 'LoRA' while it ships merged full weights.
Recommendations: (1) multi-trial verification pass on the 91.8%/45.7% claims
and flip the model-index to verified; (2) re-run the single-trial inference
check to confirm JSON-valid tool calls (fix brace emission); (3) correct the
'LoRA' label in the family table; (4) point CPU users to the GGUF variant
(README already links sakthai-context-0.5b-merged) and consider bundling a
Q4 GGUF here for a zero-hop edge path.
eval_metadata:
model: Nanthasit/sakthai-context-0.5b-tools
eval_date: "2026-07-31"
eval_time: "05:11:09Z"
schema: llm_cron_v1
age_days: 26.28
days_since_last_update: 0.016
download_velocity: 3.58
cron_run: 14