target_model: id: Nanthasit/sakthai-context-0.5b-tools pipeline_tag: text-generation library_name: transformers base_model: Qwen/Qwen2.5-0.5B-Instruct downloads: 94 likes: 0 private: false gated: false created: 2026-07-04T22:27:20.000Z last_modified: 2026-07-31T04:47:54.000Z model_age_days: 26.28 model_type: merged full SFT (494M params, bfloat16, ~942 MB safetensors) has_weights: true architecture: model_type: qwen2 architectures: ["Qwen2ForCausalLM"] hidden_size: 896 num_hidden_layers: 24 num_attention_heads: 14 num_key_value_heads: 2 intermediate_size: 4864 vocab_size: 151936 max_position_embeddings: 32768 total_parameters: 494032768 dtype: bfloat16 rope_theta: 1000000 rope_type: default tie_word_embeddings: true attention: full (24/24 full_attention layers; sliding_window null — no SWA) transformers_version: 5.14.1 repo_summary: siblings_count: 9 total_repo_bytes: 999542790 total_gb: 0.931 has_weights: true weight_file_count: 1 weight_files: ["model.safetensors (988,097,824 bytes, ~942 MB BF16)"] config_present: true tokenizer_present: true chat_template_present: true readme_present: true readme_size_bytes: 16229 eval_files_present: 1 eval_files: [".eval_results/benchmark-20260731_015807.yaml"] note: >- Clean 9-file merged repo — single BF16 safetensors, full tokenizer + chat_template.jinja, no junk. GGUF lives in the merged sibling (sakthai-context-0.5b-merged), which this README links for CPU users. benchmarks: model_index_count: 1 metrics_count: 8 all_verified: false pending_metrics: 8 entries: - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Selection Accuracy value: 91.8 verified: false - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Arguments Accuracy value: 45.7 verified: false - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Strict Accuracy value: 45.7 verified: false - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Held-Out Tool Accuracy value: 87.8 verified: false - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Partial Arguments Credit value: 58.1 verified: false - task: text-generation dataset: Nanthasit/sakthai-bench-v2 (500 rows) metric: Degenerate Outputs value: 0 verified: false - task: text-generation dataset: sakthai-bench-v2 (200 irrelevance rows) metric: Correct Silence (no tools offered) value: 100.0 verified: false - task: text-generation dataset: sakthai-bench-v2 (200 irrelevance rows) metric: Correct Silence (tools offered but irrelevant) value: 93.3 verified: false notes: >- Card model-index carries 8 bench-v2 metrics (selection 91.8% = 2.3x over the v1 40.2% baseline; args 45.7%; held-out 87.8% on web_search + get_news_headlines; 0 degenerate). Card claims independently reproducible via eval_bench.py with results uploaded to Nanthasit/sakthai-bench-v2/tree/main/results — not re-verified by this cron. A real single-trial inference check in this repo (.eval_results/benchmark-20260731_015807.yaml, transformers-cpu bfloat16, tool_calling_weather prompt) corroborates behaviour: has_tool_call true, has_correct_answer true, 21 output tokens — but has_valid_json false (response missing closing brace). Publish a multi-trial verified pass before promoting the 91.8% claim to "verified". training: method: SFT with prompt-masked (completion-only) loss base_model: Qwen/Qwen2.5-0.5B-Instruct dataset: Nanthasit/sakthai-combined-v7 (2,050 rows after bench exclusion) lora_config: "r=16, alpha=32, dropout=0.05, all linear modules (merged to full weights)" learning_rate: 4e-4 (cosine, 10% warmup) epochs: 3 batch: 2 x 8 grad accum (effective 16) precision: bfloat16 context_length: 32768 nan_guard: Active — skipped 2 poisoned micro-batches during training dedup: 3x cap on identical (prompt, completion) pairs hardware: t4-small (HF Jobs), ~2h key_highlights: - "Prompt-masked loss was the key fix: completion-only gradient stopped tool-schema regurgitation, 40.2% → 91.8% selection (+51.6pp, 2.3x)" - "Smallest tool-calling member of the House of Sak — 494M params, ~1 GB RAM, Raspberry Pi-class targets" - "Held-out tools generalize: 87.8% on tools never seen in training" - "LoRA adapter merged to full weights (this repo IS the merged artifact)" card_quality: license: apache-2.0 base_model_documented: true base_model: Qwen/Qwen2.5-0.5B-Instruct tags_count: 28 tags: - agent - conversational - ollama - transformers - small-language-model - slm - tool-use - qwen - qwen2.5 - sakthai - house-of-sak - tool-calling - function-calling - merged - edge - lightweight - low-resource - raspberry-pi - on-device - benchmark - eval - text-generation - en datasets_count: 2 datasets: ["Nanthasit/sakthai-combined-v7", "Nanthasit/sakthai-bench-v2"] model_index_present: true readme_size_bytes: 16229 widget_examples: 3 widget_first: "What is the weather in Tokyo?" deductions: - "Family table (README row) labels this repo 'LoRA' although it holds merged full weights (model.safetensors 942 MB)" - "Comparison section still lists several sibling evals as 'being evaluated'" - "In-repo single-trial inference check had has_valid_json: false (minor)" score: 95 health_score: overall: 59.9 components: popularity: 0.9 momentum: 20 benchmarks: 88 card_quality: 95 repo_hygiene: 98 weights: popularity: 0.20 momentum: 0.20 benchmarks: 0.25 card_quality: 0.20 repo_hygiene: 0.15 sibling_comparison: rank_by_downloads: 10 total_author_models: 19 max_sibling_downloads: 1599 models_with_positive_downloads: 11 velocity_rank: 11 max_sibling_velocity: 62.62 our_velocity: 3.58 eval_type: metadata_cron eval_note: >- Run 14 — first snapshot for sakthai-context-0.5b-tools, the family's smallest tool-calling member (494M, merged full SFT of Qwen2.5-0.5B-Instruct) and the direct sibling of the #2 model sakthai-context-0.5b-merged. 94 downloads (rank 10/19), 3.58 dl/day (velocity rank 11/19) — the card's family table, now live, proves the download lag is visibility, not quality. Strengths: genuinely excellent card (8-metric model-index, 91.8% selection = 2.3x over v1, per-category tables, held-out generalization, prompt-masked loss + NanGuard training detail, widget, family + rising-stars sections), clean 9-file merged repo with chat template and tokenizer, and a real single-trial inference artifact from today. Weaknesses: popularity component raw-count-capped (0.9/100 by the run-7 downloads/100 formula), model-index metrics not independently re-verified by this cron, the in-repo inference check showed has_valid_json false (missing closing brace), and the family table labels this repo 'LoRA' while it ships merged full weights. Recommendations: (1) multi-trial verification pass on the 91.8%/45.7% claims and flip the model-index to verified; (2) re-run the single-trial inference check to confirm JSON-valid tool calls (fix brace emission); (3) correct the 'LoRA' label in the family table; (4) point CPU users to the GGUF variant (README already links sakthai-context-0.5b-merged) and consider bundling a Q4 GGUF here for a zero-hop edge path. eval_metadata: model: Nanthasit/sakthai-context-0.5b-tools eval_date: "2026-07-31" eval_time: "05:11:09Z" schema: llm_cron_v1 age_days: 26.28 days_since_last_update: 0.016 download_velocity: 3.58 cron_run: 14