Files
sakthai-plus-1.5b/.eval_results/cron-eval-sakthai-plus-1.5b-2026-07-31-1.yaml
ModelHub XC 1fe2da0c55 初始化项目,由ModelHub XC社区提供模型
Model: Nanthasit/sakthai-plus-1.5b
Source: Original Platform
2026-08-21 08:25:19 +08:00

192 lines
6.9 KiB
YAML

target_model:
id: Nanthasit/sakthai-plus-1.5b
pipeline_tag: text-generation
library_name: transformers
base_model: Qwen/Qwen2.5-1.5B-Instruct
downloads: 0
likes: 0
private: false
gated: false
last_modified: "2026-07-31T04:40:16.000Z"
model_age_days: 0.6961
model_type: llm
has_weights: true
architecture:
base_model_type: qwen2
base_architectures: ["Qwen2ForCausalLM"]
base_hidden_size: 1536
base_num_hidden_layers: 28
base_num_attention_heads: 12
base_num_key_value_heads: 2
base_intermediate_size: 8960
base_vocab_size: 151936
base_max_position_embeddings: 32768
base_total_parameters: 1540000000
base_dtype: bfloat16
tie_word_embeddings: true
repo_summary:
siblings_count: 18
total_repo_bytes: 3098931104
total_gb: 3.099
has_weights: true
weight_file_count: 1
weight_bytes: 3087467144
weight_files: ["model.safetensors"]
config_present: true
tokenizer_present: true
chat_template_present: true
readme_present: true
readme_size_bytes: 8079
weight_note: "Single merged full-weight checkpoint (2.88 GiB, bf16) — no PEFT dependency at inference. Standard Qwen2.5 layout: config.json, generation_config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, README.md."
benchmarks:
model_index_count: 0
metrics_count: 9
all_verified: false
pending_metrics: 6
entries:
- dataset: lighteval
metric: winogrande
value: 59.6
verified: false
note: "entry in .eval_results/lighteval.yaml — not re-run, unverified"
- dataset: lighteval
metric: gsm8k
value: 50.9
verified: false
note: "entry in .eval_results/lighteval.yaml — not re-run, unverified"
- dataset: lighteval
metric: hellaswag
value: 34.0
verified: false
note: "entry in .eval_results/lighteval.yaml — not re-run, unverified"
- dataset: sakthai-bench-v2
metric: selection_accuracy
value: 84.8
verified: false
note: "entry in .eval_results/sakthai-bench-v2.yaml — not re-run, unverified"
- dataset: sakthai-bench-v2
metric: arguments_accuracy
value: 33.7
verified: false
note: "entry in .eval_results/sakthai-bench-v2.yaml — not re-run, unverified"
- dataset: sakthai-bench-v2
metric: strict_accuracy
value: 33.7
verified: false
note: "entry in .eval_results/sakthai-bench-v2.yaml — not re-run, unverified"
- dataset: llama.cpp tool-calling (3-trial, q4_k_m)
metric: tool_call_success
value: 1.0
verified: true
note: ".eval_results/benchmark-20260731_043634.yaml — send_email prompt, 3/3 tool calls"
- dataset: llama.cpp tool-calling (3-trial, q4_k_m)
metric: valid_json
value: 1.0
verified: true
note: ".eval_results/benchmark-20260731_043634.yaml — 3/3 valid JSON tool args"
- dataset: llama.cpp tool-calling (3-trial, q4_k_m)
metric: correct_answer
value: 1.0
verified: true
note: ".eval_results/benchmark-20260731_043634.yaml — 3/3 correct answers, avg 20.7 tps (CPU, 2 threads)"
notes: >
README frontmatter has NO model-index (model_index_count: 0) despite
.eval_results/ carrying three benchmark sources: lighteval + sakthai-bench-v2
(all unverified) and a fresh 2026-07-31 3-trial llama.cpp q4_k_m run that
passed tool-calling 3/3 (tool_call, valid JSON, correct answer). Card text
still says 'Benchmarks are pending' — stale relative to repo state.
training:
dataset: Nanthasit/sakthai-combined-v10
dataset_size: 2962
dataset_note: "v7 + v8 combined (2,962 rows, published as sakthai-combined-v10); card also cites sakthai-combined-v7"
training_method: "rsLoRA (rank-stabilized) → merged to full weights"
eval_split: "none documented on card (benchmarks pending)"
lora_config:
r: 16
alpha: 32
dropout: 0.05
target_modules: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]
optimizer: "AdamW (8-bit)"
learning_rate: 0.0002
epochs: 3
precision: bf16
hardware: "Free T4 GPU (Kaggle / Colab / HF Inference), $0 budget"
framework: "TRL + Transformers"
key_improvements:
- "rsLoRA instead of standard LoRA — better rank utilization"
- "All 7 linear modules adapted (vs 4 in v1)"
- "Dropout reduced 0.1 → 0.05"
- "48% more training data (v7 + v8)"
- "Merged full-weight checkpoint — no PEFT dependency at inference"
card_quality:
license: apache-2.0
base_model_documented: true
base_model: Qwen/Qwen2.5-1.5B-Instruct
tags_count: 12
tags: [qwen2.5, sakthai, house-of-sak, tool-calling, function-calling, agent, instruct, finetuned, sft, rslora, text-generation, merged]
datasets_cited: ["Nanthasit/sakthai-combined-v7", "Nanthasit/sakthai-combined-v10"]
model_index_present: false
readme_size_bytes: 8079
widget_example: "none (no widget block in frontmatter)"
deductions:
- "README states 'Benchmarks are pending' but .eval_results/ already carries lighteval + sakthai-bench-v2 scores (unverified) and a verified 3/3 llama.cpp tool-calling run — card text lags repo state"
- "No model-index or widget in README frontmatter — metrics will not render on the Hub widget"
- "0 downloads / 0 likes — brand-new (created 2026-07-30), family table lists it as 'New'"
score: 78
health_score:
overall: 38.2
components:
popularity: 0
momentum: 0
benchmarks: 33.3
card_quality: 78
repo_hygiene: 95
weights:
popularity: 0.20
momentum: 0.20
benchmarks: 0.25
card_quality: 0.20
repo_hygiene: 0.15
sibling_comparison:
rank_by_downloads: 12
total_author_models: 19
max_sibling_downloads: 1599
models_with_positive_downloads: 11
velocity_rank: 12
max_sibling_velocity: 62.56
our_velocity: 0.0
eval_type: metadata_cron
eval_note: >
First cron eval snapshot for sakthai-plus-1.5b (run 16 of hf-eval-updater).
The next-generation SakThai tool-calling flagship: Qwen2.5-1.5B-Instruct
rsLoRA fine-tune (r=16, alpha=32, dropout=0.05) across all 7 linear modules
on sakthai-combined-v10 (v7+v8, 2,962 rows), merged to a full bf16
checkpoint (2.88 GiB). Repo hygiene is excellent (95/100) and config is a
clean Qwen2-1.5B (28 layers, 12 heads / 2 KV, 32K context). The one verified
benchmark — a 2026-07-31 3-trial llama.cpp q4_k_m tool-calling run — passed
3/3 (tool call, valid JSON, correct answer, 20.7 tps). Honest caveats:
README claims 'benchmarks pending' while .eval_results/ already holds 6
unverified lighteval + sakthai-bench-v2 numbers, and the card has no
model-index/widget. Brand-new model: 0 downloads, velocity 0 — sits at
downloads rank 12/19 (11 siblings positive). Next cycle: run
sakthai-bench-v2 properly, add model-index + widget to the card, and update
the 'pending' benchmark text to match repo state.
eval_metadata:
model: Nanthasit/sakthai-plus-1.5b
eval_date: "2026-07-31"
eval_time: "05:40:24Z"
schema: llm_cron_v1
age_days: 0.6961
days_since_last_update: 0.0473
download_velocity: 0.0
cron_run: 16