Files
tinyllama-730M-test/README.md
ModelHub XC a60c1b6753 初始化项目,由ModelHub XC社区提供模型
Model: Josephgflowers/tinyllama-730M-test
Source: Original Platform
2026-08-14 01:09:13 +08:00

3.9 KiB

license, widget, model-index
license widget model-index
mit
text
<|system|> You are a helpful assistant</s> <|user|> What is your name? Tell me about yourself.</s> <|assistant|>
name results
tinyllama-730M-test
task dataset metrics source
type name
text-generation Text Generation
name type config split args
AI2 Reasoning Challenge (25-Shot) ai2_arc ARC-Challenge test
num_few_shot
25
type value name
acc_norm 25.09 normalized accuracy
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard
task dataset metrics source
type name
text-generation Text Generation
name type split args
HellaSwag (10-Shot) hellaswag validation
num_few_shot
10
type value name
acc_norm 33.82 normalized accuracy
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard
task dataset metrics source
type name
text-generation Text Generation
name type config split args
MMLU (5-Shot) cais/mmlu all test
num_few_shot
5
type value name
acc 24.43 accuracy
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard
task dataset metrics source
type name
text-generation Text Generation
name type config split args
TruthfulQA (0-shot) truthful_qa multiple_choice validation
num_few_shot
0
type value
mc2 42.9
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard
task dataset metrics source
type name
text-generation Text Generation
name type config split args
Winogrande (5-shot) winogrande winogrande_xl validation
num_few_shot
5
type value name
acc 51.07 accuracy
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard
task dataset metrics source
type name
text-generation Text Generation
name type config split args
GSM8k (5-shot) gsm8k main test
num_few_shot
5
type value name
acc 0.0 accuracy
url name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=Josephgflowers/tinyllama-730M-test Open LLM Leaderboard

I cut my TinyLlama 1.1B cinder v 2 down from 22 layers to 14. At 14 there was no coherent text but there were emerging ideas of a response. 1000 steps on step-by-step dataset. 6000 on Reason-with-cinder. The loss was still over 1 and the learning rate was still over 4. This model needs significat training. I am putting it up as a base model that needs work. If you continue training please let me know on the tinyllama discord, I have some interesting plans for this model.

Open LLM Leaderboard Evaluation Results

Detailed results can be found here

Metric Value
Avg. 29.55
AI2 Reasoning Challenge (25-Shot) 25.09
HellaSwag (10-Shot) 33.82
MMLU (5-Shot) 24.43
TruthfulQA (0-shot) 42.90
Winogrande (5-shot) 51.07
GSM8k (5-shot) 0.00