license, widget, model-index
license
widget
model-index
mit
text
<|system|>
You are a helpful assistant</s>
<|user|>
What is your name? Tell me about yourself.</s>
<|assistant|>
name
results
tinyllama-730M-test
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
AI2 Reasoning Challenge (25-Shot)
ai2_arc
ARC-Challenge
test
type
value
name
acc_norm
25.09
normalized accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
split
args
HellaSwag (10-Shot)
hellaswag
validation
type
value
name
acc_norm
33.82
normalized accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
MMLU (5-Shot)
cais/mmlu
all
test
type
value
name
acc
24.43
accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
TruthfulQA (0-shot)
truthful_qa
multiple_choice
validation
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
Winogrande (5-shot)
winogrande
winogrande_xl
validation
type
value
name
acc
51.07
accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
GSM8k (5-shot)
gsm8k
main
test
type
value
name
acc
0.0
accuracy
I cut my TinyLlama 1.1B cinder v 2 down from 22 layers to 14. At 14 there was no coherent text but there were emerging ideas of a response. 1000 steps on step-by-step dataset.
6000 on Reason-with-cinder. The loss was still over 1 and the learning rate was still over 4. This model needs significat training. I am putting it up as a base model that
needs work. If you continue training please let me know on the tinyllama discord, I have some interesting plans for this model.
Detailed results can be found here
Metric
Value
Avg.
29.55
AI2 Reasoning Challenge (25-Shot)
25.09
HellaSwag (10-Shot)
33.82
MMLU (5-Shot)
24.43
TruthfulQA (0-shot)
42.90
Winogrande (5-shot)
51.07
GSM8k (5-shot)
0.00