language, license, library_name, datasets, model-index
language
license
library_name
datasets
model-index
apache-2.0
transformers
totally-not-an-llm/EverythingLM-data-V3
name
results
open-llama-3b-everythingLM-2048
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
AI2 Reasoning Challenge (25-Shot)
ai2_arc
ARC-Challenge
test
type
value
name
acc_norm
42.75
normalized accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
split
args
HellaSwag (10-Shot)
hellaswag
validation
type
value
name
acc_norm
71.72
normalized accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
MMLU (5-Shot)
cais/mmlu
all
test
type
value
name
acc
27.16
accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
TruthfulQA (0-shot)
truthful_qa
multiple_choice
validation
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
Winogrande (5-shot)
winogrande
winogrande_xl
validation
type
value
name
acc
66.3
accuracy
task
dataset
metrics
source
type
name
text-generation
Text Generation
name
type
config
split
args
GSM8k (5-shot)
gsm8k
main
test
type
value
name
acc
1.52
accuracy
Trained on 2 epochs on the EverythingLM-data-V3 dataset.
This model uses the alpaca prompt format:
Detailed results can be found here
Metric
Value
Avg.
40.62
AI2 Reasoning Challenge (25-Shot)
42.75
HellaSwag (10-Shot)
71.72
MMLU (5-Shot)
27.16
TruthfulQA (0-shot)
34.26
Winogrande (5-shot)
66.30
GSM8k (5-shot)
1.52