Files
RoLlama3-8b-Instruct-2024-0…/README.md
ModelHub XC 9232bdd8e5 初始化项目,由ModelHub XC社区提供模型
Model: OpenLLM-Ro/RoLlama3-8b-Instruct-2024-06-28
Source: Original Platform
2026-07-28 04:45:13 +08:00

26 KiB

license, language, base_model, datasets, model-index
license language base_model datasets model-index
cc-by-nc-4.0
ro
meta-llama/Meta-Llama-3-8B
OpenLLM-Ro/ro_sft_alpaca
OpenLLM-Ro/ro_sft_alpaca_gpt4
OpenLLM-Ro/ro_sft_dolly
OpenLLM-Ro/ro_sft_selfinstruct_gpt4
OpenLLM-Ro/ro_sft_norobots
OpenLLM-Ro/ro_sft_orca
OpenLLM-Ro/ro_sft_camel
name results
OpenLLM-Ro/RoLlama3-8b-Instruct-2024-06-28
task dataset metrics
type
text-generation
name type
RoMT-Bench RoMT-Bench
name type value
Score Score 5.15
task dataset metrics
type
text-generation
name type
RoCulturaBench RoCulturaBench
name type value
Score Score 3.71
task dataset metrics
type
text-generation
name type
Romanian_Academic_Benchmarks Romanian_Academic_Benchmarks
name type value
Average accuracy accuracy 50.56
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_arc_challenge OpenLLM-Ro/ro_arc_challenge
name type value
Average accuracy accuracy 44.70
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_mmlu OpenLLM-Ro/ro_mmlu
name type value
Average accuracy accuracy 52.19
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_winogrande OpenLLM-Ro/ro_winogrande
name type value
Average accuracy accuracy 67.23
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_hellaswag OpenLLM-Ro/ro_hellaswag
name type value
Average accuracy accuracy 57.69
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_gsm8k OpenLLM-Ro/ro_gsm8k
name type value
Average accuracy accuracy 30.23
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_truthfulqa OpenLLM-Ro/ro_truthfulqa
name type value
Average accuracy accuracy 51.34
task dataset metrics
type
text-generation
name type
LaRoSeDa_binary LaRoSeDa_binary
name type value
Average macro-f1 macro-f1 97.52
task dataset metrics
type
text-generation
name type
LaRoSeDa_multiclass LaRoSeDa_multiclass
name type value
Average macro-f1 macro-f1 67.41
task dataset metrics
type
text-generation
name type
LaRoSeDa_binary_finetuned LaRoSeDa_binary_finetuned
name type value
Average macro-f1 macro-f1 94.15
task dataset metrics
type
text-generation
name type
LaRoSeDa_multiclass_finetuned LaRoSeDa_multiclass_finetuned
name type value
Average macro-f1 macro-f1 87.13
task dataset metrics
type
text-generation
name type
WMT_EN-RO WMT_EN-RO
name type value
Average bleu bleu 24.01
task dataset metrics
type
text-generation
name type
WMT_RO-EN WMT_RO-EN
name type value
Average bleu bleu 27.36
task dataset metrics
type
text-generation
name type
WMT_EN-RO_finetuned WMT_EN-RO_finetuned
name type value
Average bleu bleu 26.53
task dataset metrics
type
text-generation
name type
WMT_RO-EN_finetuned WMT_RO-EN_finetuned
name type value
Average bleu bleu 40.36
task dataset metrics
type
text-generation
name type
XQuAD XQuAD
name type value
Average exact_match exact_match 39.43
task dataset metrics
type
text-generation
name type
XQuAD XQuAD
name type value
Average f1 f1 59.50
task dataset metrics
type
text-generation
name type
XQuAD_finetuned XQuAD_finetuned
name type value
Average exact_match exact_match 44.45
task dataset metrics
type
text-generation
name type
XQuAD_finetuned XQuAD_finetuned
name type value
Average f1 f1 59.76
task dataset metrics
type
text-generation
name type
STS STS
name type value
Average spearman spearman 77.20
task dataset metrics
type
text-generation
name type
STS STS
name type value
Average pearson pearson 77.87
task dataset metrics
type
text-generation
name type
STS_finetuned STS_finetuned
name type value
Average spearman spearman 85.80
task dataset metrics
type
text-generation
name type
STS_finetuned STS_finetuned
name type value
Average pearson pearson 86.05
task dataset metrics
type
text-generation
name type
RoMT-Bench RoMT-Bench
name type value
First turn Score 6.03
name type value
Second turn Score 4.28
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_arc_challenge OpenLLM-Ro/ro_arc_challenge
name type value
0-shot accuracy 41.90
name type value
1-shot accuracy 44.30
name type value
3-shot accuracy 44.56
name type value
5-shot accuracy 45.50
name type value
10-shot accuracy 46.10
name type value
25-shot accuracy 45.84
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_mmlu OpenLLM-Ro/ro_mmlu
name type value
0-shot accuracy 50.85
name type value
1-shot accuracy 51.24
name type value
3-shot accuracy 53.30
name type value
5-shot accuracy 53.39
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_winogrande OpenLLM-Ro/ro_winogrande
name type value
0-shot accuracy 65.19
name type value
1-shot accuracy 66.54
name type value
3-shot accuracy 67.88
name type value
5-shot accuracy 69.30
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_hellaswag OpenLLM-Ro/ro_hellaswag
name type value
0-shot accuracy 56.12
name type value
1-shot accuracy 57.37
name type value
3-shot accuracy 57.92
name type value
5-shot accuracy 58.18
name type value
10-shot accuracy 58.85
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_gsm8k OpenLLM-Ro/ro_gsm8k
name type value
1-shot accuracy 29.42
name type value
3-shot accuracy 30.02
name type value
5-shot accuracy 31.24
task dataset metrics
type
text-generation
name type
LaRoSeDa_binary LaRoSeDa_binary
name type value
0-shot macro-f1 97.43
name type value
1-shot macro-f1 96.60
name type value
3-shot macro-f1 97.90
name type value
5-shot macro-f1 98.13
task dataset metrics
type
text-generation
name type
LaRoSeDa_multiclass LaRoSeDa_multiclass
name type value
0-shot macro-f1 63.77
name type value
1-shot macro-f1 68.91
name type value
3-shot macro-f1 66.36
name type value
5-shot macro-f1 70.61
task dataset metrics
type
text-generation
name type
WMT_EN-RO WMT_EN-RO
name type value
0-shot bleu 6.92
name type value
1-shot bleu 29.33
name type value
3-shot bleu 29.79
name type value
5-shot bleu 30.02
task dataset metrics
type
text-generation
name type
WMT_RO-EN WMT_RO-EN
name type value
0-shot bleu 4.50
name type value
1-shot bleu 30.30
name type value
3-shot bleu 36.96
name type value
5-shot bleu 37.70
task dataset metrics
type
text-generation
name type
XQuAD_EM XQuAD_EM
name type value
0-shot exact_match 4.45
name type value
1-shot exact_match 48.24
name type value
3-shot exact_match 52.03
name type value
5-shot exact_match 53.03
task dataset metrics
type
text-generation
name type
XQuAD_F1 XQuAD_F1
name type value
0-shot f1 26.08
name type value
1-shot f1 68.40
name type value
3-shot f1 71.92
name type value
5-shot f1 71.60
task dataset metrics
type
text-generation
name type
STS_Spearman STS_Spearman
name type value
1-shot spearman 77.76
name type value
3-shot spearman 76.72
name type value
5-shot spearman 77.12
task dataset metrics
type
text-generation
name type
STS_Pearson STS_Pearson
name type value
1-shot pearson 77.83
name type value
3-shot pearson 77.64
name type value
5-shot pearson 78.13

Model Card for Model ID

Built with Meta Llama 3

RoLlama3 is a family of pretrained and fine-tuned generative text models for Romanian. This is the repository for the instruct 8B model. Links to other models can be found at the bottom of this page.

Model Details

Model Description

OpenLLM-Ro represents the first open-source effort to build a LLM specialized for Romanian. OpenLLM-Ro developed and publicly releases a collection of Romanian LLMs, both in the form of foundational model and instruct and chat variants.

  • Developed by: OpenLLM-Ro

Model Sources

Intended Use

Intended Use Cases

RoLlama3 is intented for research use in Romanian. Base models can be adapted for a variety of natural language tasks while instruction and chat tuned models are intended for assistant-like chat.

Out-of-Scope Use

Use in any manner that violates the license, any applicable laws or regluations, use in languages other than Romanian.

How to Get Started with the Model

Use the code below to get started with the model.

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("OpenLLM-Ro/RoLlama3-8b-Instruct-2024-06-28")
model = AutoModelForCausalLM.from_pretrained("OpenLLM-Ro/RoLlama3-8b-Instruct-2024-06-28")

instruction = "Ce jocuri de societate pot juca cu prietenii mei?"
chat = [
        {"role": "system", "content": "Ești un asistent folositor, respectuos și onest. Încearcă să ajuți cât mai mult prin informațiile oferite, excluzând răspunsuri toxice, rasiste, sexiste, periculoase și ilegale."},
        {"role": "user", "content": instruction},
        ]
prompt = tokenizer.apply_chat_template(chat, tokenize=False, system_message="")

inputs = tokenizer.encode(prompt, add_special_tokens=False, return_tensors="pt")
outputs = model.generate(input_ids=inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))

Academic Benchmarks

Model Average ARC MMLU Winogrande Hellaswag GSM8k TruthfulQA
Llama-3-8B-Instruct50.6243.6952.0459.3353.1943.8751.59
RoLlama3-8b-Instruct-2024-06-2850.5644.7052.1967.2357.6930.2351.34
RoLlama3-8b-Instruct-2024-10-0952.2147.9453.5066.0659.7240.1645.90
RoLlama3-8b-Instruct-DPO-2024-10-0949.9646.2953.2965.5758.1534.7741.70

Downstream tasks

LaRoSeDa WMT
Few-shot Finetuned Few-shot Finetuned
Model Binary
(Macro F1)
Multiclass
(Macro F1)
Binary
(Macro F1)
Multiclass
(Macro F1)
EN-RO
(Bleu)
RO-EN
(Bleu)
EN-RO
(Bleu)
RO-EN
(Bleu)
Llama-3-8B-Instruct95.8856.2198.5386.1918.8830.9828.0240.28
RoLlama3-8b-Instruct-2024-06-2897.5267.4194.1587.1324.0127.3626.5340.36
RoLlama3-8b-Instruct-2024-10-0995.5861.2096.4687.2622.9224.2827.3140.52
RoLlama3-8b-Instruct-DPO-2024-10-0997.4854.00--22.0923.00--
XQuAD STS
Few-shot Finetuned Few-shot Finetuned
Model (EM) (F1) (EM) (F1) (Spearman) (Pearson) (Spearman) (Pearson)
Llama-3-8B-Instruct39.4758.6767.6582.7773.0472.3683.4984.06
RoLlama3-8b-Instruct-2024-06-2839.4359.5044.4559.7677.2077.8785.8086.05
RoLlama3-8b-Instruct-2024-10-0918.8931.7950.8465.1877.6076.8686.7087.09
RoLlama3-8b-Instruct-DPO-2024-10-0926.0542.77--79.6479.52--

MT-Bench

Model Average 1st turn 2nd turn Answers in Ro
Llama-3-8B-Instruct5.966.165.76158/160
RoLlama3-8b-Instruct-2024-06-285.156.034.28160/160
RoLlama3-8b-Instruct-2024-10-095.386.094.67160/160
RoLlama3-8b-Instruct-DPO-2024-10-095.876.225.49160/160

RoCulturaBench

Model Average Answers in Ro
Llama-3-8B-Instruct4.62100/100
RoLlama3-8b-Instruct-2024-06-283.71100/100
RoLlama3-8b-Instruct-2024-10-093.81100/100
RoLlama3-8b-Instruct-DPO-2024-10-094.40100/100

RoLlama3 Model Family

Model Link
RoLlama3-8b-Instruct-2024-06-28 link
RoLlama3-8b-Instruct-2024-10-09 link
RoLlama3-8b-Instruct-DPO-2024-10-09 link

Citation

@inproceedings{masala-etal-2024-vorbesti,
    title = "``Vorbe\c{s}ti Rom{\^a}ne\c{s}te?'' A Recipe to Train Powerful {R}omanian {LLM}s with {E}nglish Instructions",
    author = "Masala, Mihai and Ilie-Ablachim, Denis and Dima, Alexandru and Corlatescu, Dragos Georgian and Zavelca, Miruna-Andreea and Olaru, Ovio and Terian, Simina-Maria and Terian, Andrei and Leordeanu, Marius and Velicu, Horia and Popescu, Marius and Dascalu, Mihai and Rebedea, Traian",
    editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.681/",
    doi = "10.18653/v1/2024.findings-emnlp.681",
    pages = "11632--11647"
}