Files
ModelHub XC f760bdb35d 初始化项目,由ModelHub XC社区提供模型
Model: OpenLLM-Ro/RoLlama3.1-8b-Instruct-DPO
Source: Original Platform
2026-08-04 16:33:18 +08:00

25 KiB

license, language, base_model, datasets, model-index
license language base_model datasets model-index
cc-by-nc-4.0
ro
OpenLLM-Ro/RoLlama3.1-8b-Instruct-2025-04-23
OpenLLM-Ro/ro_dpo_helpsteer
OpenLLM-Ro/ro_dpo_ultrafeedback
OpenLLM-Ro/ro_dpo_magpie
OpenLLM-Ro/ro_dpo_argilla_magpie
OpenLLM-Ro/ro_dpo_helpsteer2
name results
OpenLLM-Ro/RoLlama3.1-8b-Instruct-DPO-2025-04-23
task dataset metrics
type
text-generation
name type
RoMT-Bench RoMT-Bench
name type value
Score Score 7.00
task dataset metrics
type
text-generation
name type
RoCulturaBench RoCulturaBench
name type value
Score Score 4.73
task dataset metrics
type
text-generation
name type
Romanian_Academic_Benchmarks Romanian_Academic_Benchmarks
name type value
Average accuracy accuracy 53.76
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_arc_challenge OpenLLM-Ro/ro_arc_challenge
name type value
Average accuracy accuracy 51.09
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_mmlu OpenLLM-Ro/ro_mmlu
name type value
Average accuracy accuracy 56.22
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_winogrande OpenLLM-Ro/ro_winogrande
name type value
Average accuracy accuracy 66.77
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_hellaswag OpenLLM-Ro/ro_hellaswag
name type value
Average accuracy accuracy 59.38
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_gsm8k OpenLLM-Ro/ro_gsm8k
name type value
Average accuracy accuracy 31.54
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_truthfulqa OpenLLM-Ro/ro_truthfulqa
name type value
Average accuracy accuracy 57.56
task dataset metrics
type
text-generation
name type
LaRoSeDa_binary LaRoSeDa_binary
name type value
Average macro-f1 macro-f1 96.87
task dataset metrics
type
text-generation
name type
LaRoSeDa_multiclass LaRoSeDa_multiclass
name type value
Average macro-f1 macro-f1 60.75
task dataset metrics
type
text-generation
name type
WMT_EN-RO WMT_EN-RO
name type value
Average bleu bleu 20.30
task dataset metrics
type
text-generation
name type
WMT_RO-EN WMT_RO-EN
name type value
Average bleu bleu 18.57
task dataset metrics
type
text-generation
name type
XQuAD XQuAD
name type value
Average exact_match exact_match 9.22
task dataset metrics
type
text-generation
name type
XQuAD XQuAD
name type value
Average f1 f1 22.75
task dataset metrics
type
text-generation
name type
STS STS
name type value
Average spearman spearman 30.82
task dataset metrics
type
text-generation
name type
STS STS
name type value
Average pearson pearson 20.25
task dataset metrics
type
text-generation
name type
RoMT-Bench RoMT-Bench
name type value
First turn Score 7.30
name type value
Second turn Score 6.70
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_arc_challenge OpenLLM-Ro/ro_arc_challenge
name type value
0-shot accuracy 51.59
name type value
1-shot accuracy 52.10
name type value
3-shot accuracy 50.99
name type value
5-shot accuracy 50.81
name type value
10-shot accuracy 49.70
name type value
25-shot accuracy 51.33
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_mmlu OpenLLM-Ro/ro_mmlu
name type value
0-shot accuracy 56.88
name type value
1-shot accuracy 55.61
name type value
3-shot accuracy 56.06
name type value
5-shot accuracy 56.31
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_winogrande OpenLLM-Ro/ro_winogrande
name type value
0-shot accuracy 65.67
name type value
1-shot accuracy 66.30
name type value
3-shot accuracy 67.40
name type value
5-shot accuracy 67.72
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_hellaswag OpenLLM-Ro/ro_hellaswag
name type value
0-shot accuracy 60.53
name type value
1-shot accuracy 60.37
name type value
3-shot accuracy 58.20
name type value
5-shot accuracy 58.18
name type value
10-shot accuracy 59.61
task dataset metrics
type
text-generation
name type
OpenLLM-Ro/ro_gsm8k OpenLLM-Ro/ro_gsm8k
name type value
1-shot accuracy 25.09
name type value
3-shot accuracy 30.02
name type value
5-shot accuracy 39.50
task dataset metrics
type
text-generation
name type
LaRoSeDa_binary LaRoSeDa_binary
name type value
0-shot macro-f1 95.39
name type value
1-shot macro-f1 95.90
name type value
3-shot macro-f1 98.00
name type value
5-shot macro-f1 98.17
task dataset metrics
type
text-generation
name type
LaRoSeDa_multiclass LaRoSeDa_multiclass
name type value
0-shot macro-f1 60.30
name type value
1-shot macro-f1 64.73
name type value
3-shot macro-f1 58.69
name type value
5-shot macro-f1 59.30
task dataset metrics
type
text-generation
name type
WMT_EN-RO WMT_EN-RO
name type value
0-shot bleu 5.46
name type value
1-shot bleu 26.08
name type value
3-shot bleu 25.90
name type value
5-shot bleu 23.76
task dataset metrics
type
text-generation
name type
WMT_RO-EN WMT_RO-EN
name type value
0-shot bleu 2.74
name type value
1-shot bleu 20.95
name type value
3-shot bleu 31.53
name type value
5-shot bleu 19.05
task dataset metrics
type
text-generation
name type
XQuAD_EM XQuAD_EM
name type value
0-shot exact_match 12.27
name type value
1-shot exact_match 17.98
name type value
3-shot exact_match 5.04
name type value
5-shot exact_match 1.60
task dataset metrics
type
text-generation
name type
XQuAD_F1 XQuAD_F1
name type value
0-shot f1 26.24
name type value
1-shot f1 32.54
name type value
3-shot f1 18.00
name type value
5-shot f1 14.22
task dataset metrics
type
text-generation
name type
STS_Spearman STS_Spearman
name type value
1-shot spearman 76.70
name type value
3-shot spearman 2.82
name type value
5-shot spearman 12.95
task dataset metrics
type
text-generation
name type
STS_Pearson STS_Pearson
name type value
1-shot pearson 77.30
name type value
3-shot pearson -14.56
name type value
5-shot pearson -1.99

Model Card for Model ID

Built with Meta Llama 3.1

This model points/is identical to RoLlama3.1-8b-Instruct-DPO-2025-04-23.

RoLlama3.1 is a family of pretrained and fine-tuned generative text models for Romanian. This is the repository for the human aligned instruct 8B model. Links to other models can be found at the bottom of this page.

Model Details

Model Description

OpenLLM-Ro represents the first open-source effort to build a LLM specialized for Romanian. OpenLLM-Ro developed and publicly releases a collection of Romanian LLMs, both in the form of foundational model and instruct and chat variants.

  • Developed by: OpenLLM-Ro

Model Sources

Intended Use

Intended Use Cases

RoLlama3.1 is intented for research use in Romanian. Base models can be adapted for a variety of natural language tasks while instruction and chat tuned models are intended for assistant-like chat.

Out-of-Scope Use

Use in any manner that violates the license, any applicable laws or regluations, use in languages other than Romanian.

How to Get Started with the Model

Use the code below to get started with the model.

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("OpenLLM-Ro/RoLlama3.1-8b-Instruct-DPO")
model = AutoModelForCausalLM.from_pretrained("OpenLLM-Ro/RoLlama3.1-8b-Instruct-DPO")

instruction = "Ce jocuri de societate pot juca cu prietenii mei?"
chat = [
        {"role": "system", "content": "Ești un asistent folositor, respectuos și onest. Încearcă să ajuți cât mai mult prin informațiile oferite, excluzând răspunsuri toxice, rasiste, sexiste, periculoase și ilegale."},
        {"role": "user", "content": instruction},
        ]
prompt = tokenizer.apply_chat_template(chat, tokenize=False, system_message="")

inputs = tokenizer.encode(prompt, add_special_tokens=False, return_tensors="pt")
outputs = model.generate(input_ids=inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))

Academic Benchmarks

Model Average ARC MMLU Winogrande Hellaswag GSM8k TruthfulQA
Llama-3.1-8B-Instruct49.8742.8653.7359.7156.8235.5650.54
RoLlama3.1-8b-Instruct-2024-10-0953.0347.6954.5765.8459.9444.3045.82
RoLlama3.1-8b-Instruct-2025-04-2353.3648.9755.1766.5260.7342.0346.71
RoLlama3.1-8b-Instruct-DPO-2024-10-0952.7444.8455.0665.8758.6744.1747.82
RoLlama3.1-8b-Instruct-DPO-2025-04-2353.7651.0956.2266.7759.3831.5457.56

Downstream tasks

LaRoSeDa WMT
Few-shot Finetuned Few-shot Finetuned
Model Binary
(Macro F1)
Multiclass
(Macro F1)
Binary
(Macro F1)
Multiclass
(Macro F1)
EN-RO
(Bleu)
RO-EN
(Bleu)
EN-RO
(Bleu)
RO-EN
(Bleu)
Llama-3.1-8B-Instruct95.7459.4998.5782.4119.0127.7729.0239.80
RoLlama3.1-8b-Instruct-2024-10-0994.5660.1095.1287.5321.8823.9928.2740.44
RoLlama3.1-8b-Instruct-2025-04-2395.3260.84--23.1825.11--
RoLlama3.1-8b-Instruct-DPO-2024-10-0996.1055.37--21.2921.86--
RoLlama3.1-8b-Instruct-DPO-2025-04-2396.8760.75--20.3018.57--
XQuAD STS
Few-shot Finetuned Few-shot Finetuned
Model (EM) (F1) (EM) (F1) (Spearman) (Pearson) (Spearman) (Pearson)
Llama-3.1-8B-Instruct44.9664.4569.5084.3172.1171.6484.5984.96
RoLlama3.1-8b-Instruct-2024-10-0913.5923.5649.4162.9375.8976.0086.8687.05
RoLlama3.1-8b-Instruct-2025-04-2310.7419.75--73.5374.93--
RoLlama3.1-8b-Instruct-DPO-2024-10-0921.5836.54--78.0177.98--
RoLlama3.1-8b-Instruct-DPO-2025-04-239.2222.75--30.8220.25--

MT-Bench

Model Average 1st turn 2nd turn Answers in Ro
Llama-3.1-8B-Instruct5.695.855.53160/160
RoLlama3.1-8b-Instruct-2024-10-095.425.954.89160/160
RoLlama3.1-8b-Instruct-2025-04-236.436.786.09160/160
RoLlama3.1-8b-Instruct-DPO-2024-10-096.216.745.69160/160
RoLlama3.1-8b-Instruct-DPO-2025-04-237.007.306.70160/160

RoCulturaBench

Model Average Answers in Ro
Llama-3.1-8B-Instruct3.54100/100
RoLlama3.1-8b-Instruct-2024-10-093.55100/100
RoLlama3.1-8b-Instruct-2025-04-234.28100/100
RoLlama3.1-8b-Instruct-DPO-2024-10-094.42100/100
RoLlama3.1-8b-Instruct-DPO-2025-04-234.73100/100

RoLlama3.1 Model Family

Model Link
RoLlama3.1-8b-Instruct-2024-10-09 link
RoLlama3.1-8b-Instruct-2025-04-23 link
RoLlama3.1-8b-Instruct-DPO-2024-10-09 link
RoLlama3.1-8b-Instruct-DPO-2025-04-23 link

Citation

@inproceedings{masala-etal-2024-vorbesti,
    title = "``Vorbe\c{s}ti Rom{\^a}ne\c{s}te?'' A Recipe to Train Powerful {R}omanian {LLM}s with {E}nglish Instructions",
    author = "Masala, Mihai and Ilie-Ablachim, Denis and Dima, Alexandru and Corlatescu, Dragos Georgian and Zavelca, Miruna-Andreea and Olaru, Ovio and Terian, Simina-Maria and Terian, Andrei and Leordeanu, Marius and Velicu, Horia and Popescu, Marius and Dascalu, Mihai and Rebedea, Traian",
    editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.681/",
    doi = "10.18653/v1/2024.findings-emnlp.681",
    pages = "11632--11647"
}