初始化项目，由ModelHub XC社区提供模型

Model: prometheus-eval/prometheus-7b-v2.0 Source: Original Platform
2026-05-21 14:24:19 +08:00
commit 5d644f2848
18 changed files with 91483 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,35 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,182 @@
 ---
 tags:
 - text2text-generation
 datasets:
 - prometheus-eval/Feedback-Collection
 - prometheus-eval/Preference-Collection
 license: apache-2.0
 language:
 - en
 pipeline_tag: text2text-generation
 library_name: transformers
 metrics:
 - pearsonr
 - spearmanr
 - kendall-tau
 - accuracy
 ---
 ## Links for Reference
 - **Homepage: In Progress** 
 - **Repository:https://github.com/prometheus-eval/prometheus-eval** 
 - **Paper:https://arxiv.org/abs/2405.01535** 
 - **Point of Contact:seungone@cmu.edu** 
 # TL;DR
 Prometheus 2 is an alternative of GPT-4 evaluation when doing fine-grained evaluation of an underlying LLM & a Reward model for Reinforcement Learning from Human Feedback (RLHF).
 ![plot](./finegrained_eval.JPG)
 Prometheus 2 is a language model using [Mistral-Instruct](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2) as a base model. 
 It is fine-tuned on 100K feedback within the [Feedback Collection](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and 200K feedback within the [Preference Collection](https://huggingface.co/datasets/prometheus-eval/Preference-Collection).
 It is also made by weight merging to support both absolute grading (direct assessment) and relative grading (pairwise ranking).
 The surprising thing is that we find weight merging also improves performance on each format.
 # Model Details
 ## Model Description
 - **Model type:** Language model
 - **Language(s) (NLP):** English
 - **License:** Apache 2.0
 - **Related Models:** [All Prometheus Checkpoints](https://huggingface.co/models?search=prometheus-eval/Prometheus)
 - **Resources for more information:**
  - [Research paper](https://arxiv.org/abs/2405.01535)
  - [GitHub Repo](https://github.com/prometheus-eval/prometheus-eval)
 Prometheus is trained with two different sizes (7B and 8x7B).
 You could check the 8x7B sized LM on [this page](https://huggingface.co/prometheus-eval/prometheus-2-8x7b-v2.0).
 Also, check out our dataset as well on [this page](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and [this page](https://huggingface.co/datasets/prometheus-eval/Preference-Collection).
 ## Prompt Format
 We have made wrapper functions and classes to conveniently use Prometheus 2 at [our github repository](https://github.com/prometheus-eval/prometheus-eval).
 We highly recommend you use it!
 However, if you just want to use the model for your use case, please refer to the prompt format below.
 Note that absolute grading and relative grading requires different prompt templates and system prompts.
 ### Absolute Grading (Direct Assessment)
 Prometheus requires 4 components in the input: An instruction, a response to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below.
 You should fill in the instruction, response, reference answer, criteria description, and score description for score in range of 1 to 5.
 Fix the components with \{text\} inside.
 ```
 ###Task Description:
 An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given.
 1. Write a detailed feedback that assess the quality of the response strictly based on the given score rubric, not evaluating in general.
 2. After writing a feedback, write a score that is an integer between 1 and 5. You should refer to the score rubric.
 3. The output format should look as follows: \"Feedback: (write a feedback for criteria) [RESULT] (an integer number between 1 and 5)\"
 4. Please do not generate any other opening, closing, and explanations.
 ###The instruction to evaluate:
 {orig_instruction}
 ###Response to evaluate:
 {orig_response}
 ###Reference Answer (Score 5):
 {orig_reference_answer}
 ###Score Rubrics:
 [{orig_criteria}]
 Score 1: {orig_score1_description}
 Score 2: {orig_score2_description}
 Score 3: {orig_score3_description}
 Score 4: {orig_score4_description}
 Score 5: {orig_score5_description}
 ###Feedback: 
 ```
 After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors).
 You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).
 ```
 conv = get_conv_template("mistral")
 conv.set_system_message("You are a fair judge assistant tasked with providing clear, objective feedback based on specific criteria, ensuring each assessment reflects the absolute standards set for performance.")
 conv.append_message(conv.roles[0], dialogs['instruction'])
 conv.append_message(conv.roles[1], None)
 prompt = conv.get_prompt()
 x = tokenizer(prompt,truncation=False)
 ```
 As a result, a feedback and score decision will be generated, divided by a separating phrase ```[RESULT]``` 
 ### Relative Grading (Pairwise Ranking)
 Prometheus requires 4 components in the input: An instruction, 2 responses to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below.
 You should fill in the instruction, 2 responses, reference answer, and criteria description.
 Fix the components with \{text\} inside.
 ```
 ###Task Description:
 An instruction (might include an Input inside it), two responses to evaluate (denoted as Response A and Response B), a reference answer, and an evaluation criteria are given.
 1. Write a detailed feedback that assess the quality of the two responses strictly based on the given evaluation criteria, not evaluating in general.
 2. Make comparisons between Response A, Response B, and the Reference Answer. Instead of examining Response A and Response B separately, go straight to the point and mention about the commonalities and differences between them.
 3. After writing the feedback, indicate the better response, either "A" or "B".
 4. The output format should look as follows: "Feedback: (write a feedback for criteria) [RESULT] (Either "A" or "B")"
 5. Please do not generate any other opening, closing, and explanations.
 ###Instruction:
 {orig_instruction}
 ###Response A:
 {orig_response_A}
 ###Response B:
 {orig_response_B}
 ###Reference Answer:
 {orig_reference_answer}
 ###Score Rubric:
 {orig_criteria}
 ###Feedback: 
 ```
 After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors).
 You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).
 ```
 conv = get_conv_template("mistral")
 conv.set_system_message("You are a fair judge assistant assigned to deliver insightful feedback that compares individual performances, highlighting how each stands relative to others within the same cohort.")
 conv.append_message(conv.roles[0], dialogs['instruction'])
 conv.append_message(conv.roles[1], None)
 prompt = conv.get_prompt()
 x = tokenizer(prompt,truncation=False)
 ```
 As a result, a feedback and score decision will be generated, divided by a separating phrase ```[RESULT]``` 
 ## License
 Feedback Collection, Preference Collection, and Prometheus 2 are subject to OpenAI's Terms of Use for the generated data. If you suspect any violations, please reach out to us.
 # Citation
 If you find the following model helpful, please consider citing our paper!
 **BibTeX:**
 ```bibtex
@misc{kim2023prometheus,
    title={Prometheus: Inducing Fine-grained Evaluation Capability in Language Models},
    author={Seungone Kim and Jamin Shin and Yejin Cho and Joel Jang and Shayne Longpre and Hwaran Lee and Sangdoo Yun and Seongjin Shin and Sungdong Kim and James Thorne and Minjoon Seo},
    year={2023},
    eprint={2310.08491},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
 }
 ```
 ```bibtex
@misc{kim2024prometheus,
    title={Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models},
    author={Seungone Kim and Juyoung Suk and Shayne Longpre and Bill Yuchen Lin and Jamin Shin and Sean Welleck and Graham Neubig and Moontae Lee and Kyungjae Lee and Minjoon Seo},
    year={2024},
    eprint={2405.01535},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
 }
 ```
--- a/config.json
+++ b/config.json
@@ -0,0 +1,26 @@
 {
  "_name_or_path": "kaist-ai/prometheus-7b-v1.5-beta-1",
  "architectures": [
    "MistralForCausalLM"
  ],
  "attention_dropout": 0.0,
  "bos_token_id": 1,
  "eos_token_id": 2,
  "hidden_act": "silu",
  "hidden_size": 4096,
  "initializer_range": 0.02,
  "intermediate_size": 14336,
  "max_position_embeddings": 32768,
  "model_type": "mistral",
  "num_attention_heads": 32,
  "num_hidden_layers": 32,
  "num_key_value_heads": 8,
  "rms_norm_eps": 1e-05,
  "rope_theta": 1000000.0,
  "sliding_window": null,
  "tie_word_embeddings": false,
  "torch_dtype": "bfloat16",
  "transformers_version": "4.35.2",
  "use_cache": true,
  "vocab_size": 32000
 }
--- a/finegrained_eval.JPG
+++ b/finegrained_eval.JPG
--- a/mergekit_config.yml
+++ b/mergekit_config.yml
@@ -0,0 +1,10 @@
 models:
  - model: kaist-ai/prometheus-7b-v1.5-beta-1
    parameters:
      weight: 1.0
  - model: kaist-ai/prometheus-7b-v1.5-beta-2
    parameters:
      weight: 1.0
 merge_method: linear
 dtype: bfloat16
--- a/model-00001-of-00008.safetensors
+++ b/model-00001-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:26afbe36c7d04dd26488eff13c5b7975f2e2a658eb6f3517c3f58e3c61b2ba2d
 size 1979773128
--- a/model-00002-of-00008.safetensors
+++ b/model-00002-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:68754ee9a08370759db5ccfa647915f3c6dca186d587a34b8223a32ba6723d54
 size 1946235640
--- a/model-00003-of-00008.safetensors
+++ b/model-00003-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:7044141b330ff94fd52f1fb18799156071ea212b045050f9f66b16f010867a58
 size 1973490216
--- a/model-00004-of-00008.safetensors
+++ b/model-00004-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:b6f73b31c39c598501ebb1c9237c3efcac3798ea2b77c07e81b5f1d5a8a82f56
 size 1979781464
--- a/model-00005-of-00008.safetensors
+++ b/model-00005-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:1ff853ff8692412e3bc26acccc0c0b3430b8a11c531fcbd3a811f7410be2679c
 size 1946243984
--- a/model-00006-of-00008.safetensors
+++ b/model-00006-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:c3bedf39d42e46871e188bc50003f28a6100c82109c822b9cdeab3fc3468db45
 size 1923166040
--- a/model-00007-of-00008.safetensors
+++ b/model-00007-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:96db009cb729bdec84a6df74d4f7413479f4ab98ff8df00a75ebbe9a93ae05de
 size 1946243984
--- a/model-00008-of-00008.safetensors
+++ b/model-00008-of-00008.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:714069c1ba0aa520a27a89ffc45cf1bc2af0adfaf3b91deaed5b7e95b51fbf06
 size 788563544
--- a/model.safetensors.index.json
+++ b/model.safetensors.index.json
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
@@ -0,0 +1,30 @@
 {
  "bos_token": {
    "content": "<s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "eos_token": {
    "content": "</s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "pad_token": {
    "content": "</s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "unk_token": {
    "content": "<unk>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  }
 }
--- a/tokenizer.json
+++ b/tokenizer.json
--- a/tokenizer.model
+++ b/tokenizer.model
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,45 @@
 {
  "added_tokens_decoder": {
    "0": {
      "content": "<unk>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "1": {
      "content": "<s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "2": {
      "content": "</s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    }
  },
  "additional_special_tokens": [],
  "bos_token": "<s>",
  "chat_template": "{%- if messages[0]['role'] == 'system' %}\n    {%- set system_message = messages[0]['content'] %}\n    {%- set loop_messages = messages[1:] %}\n{%- else %}\n    {%- set loop_messages = messages %}\n{%- endif %}\n\n{{- bos_token }}\n{%- for message in loop_messages %}\n    {%- if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}\n        {{- raise_exception('After the optional system message, conversation roles must alternate user/assistant/user/assistant/...') }}\n    {%- endif %}\n    {%- if message['role'] == 'user' %}\n        {%- if loop.first and system_message is defined %}\n            {{- ' [INST] ' + system_message + '\\n\\n' + message['content'] + ' [/INST]' }}\n        {%- else %}\n            {{- ' [INST] ' + message['content'] + ' [/INST]' }}\n        {%- endif %}\n    {%- elif message['role'] == 'assistant' %}\n        {{- ' ' + message['content'] + eos_token}}\n    {%- else %}\n        {{- raise_exception('Only user and assistant roles are supported, with the exception of an initial optional system message!') }}\n    {%- endif %}\n{%- endfor %}\n",
  "clean_up_tokenization_spaces": false,
  "eos_token": "</s>",
  "legacy": true,
  "max_length": 4096,
  "model_max_length": 1000000000000000019884624838656,
  "pad_token": "</s>",
  "sp_model_kwargs": {},
  "spaces_between_special_tokens": false,
  "stride": 0,
  "tokenizer_class": "LlamaTokenizer",
  "truncation_side": "right",
  "truncation_strategy": "longest_first",
  "unk_token": "<unk>",
  "use_default_system_prompt": false
 }