commit cd82c37be3afffc90b04e2cdaccb216409d06271 Author: ModelHub XC Date: Fri Oct 2 07:09:18 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: adgomant/adele-judge-qwen3-14B-cre Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..dffe933 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,37 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +adapter/tokenizer.json filter=lfs diff=lfs merge=lfs -text +tokenizer.json filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..2cd3fc0 --- /dev/null +++ b/README.md @@ -0,0 +1,151 @@ +--- +library_name: transformers +tags: +- text-generation +- peft +- adele +- judge +base_model: Qwen/Qwen3-14B +datasets: +- CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0 +--- + +# ADeLe Distilled Judge + +This repository contains an ADeLe-suite-specific distilled judge. It scores a model response against a question and reference answer with an ordinal score from 1 to 5, then derives binary correctness with the ADeLe threshold. + +The repository root contains a merged Transformers model for standard loading. The original LoRA adapter is also included under `adapter/` for provenance and reuse. + +## Intended Use + +Use this model to score ADeLe-style examples where a question, reference answer, and model response are available. It is intended for out-of-model evaluation within the ADeLe benchmark suite, not as a general-purpose evaluator. + +## Input Format + +The recommended helper accepts: + +- `question` +- `reference_answer` or `ground_truth` +- `model_response` + +## Score Rubric + +Allowed scores: 1, 2, 3, 4, 5 + +- 1: surely incorrect +- 2: likely incorrect +- 3: minimally correct or sufficient +- 4: likely correct +- 5: surely correct + +Binary label: scores greater than or equal to 3 are `CORRECT`; lower scores are `INCORRECT`. + +## Training And Validation Data + +| Split | Examples | Models | +| --- | --- | --- | +| train | 239,420 | 16 | +| validation | 45,738 | 3 | + +- `train` models: `DK-R1-Dist-Qwen-1.5B`, `DK-R1-Dist-Qwen-32B`, `DK-R1-Dist-Qwen-7B`, `gemini-2.5-flash`, `gemini-3.1-pro`, `gpt-35-turbo`, `gpt-5.2`, `gpt4o`, `llama3d1-405b`, `llama3d2-11b`, `llama3d2-1b`, `llama3d2-90b`, `llama4-17B-128E`, `o1-mini`, `o1_re=low`, `o3-mini` +- `validation` models: `DK-R1-Dist-Qwen-14B`, `gemini-3-flash`, `llama3d2-3b` + +## Data Quality And Label Construction + +Training labels are distilled from two proprietary judge scores used by the ADeLe evaluation pipeline to derive the official correctness signal. The configured source columns are `score_gpt4o` and `score_sonnet`. + +- Ordinal target: `floor(mean(score_gpt4o, score_sonnet))`. +- Binary target: `CORRECT` when the ordinal target is >= `3`. +- Judge-agreement filter: keep examples with `abs(score_gpt4o - score_sonnet) <= 1`. +- Response-length filter: keep responses with at most `4096` base-tokenizer tokens before prompt formatting. +- Sequence-length filter: keep full chat-formatted examples within `max_seq_length=8192`. + +## Validation Results + +Source artifact: `validation_trainer_metrics.json`. + +| Metric | Value | +| --- | --- | +| Epoch | 1.0000 | +| Binary accuracy | 0.9894 | +| Binary macro F1 | 0.9880 | +| Precision, CORRECT | 0.9932 | +| Recall, CORRECT | 0.9909 | +| Precision, INCORRECT | 0.9817 | +| Recall, INCORRECT | 0.9863 | +| False negative rate, CORRECT | 0.0091 | +| False positive rate, CORRECT | 0.0137 | +| Ordinal accuracy | 0.9639 | +| Ordinal macro F1 | 0.7351 | +| Mean confidence | 0.9604 | + +## Recommended Inference + +Do not use free-form generation as the primary prediction method. The recommended path scores the restricted continuations `"1"`, `"2"`, `"3"`, `"4"`, and `"5"`. + +```python +from transformers import pipeline + +judge = pipeline( + "adele-judge", + model="adgomant/adele-judge-qwen3-14-cre", + trust_remote_code=True, + device_map="auto", +) +result = judge( + {"question": "...", "reference_answer": "...", "model_response": "..."} +) +print(result) + +results = judge([ + {"question": "...", "reference_answer": "...", "model_response": "..."}, + {"question": "...", "ground_truth": "...", "model_response": "..."}, +], batch_size=8) +``` + +The result has this shape: + +```python +{ + "score": 4, + "label": "CORRECT", + "probs": {"1": 0.01, "2": 0.02, "3": 0.08, "4": 0.70, "5": 0.19}, + "logprobs": {"1": -5.0, "2": -4.2, "3": -2.9, "4": -0.8, "5": -2.1}, + "confidence": 0.70, + "margin": 1.3, + "entropy": 0.82, +} +``` + +## Standard Transformers Loading + +```python +from transformers import AutoModelForCausalLM, AutoTokenizer + +tokenizer = AutoTokenizer.from_pretrained("adgomant/adele-judge-qwen3-14-cre", trust_remote_code=True) +model = AutoModelForCausalLM.from_pretrained("adgomant/adele-judge-qwen3-14-cre", trust_remote_code=True) +``` + +`generation_config.json` uses safe one-token defaults for debugging, but `generate()` is not the recommended scoring method. + +## Metadata + +Training, filtering, split, tokenization, and metric artifacts available at packaging time are stored in `adele_judge_metadata.json`. + +The model is trained on distilled judge targets. These targets are useful for reproducing the ADeLe paper-style correctness signal at lower inference cost, but they should not be interpreted as independent human annotations. + +## References + +- ADeLe project page: [ADeLe v1.0](https://kinds-of-intelligence-cfi.github.io/ADELE/). +- ADeLe paper and official correctness definition: [General scales unlock AI evaluation with explanatory and predictive power](https://www.nature.com/articles/s41586-026-10303-2). +- Official ADeLe dataset: [CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0](https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0). +- Official instance-level model-response data used for distillation: [https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation/tree/main/ADeLe_battery_data/subject_specific_instance_level_data](https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation/tree/main/ADeLe_battery_data/subject_specific_instance_level_data). +- Training and Hub packaging implementation: [https://github.com/adgomant/adele-judge](https://github.com/adgomant/adele-judge). + +## Limitations + +- ADeLe-specific judge; not a general-purpose evaluator. +- Distilled from proprietary judge labels and inherits their noise, calibration, and biases. +- Intended for scoring responses against a reference answer. +- It should not produce explanations; the expected output is a single score. +- Validation is out-of-model within the ADeLe suite, so transfer outside that suite should be measured before relying on it. \ No newline at end of file diff --git a/adapter/README.md b/adapter/README.md new file mode 100644 index 0000000..7caaeb8 --- /dev/null +++ b/adapter/README.md @@ -0,0 +1,207 @@ +--- +base_model: Qwen/Qwen3-14B +library_name: peft +pipeline_tag: text-generation +tags: +- base_model:adapter:Qwen/Qwen3-14B +- lora +- transformers +--- + +# Model Card for Model ID + + + + + +## Model Details + +### Model Description + + + + + +- **Developed by:** [More Information Needed] +- **Funded by [optional]:** [More Information Needed] +- **Shared by [optional]:** [More Information Needed] +- **Model type:** [More Information Needed] +- **Language(s) (NLP):** [More Information Needed] +- **License:** [More Information Needed] +- **Finetuned from model [optional]:** [More Information Needed] + +### Model Sources [optional] + + + +- **Repository:** [More Information Needed] +- **Paper [optional]:** [More Information Needed] +- **Demo [optional]:** [More Information Needed] + +## Uses + + + +### Direct Use + + + +[More Information Needed] + +### Downstream Use [optional] + + + +[More Information Needed] + +### Out-of-Scope Use + + + +[More Information Needed] + +## Bias, Risks, and Limitations + + + +[More Information Needed] + +### Recommendations + + + +Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations. + +## How to Get Started with the Model + +Use the code below to get started with the model. + +[More Information Needed] + +## Training Details + +### Training Data + + + +[More Information Needed] + +### Training Procedure + + + +#### Preprocessing [optional] + +[More Information Needed] + + +#### Training Hyperparameters + +- **Training regime:** [More Information Needed] + +#### Speeds, Sizes, Times [optional] + + + +[More Information Needed] + +## Evaluation + + + +### Testing Data, Factors & Metrics + +#### Testing Data + + + +[More Information Needed] + +#### Factors + + + +[More Information Needed] + +#### Metrics + + + +[More Information Needed] + +### Results + +[More Information Needed] + +#### Summary + + + +## Model Examination [optional] + + + +[More Information Needed] + +## Environmental Impact + + + +Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700). + +- **Hardware Type:** [More Information Needed] +- **Hours used:** [More Information Needed] +- **Cloud Provider:** [More Information Needed] +- **Compute Region:** [More Information Needed] +- **Carbon Emitted:** [More Information Needed] + +## Technical Specifications [optional] + +### Model Architecture and Objective + +[More Information Needed] + +### Compute Infrastructure + +[More Information Needed] + +#### Hardware + +[More Information Needed] + +#### Software + +[More Information Needed] + +## Citation [optional] + + + +**BibTeX:** + +[More Information Needed] + +**APA:** + +[More Information Needed] + +## Glossary [optional] + + + +[More Information Needed] + +## More Information [optional] + +[More Information Needed] + +## Model Card Authors [optional] + +[More Information Needed] + +## Model Card Contact + +[More Information Needed] +### Framework versions + +- PEFT 0.19.1 \ No newline at end of file diff --git a/adapter/adapter_config.json b/adapter/adapter_config.json new file mode 100644 index 0000000..7776a73 --- /dev/null +++ b/adapter/adapter_config.json @@ -0,0 +1,48 @@ +{ + "alora_invocation_tokens": null, + "alpha_pattern": {}, + "arrow_config": null, + "auto_mapping": null, + "base_model_name_or_path": "Qwen/Qwen3-14B", + "bias": "none", + "corda_config": null, + "ensure_weight_tying": false, + "eva_config": null, + "exclude_modules": null, + "fan_in_fan_out": false, + "inference_mode": true, + "init_lora_weights": true, + "layer_replication": null, + "layers_pattern": null, + "layers_to_transform": null, + "loftq_config": {}, + "lora_alpha": 64, + "lora_bias": false, + "lora_dropout": 0.0, + "lora_ga_config": null, + "megatron_config": null, + "megatron_core": "megatron.core", + "modules_to_save": null, + "peft_type": "LORA", + "peft_version": "0.19.1", + "qalora_group_size": 16, + "r": 32, + "rank_pattern": {}, + "revision": null, + "target_modules": [ + "gate_proj", + "v_proj", + "k_proj", + "up_proj", + "down_proj", + "o_proj", + "q_proj" + ], + "target_parameters": null, + "task_type": "CAUSAL_LM", + "trainable_token_indices": null, + "use_bdlora": null, + "use_dora": false, + "use_qalora": false, + "use_rslora": false +} \ No newline at end of file diff --git a/adapter/adapter_model.safetensors b/adapter/adapter_model.safetensors new file mode 100644 index 0000000..fbf80ce --- /dev/null +++ b/adapter/adapter_model.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:68ee829cf0d836f834e05f7fd8048ed8728a005134279148191bcc08ac9559bb +size 513877864 diff --git a/adapter/chat_template.jinja b/adapter/chat_template.jinja new file mode 100644 index 0000000..01be9b3 --- /dev/null +++ b/adapter/chat_template.jinja @@ -0,0 +1,89 @@ +{%- if tools %} + {{- '<|im_start|>system\n' }} + {%- if messages[0].role == 'system' %} + {{- messages[0].content + '\n\n' }} + {%- endif %} + {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within XML tags:\n" }} + {%- for tool in tools %} + {{- "\n" }} + {{- tool | tojson }} + {%- endfor %} + {{- "\n\n\nFor each function call, return a json object with function name and arguments within XML tags:\n\n{\"name\": , \"arguments\": }\n<|im_end|>\n" }} +{%- else %} + {%- if messages[0].role == 'system' %} + {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }} + {%- endif %} +{%- endif %} +{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %} +{%- for message in messages[::-1] %} + {%- set index = (messages|length - 1) - loop.index0 %} + {%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('') and message.content.endswith('')) %} + {%- set ns.multi_step_tool = false %} + {%- set ns.last_query_index = index %} + {%- endif %} +{%- endfor %} +{%- for message in messages %} + {%- if message.content is string %} + {%- set content = message.content %} + {%- else %} + {%- set content = '' %} + {%- endif %} + {%- if (message.role == "user") or (message.role == "system" and not loop.first) %} + {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }} + {%- elif message.role == "assistant" %} + {%- set reasoning_content = '' %} + {%- if message.reasoning_content is string %} + {%- set reasoning_content = message.reasoning_content %} + {%- else %} + {%- if '' in content %} + {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %} + {%- set content = content.split('')[-1].lstrip('\n') %} + {%- endif %} + {%- endif %} + {%- if loop.index0 > ns.last_query_index %} + {%- if loop.last or (not loop.last and reasoning_content) %} + {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content.strip('\n') + '\n\n\n' + content.lstrip('\n') }} + {%- else %} + {{- '<|im_start|>' + message.role + '\n' + content }} + {%- endif %} + {%- else %} + {{- '<|im_start|>' + message.role + '\n' + content }} + {%- endif %} + {%- if message.tool_calls %} + {%- for tool_call in message.tool_calls %} + {%- if (loop.first and content) or (not loop.first) %} + {{- '\n' }} + {%- endif %} + {%- if tool_call.function %} + {%- set tool_call = tool_call.function %} + {%- endif %} + {{- '\n{"name": "' }} + {{- tool_call.name }} + {{- '", "arguments": ' }} + {%- if tool_call.arguments is string %} + {{- tool_call.arguments }} + {%- else %} + {{- tool_call.arguments | tojson }} + {%- endif %} + {{- '}\n' }} + {%- endfor %} + {%- endif %} + {{- '<|im_end|>\n' }} + {%- elif message.role == "tool" %} + {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %} + {{- '<|im_start|>user' }} + {%- endif %} + {{- '\n\n' }} + {{- content }} + {{- '\n' }} + {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %} + {{- '<|im_end|>\n' }} + {%- endif %} + {%- endif %} +{%- endfor %} +{%- if add_generation_prompt %} + {{- '<|im_start|>assistant\n' }} + {%- if enable_thinking is defined and enable_thinking is false %} + {{- '\n\n\n\n' }} + {%- endif %} +{%- endif %} \ No newline at end of file diff --git a/adapter/tokenizer.json b/adapter/tokenizer.json new file mode 100644 index 0000000..c7afbed --- /dev/null +++ b/adapter/tokenizer.json @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 +size 11422650 diff --git a/adapter/tokenizer_config.json b/adapter/tokenizer_config.json new file mode 100644 index 0000000..7d75d3b --- /dev/null +++ b/adapter/tokenizer_config.json @@ -0,0 +1,29 @@ +{ + "add_prefix_space": false, + "backend": "tokenizers", + "bos_token": null, + "clean_up_tokenization_spaces": false, + "eos_token": "<|im_end|>", + "errors": "replace", + "extra_special_tokens": [ + "<|im_start|>", + "<|im_end|>", + "<|object_ref_start|>", + "<|object_ref_end|>", + "<|box_start|>", + "<|box_end|>", + "<|quad_start|>", + "<|quad_end|>", + "<|vision_start|>", + "<|vision_end|>", + "<|vision_pad|>", + "<|image_pad|>", + "<|video_pad|>" + ], + "is_local": false, + "model_max_length": 131072, + "pad_token": "<|endoftext|>", + "split_special_tokens": false, + "tokenizer_class": "Qwen2Tokenizer", + "unk_token": null +} diff --git a/adapter/training_args.bin b/adapter/training_args.bin new file mode 100644 index 0000000..30c835c --- /dev/null +++ b/adapter/training_args.bin @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f7b8e0e2fb5df1b2a243361e988314c091bebb39a1d6987c307d909b5412bd56 +size 5201 diff --git a/adele_judge_config.json b/adele_judge_config.json new file mode 100644 index 0000000..0327618 --- /dev/null +++ b/adele_judge_config.json @@ -0,0 +1,28 @@ +{ + "inference": { + "allowed_scores": [ + "1", + "2", + "3", + "4", + "5" + ], + "binary_threshold": 3, + "method": "restricted_continuation_logprobs_fast" + }, + "model": { + "base_model": "Qwen/Qwen3-14B", + "revision": null, + "thinking_mode": { + "apply_if_supported": true, + "enabled": false + } + }, + "prompt": { + "system_prompt": "You are an answer quality evaluator. You will receive a question, a reference\nanswer, and a model response, each delimited by ### headers.\n\nRate the model response on the following scale:\n1 - Surely incorrect: the response is wrong or completely misses the point.\n2 - Likely incorrect: the response has significant errors or critical omissions.\n3 - Minimally correct / sufficient: the response is partially correct or borderline.\n4 - Likely correct: the response is mostly correct with minor issues.\n5 - Surely correct: the response is fully correct and complete.\n\nRules:\n- Base your rating strictly on correctness relative to the reference answer.\n- Ignore writing style, verbosity, or formatting unless they affect correctness.\n- If the final answer is correct but the reasoning contains substantive errors, lower the rating according to the severity of those errors.\n- Output only a single integer from 1 to 5. No explanation. No punctuation.\n" + }, + "training": { + "max_seq_length": 4096, + "objective": "restricted_score_ce" + } +} \ No newline at end of file diff --git a/adele_judge_metadata.json b/adele_judge_metadata.json new file mode 100644 index 0000000..eb57d64 --- /dev/null +++ b/adele_judge_metadata.json @@ -0,0 +1,1059 @@ +{ + "adapter_path": "runs/qwen3_14b_restricted_score_ce/adapter", + "allowed_scores": [ + "1", + "2", + "3", + "4", + "5" + ], + "artifacts": { + "config.yaml": { + "data": { + "columns": { + "benchmark": "benchmark", + "example_id": "instance_id", + "judge_1_score": "score_gpt4o", + "judge_2_score": "score_sonnet", + "model_id": "model_id", + "question": "question", + "reference_answer": "ground_truth", + "response": "response", + "source": "source", + "task": "task" + }, + "filters": { + "max_disagreement": 1, + "max_response_tokens": 4096, + "on_sequence_overflow": "skip" + }, + "path": "data/processed/response_scores.parquet", + "prepared_dir": null, + "preprocessing_num_workers": 40, + "token_length_batch_size": 2048, + "tokenizers_parallelism": true + }, + "distributed": { + "backend": "nccl", + "deepspeed": { + "config_overrides": {}, + "gradient_clipping": "auto", + "offload_optimizer_device": "none", + "offload_param_device": "none", + "stage3_gather_16bit_weights_on_model_save": true, + "zero_stage": 2 + }, + "enabled": true, + "find_unused_parameters": false, + "fsdp": { + "activation_checkpointing": true, + "sharding_strategy": "full_shard", + "transformer_layer_cls_to_wrap": null, + "use_orig_params": true + }, + "gradient_checkpointing": false, + "mixed_precision": "bf16", + "strategy": "ddp" + }, + "evaluation": { + "length_buckets": [ + 0, + 256, + 512, + 1024, + 2048, + 3072, + 4096, + 1000000000 + ] + }, + "hub": { + "commit_message": "Upload ADeLe distilled judge", + "create_pr": false, + "local_checkpoint_dir": null, + "max_shard_size": "5GB", + "output_staging_dir": null, + "private": false, + "repo_id": null + }, + "inference": { + "allow_base_model": true, + "allowed_scores": [ + "1", + "2", + "3", + "4", + "5" + ], + "batch_size": 64, + "binary_threshold": 3, + "generation_fallback": false, + "method": "restricted_continuation_logprobs_fast", + "require_adapter": false + }, + "model": { + "adapter_path": null, + "attn_implementation": "sdpa", + "model_name_or_path": "Qwen/Qwen3-14B", + "revision": null, + "thinking_mode": { + "apply_if_supported": true, + "enabled": false + }, + "trust_remote_code": true + }, + "project": { + "output_dir": "runs/qwen3_14b_restricted_score_ce", + "run_name": "qwen3_14b_restricted_score_ce", + "seed": 42 + }, + "prompt": { + "system_prompt": "You are an answer quality evaluator. You will receive a question, a reference\nanswer, and a model response, each delimited by ### headers.\n\nRate the model response on the following scale:\n1 - Surely incorrect: the response is wrong or completely misses the point.\n2 - Likely incorrect: the response has significant errors or critical omissions.\n3 - Minimally correct / sufficient: the response is partially correct or borderline.\n4 - Likely correct: the response is mostly correct with minor issues.\n5 - Surely correct: the response is fully correct and complete.\n\nRules:\n- Base your rating strictly on correctness relative to the reference answer.\n- Ignore writing style, verbosity, or formatting unless they affect correctness.\n- If the final answer is correct but the reasoning contains substantive errors, lower the rating according to the severity of those errors.\n- Output only a single integer from 1 to 5. No explanation. No punctuation.\n" + }, + "split": { + "held_out_model": null, + "lomo_validation_fraction": 0.05, + "lomo_validation_max_examples": 30000, + "lomo_validation_seed": 42, + "mode": "fixed_by_model", + "train_models": "auto_except_val_test", + "validation_models": [ + "gemini-3-flash", + "DK-R1-Dist-Qwen-14B", + "llama3d2-3b" + ] + }, + "training": { + "cache_tokenized_datasets": true, + "class_weighting": null, + "dtype": "bfloat16", + "eval_steps": 500, + "eval_subset_seed": 42, + "eval_subset_size": null, + "eval_subset_strategy": "stratified", + "eval_subset_stratify_columns": [ + "model_id", + "target_score" + ], + "gradient_accumulation_steps": 4, + "learning_rate": 3e-05, + "length_column_name": "length", + "load_in_4bit": true, + "logging_steps": 10, + "lora_alpha": 64, + "lora_dropout": 0.0, + "lora_r": 32, + "loss": { + "class_weights": null, + "lambda_binary": 0.5, + "type": "ce_5way" + }, + "lr_scheduler_type": "cosine", + "max_grad_norm": 1.0, + "max_seq_length": 4096, + "num_train_epochs": 1, + "objective": "restricted_score_ce", + "optim": "adamw_8bit", + "packing": false, + "per_device_eval_batch_size": 2, + "per_device_train_batch_size": 2, + "resume_from_checkpoint": null, + "save_steps": 500, + "save_total_limit": 10, + "score_class_weights": null, + "seed": 42, + "target_modules": "auto", + "train_sampling_strategy": "random", + "warmup_ratio": 0.03, + "weight_decay": 0.0 + } + }, + "dataset_filtering_report.json": { + "after_disagreement_filter": 287811, + "after_response_length_filter": 285310, + "effective_prompt_budget_tokens": 0, + "examples_after_sequence_filter": 285158, + "examples_before_sequence_filter": 285310, + "filter_stage_distributions": { + "after_disagreement_filter": { + "benchmark": { + "ChemLLMBench": 28771, + "Civil Service Examination": 7234, + "Data Analysis": 617, + "Date Arithmetic": 9167, + "GRE & GMAT": 3737, + "LSAT": 16530, + "Language": 470, + "MCTACO": 3734, + "MMLU-Pro": 98811, + "Math": 3001, + "MedCalcBench": 9804, + "MenatQA": 11777, + "OmniMath": 29517, + "Reasoning": 975, + "SAT": 7645, + "SciBench": 6360, + "TempReason": 11519, + "TimeDial": 6087, + "TimeQA": 13083, + "TruthQuest": 18972 + }, + "model_id": { + "DK-R1-Dist-Qwen-1.5B": 15314, + "DK-R1-Dist-Qwen-14B": 15352, + "DK-R1-Dist-Qwen-32B": 14922, + "DK-R1-Dist-Qwen-7B": 15215, + "gemini-2.5-flash": 14468, + "gemini-3-flash": 15106, + "gemini-3.1-pro": 15158, + "gpt-35-turbo": 15163, + "gpt-5.2": 14241, + "gpt4o": 16106, + "llama3d1-405b": 15418, + "llama3d2-11b": 15272, + "llama3d2-1b": 15274, + "llama3d2-3b": 15288, + "llama3d2-90b": 15393, + "llama4-17B-128E": 15161, + "o1-mini": 15358, + "o1_re=low": 15044, + "o3-mini": 14558 + }, + "num_examples": 287811, + "target_binary": { + "CORRECT": 192174, + "INCORRECT": 95637 + }, + "target_score": { + "1": 89963, + "2": 5674, + "3": 4685, + "4": 11341, + "5": 176148 + }, + "task": { + "AMPS_Hard": 1174, + "AQuA-RAT": 3737, + "Algebra": 5977, + "Applied Mathematics": 5444, + "Calculus": 535, + "Chemistry": 2539, + "Date Arithmetic": 9167, + "Discrete Mathematics": 5475, + "E": 6185, + "Geometry": 5928, + "I": 6651, + "LSAT-AR": 3290, + "LSAT-LR": 8608, + "LSAT-RC": 4632, + "LogiQA-en": 7234, + "MCTACO": 3734, + "Math": 1901, + "MenatQA-Counterfactual": 2172, + "MenatQA-Order": 2738, + "MenatQA-Scope": 6867, + "Number Theory": 5639, + "Physics": 1920, + "Precalculus": 519, + "S": 6136, + "SAT-En": 3643, + "SAT-Math": 4002, + "TempReason-L2": 5565, + "TempReason-L3": 5954, + "TimeDial": 6087, + "TimeQA-explicit": 6894, + "TimeQA-implicit": 6189, + "biology": 8307, + "business": 7478, + "chemistry": 6672, + "computer science": 6271, + "connections": 470, + "cta": 617, + "date": 491, + "diagnosis": 211, + "dosage": 372, + "economics": 7831, + "engineering": 5291, + "health": 7479, + "history": 5520, + "lab": 3165, + "law": 6462, + "math": 7719, + "math_comp": 1394, + "molecule_captioning": 2641, + "molecule_design": 4313, + "name_prediction": 8652, + "olympiad": 433, + "other": 7799, + "philosophy": 7334, + "physical": 3929, + "physics": 6828, + "psychology": 7820, + "reaction_prediction": 6929, + "retrosynthesis": 6236, + "risk": 1353, + "severity": 283, + "spatial": 582, + "zebra_puzzle": 393 + } + }, + "after_response_length_filter": { + "benchmark": { + "ChemLLMBench": 27913, + "Civil Service Examination": 7207, + "Data Analysis": 617, + "Date Arithmetic": 9167, + "GRE & GMAT": 3733, + "LSAT": 16471, + "Language": 468, + "MCTACO": 3733, + "MMLU-Pro": 98506, + "Math": 2951, + "MedCalcBench": 9799, + "MenatQA": 11762, + "OmniMath": 28488, + "Reasoning": 953, + "SAT": 7645, + "SciBench": 6270, + "TempReason": 11519, + "TimeDial": 6087, + "TimeQA": 13082, + "TruthQuest": 18939 + }, + "model_id": { + "DK-R1-Dist-Qwen-1.5B": 15302, + "DK-R1-Dist-Qwen-14B": 15351, + "DK-R1-Dist-Qwen-32B": 12577, + "DK-R1-Dist-Qwen-7B": 15128, + "gemini-2.5-flash": 14426, + "gemini-3-flash": 15104, + "gemini-3.1-pro": 15158, + "gpt-35-turbo": 15163, + "gpt-5.2": 14241, + "gpt4o": 16106, + "llama3d1-405b": 15418, + "llama3d2-11b": 15267, + "llama3d2-1b": 15272, + "llama3d2-3b": 15288, + "llama3d2-90b": 15393, + "llama4-17B-128E": 15159, + "o1-mini": 15358, + "o1_re=low": 15041, + "o3-mini": 14558 + }, + "num_examples": 285310, + "target_binary": { + "CORRECT": 191333, + "INCORRECT": 93977 + }, + "target_score": { + "1": 88379, + "2": 5598, + "3": 4641, + "4": 11284, + "5": 175408 + }, + "task": { + "AMPS_Hard": 1157, + "AQuA-RAT": 3733, + "Algebra": 5814, + "Applied Mathematics": 5309, + "Calculus": 516, + "Chemistry": 2502, + "Date Arithmetic": 9167, + "Discrete Mathematics": 5235, + "E": 6164, + "Geometry": 5695, + "I": 6647, + "LSAT-AR": 3234, + "LSAT-LR": 8605, + "LSAT-RC": 4632, + "LogiQA-en": 7207, + "MCTACO": 3733, + "Math": 1887, + "MenatQA-Counterfactual": 2166, + "MenatQA-Order": 2737, + "MenatQA-Scope": 6859, + "Number Theory": 5411, + "Physics": 1881, + "Precalculus": 508, + "S": 6128, + "SAT-En": 3643, + "SAT-Math": 4002, + "TempReason-L2": 5565, + "TempReason-L3": 5954, + "TimeDial": 6087, + "TimeQA-explicit": 6893, + "TimeQA-implicit": 6189, + "biology": 8294, + "business": 7469, + "chemistry": 6615, + "computer science": 6245, + "connections": 468, + "cta": 617, + "date": 491, + "diagnosis": 211, + "dosage": 372, + "economics": 7827, + "engineering": 5182, + "health": 7475, + "history": 5520, + "lab": 3163, + "law": 6454, + "math": 7692, + "math_comp": 1367, + "molecule_captioning": 2614, + "molecule_design": 4306, + "name_prediction": 8201, + "olympiad": 427, + "other": 7797, + "philosophy": 7330, + "physical": 3927, + "physics": 6787, + "psychology": 7819, + "reaction_prediction": 6768, + "retrosynthesis": 6024, + "risk": 1353, + "severity": 282, + "spatial": 572, + "zebra_puzzle": 381 + } + }, + "raw": { + "benchmark": { + "ChemLLMBench": 32694, + "Civil Service Examination": 7677, + "Data Analysis": 627, + "Date Arithmetic": 9365, + "GRE & GMAT": 3855, + "LSAT": 17145, + "Language": 550, + "MCTACO": 3891, + "MMLU-Pro": 102896, + "Math": 3263, + "MedCalcBench": 10453, + "MenatQA": 12674, + "OmniMath": 31272, + "Reasoning": 1056, + "SAT": 7752, + "SciBench": 6728, + "TempReason": 12318, + "TimeDial": 6412, + "TimeQA": 13681, + "TruthQuest": 20037 + }, + "model_id": { + "DK-R1-Dist-Qwen-1.5B": 16108, + "DK-R1-Dist-Qwen-14B": 16108, + "DK-R1-Dist-Qwen-32B": 16069, + "DK-R1-Dist-Qwen-7B": 16108, + "gemini-2.5-flash": 15785, + "gemini-3-flash": 16045, + "gemini-3.1-pro": 15981, + "gpt-35-turbo": 16100, + "gpt-5.2": 15050, + "gpt4o": 16106, + "llama3d1-405b": 16108, + "llama3d2-11b": 16108, + "llama3d2-1b": 16107, + "llama3d2-3b": 16108, + "llama3d2-90b": 16108, + "llama4-17B-128E": 16041, + "o1-mini": 16108, + "o1_re=low": 16108, + "o3-mini": 16090 + }, + "num_examples": 304346, + "target_binary": { + "CORRECT": 201528, + "INCORRECT": 102818 + }, + "target_score": { + "1": 89963, + "2": 12855, + "3": 13029, + "4": 12351, + "5": 176148 + }, + "task": { + "AMPS_Hard": 1292, + "AQuA-RAT": 3855, + "Algebra": 6331, + "Applied Mathematics": 5697, + "Calculus": 565, + "Chemistry": 2694, + "Date Arithmetic": 9365, + "Discrete Mathematics": 5861, + "E": 6532, + "Geometry": 6215, + "I": 7045, + "LSAT-AR": 3535, + "LSAT-LR": 8868, + "LSAT-RC": 4742, + "LogiQA-en": 7677, + "MCTACO": 3891, + "Math": 1988, + "MenatQA-Counterfactual": 2440, + "MenatQA-Order": 2925, + "MenatQA-Scope": 7309, + "Number Theory": 6035, + "Physics": 2046, + "Precalculus": 568, + "S": 6460, + "SAT-En": 3686, + "SAT-Math": 4066, + "TempReason-L2": 5960, + "TempReason-L3": 6358, + "TimeDial": 6412, + "TimeQA-explicit": 7137, + "TimeQA-implicit": 6544, + "biology": 8483, + "business": 7783, + "chemistry": 6981, + "computer science": 6542, + "connections": 550, + "cta": 627, + "date": 513, + "diagnosis": 266, + "dosage": 380, + "economics": 8104, + "engineering": 5614, + "health": 7793, + "history": 5729, + "lab": 3357, + "law": 6797, + "math": 8070, + "math_comp": 1480, + "molecule_captioning": 3040, + "molecule_design": 5585, + "name_prediction": 9036, + "olympiad": 491, + "other": 8129, + "philosophy": 7617, + "physical": 4029, + "physics": 7160, + "psychology": 8094, + "reaction_prediction": 7819, + "retrosynthesis": 7214, + "risk": 1587, + "severity": 321, + "spatial": 638, + "zebra_puzzle": 418 + } + } + }, + "kept_prompt_token_length": { + "max": 4093, + "mean": 853.0326345394483, + "min": 252, + "p50": 769.0, + "p75": 988.0, + "p90": 1330.0, + "p95": 1659.0, + "p99": 2596.0 + }, + "kept_sequence_length": { + "max": 4095, + "mean": 855.0326345394483, + "min": 254, + "p50": 771.0, + "p75": 990.0, + "p90": 1332.0, + "p95": 1661.0, + "p99": 2598.0 + }, + "length_filter_warnings": [ + "max_response_tokens is greater than or equal to max_seq_length; examples can pass the response cap while having no room for the formatted prompt.", + "Effective prompt budget is only 0 tokens after reserving the response cap; consider increasing training.max_seq_length." + ], + "max_disagreement": 1, + "max_response_tokens": 4096, + "max_seq_length": 4096, + "on_sequence_overflow": "skip", + "overflowed_response_token_length": { + "max": 4096, + "mean": 3629.6052631578946, + "min": 1582, + "p50": 3875.5, + "p75": 3996.5, + "p90": 4051.0, + "p95": 4075.6, + "p99": 4091.45 + }, + "raw_examples": 304346, + "removed_by_disagreement": 16535, + "removed_by_disagreement_pct": 5.432961169195587, + "removed_by_response_length": 2501, + "removed_by_response_length_pct": 0.8689730413361546, + "sequence_overflow_count": 152, + "sequence_overflow_pct": 0.05327538466930706, + "sequence_overflow_reason": "Full chat-formatted sequence exceeded max_seq_length after response filtering. This includes system prompt, question, reference answer, model response, and chat-template tokens, plus target score tokens." + }, + "inference_config.yaml": { + "data": { + "columns": { + "benchmark": "benchmark", + "example_id": "instance_id", + "judge_1_score": "score_gpt4o", + "judge_2_score": "score_sonnet", + "model_id": "model_id", + "question": "question", + "reference_answer": "ground_truth", + "response": "response", + "source": "source", + "task": "task" + }, + "filters": { + "max_disagreement": 1, + "max_response_tokens": 4096, + "on_sequence_overflow": "skip" + }, + "path": "data/processed/response_scores.parquet", + "prepared_dir": null, + "preprocessing_num_workers": 5, + "token_length_batch_size": 2048, + "tokenizers_parallelism": true + }, + "distributed": { + "backend": "nccl", + "deepspeed": { + "config_overrides": {}, + "gradient_clipping": "auto", + "offload_optimizer_device": "none", + "offload_param_device": "none", + "stage3_gather_16bit_weights_on_model_save": true, + "zero_stage": 2 + }, + "enabled": true, + "find_unused_parameters": false, + "fsdp": { + "activation_checkpointing": true, + "sharding_strategy": "full_shard", + "transformer_layer_cls_to_wrap": null, + "use_orig_params": true + }, + "gradient_checkpointing": true, + "mixed_precision": "bf16", + "strategy": "ddp" + }, + "evaluation": { + "length_buckets": [ + 0, + 256, + 512, + 1024, + 2048, + 3072, + 4096, + 1000000000 + ] + }, + "hub": { + "commit_message": "Upload ADeLe distilled judge", + "create_pr": false, + "local_checkpoint_dir": null, + "max_shard_size": "5GB", + "output_staging_dir": null, + "private": false, + "repo_id": null + }, + "inference": { + "allow_base_model": true, + "allowed_scores": [ + "1", + "2", + "3", + "4", + "5" + ], + "batch_size": 64, + "binary_threshold": 3, + "generation_fallback": false, + "method": "restricted_continuation_logprobs_fast", + "require_adapter": false + }, + "model": { + "adapter_path": "runs/qwen3_14b_restricted_score_ce/adapter", + "attn_implementation": "sdpa", + "model_name_or_path": "Qwen/Qwen3-14B", + "revision": null, + "thinking_mode": { + "apply_if_supported": true, + "enabled": false + }, + "trust_remote_code": true + }, + "project": { + "output_dir": "runs/qwen3_14b_restricted_score_ce", + "run_name": "qwen3_14b_restricted_score_ce", + "seed": 42 + }, + "prompt": { + "system_prompt": "You are an answer quality evaluator. You will receive a question, a reference\nanswer, and a model response, each delimited by ### headers.\n\nRate the model response on the following scale:\n1 - Surely incorrect: the response is wrong or completely misses the point.\n2 - Likely incorrect: the response has significant errors or critical omissions.\n3 - Minimally correct / sufficient: the response is partially correct or borderline.\n4 - Likely correct: the response is mostly correct with minor issues.\n5 - Surely correct: the response is fully correct and complete.\n\nRules:\n- Base your rating strictly on correctness relative to the reference answer.\n- Ignore writing style, verbosity, or formatting unless they affect correctness.\n- If the final answer is correct but the reasoning contains substantive errors, lower the rating according to the severity of those errors.\n- Output only a single integer from 1 to 5. No explanation. No punctuation.\n" + }, + "split": { + "held_out_model": null, + "lomo_validation_fraction": 0.05, + "lomo_validation_max_examples": 30000, + "lomo_validation_seed": 42, + "mode": "fixed_by_model", + "train_models": "auto_except_val_test", + "validation_models": [ + "gemini-3-flash", + "DK-R1-Dist-Qwen-14B", + "llama3d2-3b" + ] + }, + "training": { + "cache_tokenized_datasets": true, + "class_weighting": null, + "dtype": "bfloat16", + "eval_steps": 500, + "eval_subset_seed": 42, + "eval_subset_size": null, + "eval_subset_strategy": "stratified", + "eval_subset_stratify_columns": [ + "model_id", + "target_score" + ], + "gradient_accumulation_steps": 4, + "learning_rate": 3e-05, + "length_column_name": "length", + "load_in_4bit": true, + "logging_steps": 10, + "lora_alpha": 64, + "lora_dropout": 0.0, + "lora_r": 32, + "loss": { + "class_weights": null, + "lambda_binary": 0.5, + "type": "ce_5way" + }, + "lr_scheduler_type": "cosine", + "max_grad_norm": 1.0, + "max_seq_length": 4096, + "num_train_epochs": 1, + "objective": "restricted_score_ce", + "optim": "adamw_8bit", + "packing": false, + "per_device_eval_batch_size": 2, + "per_device_train_batch_size": 2, + "resume_from_checkpoint": null, + "save_steps": 500, + "save_total_limit": 10, + "score_class_weights": null, + "seed": 42, + "target_modules": "auto", + "train_sampling_strategy": "random", + "warmup_ratio": 0.03, + "weight_decay": 0.0 + } + }, + "length_statistics.json": { + "num_examples": 285158, + "prompt_token_length": { + "max": 4093, + "mean": 853.0326345394483, + "min": 252, + "p50": 769.0, + "p75": 988.0, + "p90": 1330.0, + "p95": 1659.0, + "p99": 2596.0 + }, + "response_token_length": { + "max": 3803, + "mean": 397.55135047938336, + "min": 0, + "p50": 307.0, + "p75": 508.0, + "p90": 789.0, + "p95": 1116.0, + "p99": 2162.0 + }, + "sequence_length": { + "max": 4095, + "mean": 855.0326345394483, + "min": 254, + "p50": 771.0, + "p75": 990.0, + "p90": 1332.0, + "p95": 1661.0, + "p99": 2598.0 + }, + "target_token_length": { + "max": 1, + "mean": 1.0, + "min": 1, + "p50": 1.0, + "p75": 1.0, + "p90": 1.0, + "p95": 1.0, + "p99": 1.0 + } + }, + "run_metadata.json": { + "distributed": { + "backend": "nccl", + "deepspeed": { + "config_overrides": {}, + "gradient_clipping": "auto", + "offload_optimizer_device": "none", + "offload_param_device": "none", + "stage3_gather_16bit_weights_on_model_save": true, + "zero_stage": 2 + }, + "effective_global_batch_size": 64, + "enabled": true, + "find_unused_parameters": false, + "fsdp": { + "activation_checkpointing": true, + "sharding_strategy": "full_shard", + "transformer_layer_cls_to_wrap": null, + "use_orig_params": true + }, + "gradient_accumulation_steps": 4, + "gradient_checkpointing": true, + "launcher": "torchrun", + "mixed_precision": "bf16", + "per_device_train_batch_size": 2, + "strategy": "ddp", + "world_size": 8 + }, + "effective_global_batch_size": 64, + "evaluation_enabled": true, + "git_commit": "70ddf4477faac49d4c1d9eb87cbefb58fab7e4a0", + "package_versions": { + "accelerate": "1.13.0", + "bitsandbytes": "0.49.2", + "datasets": "4.3.0", + "deepspeed": "unavailable", + "peft": "0.19.1", + "python": "3.11.15", + "torch": "2.7.1+cu118", + "transformers": "5.5.0", + "trl": "0.24.0", + "unsloth": "2026.4.8" + }, + "score_class_weights": null, + "score_token_ids": [ + 16, + 17, + 18, + 19, + 20 + ], + "source_split_counts": { + "test": 0, + "train": 239420, + "validation": 45738 + }, + "training_backend": "transformers_peft", + "training_examples": 239420, + "training_loss": { + "class_weights": null, + "lambda_binary": 0.5, + "type": "ce_5way" + }, + "training_mode": "standard", + "training_objective": "restricted_score_ce", + "validation_full_examples": 45738, + "validation_monitor_examples": 45738, + "warmup": { + "total_optimization_steps": 3741, + "warmup_ratio": 0.03, + "warmup_steps": 113 + }, + "world_size": 8 + }, + "score_tokenization_report.json": [ + { + "num_tokens": 1, + "score": "1", + "token_ids": [ + 16 + ], + "tokens": [ + "1" + ] + }, + { + "num_tokens": 1, + "score": "2", + "token_ids": [ + 17 + ], + "tokens": [ + "2" + ] + }, + { + "num_tokens": 1, + "score": "3", + "token_ids": [ + 18 + ], + "tokens": [ + "3" + ] + }, + { + "num_tokens": 1, + "score": "4", + "token_ids": [ + 19 + ], + "tokens": [ + "4" + ] + }, + { + "num_tokens": 1, + "score": "5", + "token_ids": [ + 20 + ], + "tokens": [ + "5" + ] + } + ], + "split_report.json": { + "test": { + "examples": 0, + "models": [], + "num_models": 0 + }, + "train": { + "examples": 239420, + "models": [ + "DK-R1-Dist-Qwen-1.5B", + "DK-R1-Dist-Qwen-32B", + "DK-R1-Dist-Qwen-7B", + "gemini-2.5-flash", + "gemini-3.1-pro", + "gpt-35-turbo", + "gpt-5.2", + "gpt4o", + "llama3d1-405b", + "llama3d2-11b", + "llama3d2-1b", + "llama3d2-90b", + "llama4-17B-128E", + "o1-mini", + "o1_re=low", + "o3-mini" + ], + "num_models": 16 + }, + "validation": { + "examples": 45738, + "models": [ + "DK-R1-Dist-Qwen-14B", + "gemini-3-flash", + "llama3d2-3b" + ], + "num_models": 3 + } + }, + "train_metrics.json": { + "epoch": 1.0, + "total_flos": 2.185247058507386e+19, + "train_loss": 0.6326745578038823, + "train_runtime": 128852.3134, + "train_samples_per_second": 1.858, + "train_steps_per_second": 0.029 + }, + "validation_trainer_metrics.json": { + "epoch": 1.0, + "eval_binary_accuracy": 0.9893523984433076, + "eval_binary_macro_f1": 0.9880204659010172, + "eval_confidence_mean": 0.9603744832085106, + "eval_confidence_p50": 0.9991409778594971, + "eval_confidence_p90": 0.9999487400054932, + "eval_expected_calibration_error_10bin": 0.004097698586483999, + "eval_f1_correct": 0.9920149535162078, + "eval_f1_incorrect": 0.9840259782858267, + "eval_f1_score_1": 0.9774369762155193, + "eval_f1_score_2": 0.38461538461538464, + "eval_f1_score_3": 0.5938009787928223, + "eval_f1_score_4": 0.730398069963812, + "eval_f1_score_5": 0.989033961060818, + "eval_false_negative_rate_correct": 0.009138552243694727, + "eval_false_positive_rate_correct": 0.013677012098895318, + "eval_loss": 0.09880143404006958, + "eval_num_examples": 45738, + "eval_ordinal_accuracy": 0.963924963924964, + "eval_ordinal_macro_f1": 0.7350570741296714, + "eval_ordinal_mae": 0.04779395688486598, + "eval_precision_correct": 0.9931711481007256, + "eval_precision_incorrect": 0.98173964264677, + "eval_precision_score_1": 0.9729145558932792, + "eval_precision_score_2": 0.41139240506329117, + "eval_precision_score_3": 0.6159052453468697, + "eval_precision_score_4": 0.7802835051546392, + "eval_precision_score_5": 0.9858030795310072, + "eval_pred_binary_counts": { + "CORRECT": 30459, + "INCORRECT": 15279 + }, + "eval_pred_score_counts": { + "1": 14805, + "2": 474, + "3": 591, + "4": 1552, + "5": 28316 + }, + "eval_recall_correct": 0.9908614477563052, + "eval_recall_incorrect": 0.9863229879011047, + "eval_recall_score_1": 0.9820016362148896, + "eval_recall_score_2": 0.3611111111111111, + "eval_recall_score_3": 0.573228346456693, + "eval_recall_score_4": 0.6865079365079365, + "eval_recall_score_5": 0.9922860900785611, + "eval_runtime": 4341.2041, + "eval_samples_per_second": 10.536, + "eval_score_entropy_mean": 0.10761874169111252, + "eval_score_margin_mean": 6.721137046813965, + "eval_steps_per_second": 0.659, + "eval_support_correct": 30530, + "eval_support_incorrect": 15208, + "eval_support_score_1": 14668, + "eval_support_score_2": 540, + "eval_support_score_3": 635, + "eval_support_score_4": 1764, + "eval_support_score_5": 28131, + "eval_target_score_counts": { + "1": 14668, + "2": 540, + "3": 635, + "4": 1764, + "5": 28131 + }, + "eval_within_1_accuracy": 0.993244129607766 + } + }, + "base_model": "Qwen/Qwen3-14B", + "binary_threshold": 3, + "git_commit": "8a6e08f6f071ce154a27e2718a55ad33895348c6", + "max_seq_length": 4096, + "package_versions": { + "accelerate": "1.13.0", + "bitsandbytes": "0.49.2", + "datasets": "4.3.0", + "deepspeed": "unavailable", + "peft": "0.19.1", + "python": "3.11.15", + "torch": "2.7.1+cu118", + "transformers": "5.5.0", + "trl": "0.24.0", + "unsloth": "2026.4.8" + }, + "repo_id": "adgomant/adele-judge-qwen3-14-cre", + "run_dir": "runs/qwen3_14b_restricted_score_ce", + "thinking_mode": { + "apply_if_supported": true, + "enabled": false + }, + "training_objective": "restricted_score_ce" +} \ No newline at end of file diff --git a/adele_judge_pipeline.py b/adele_judge_pipeline.py new file mode 100644 index 0000000..f3f047c --- /dev/null +++ b/adele_judge_pipeline.py @@ -0,0 +1,304 @@ +from __future__ import annotations + +import inspect +import json +from pathlib import Path +from typing import Any + +from transformers import Pipeline + + +THINKING_KWARG = "enable_thinking" +DEFAULT_SYSTEM_PROMPT = "Return only one score from 1 to 5. Do not explain." +DEFAULT_ALLOWED_SCORES = ["1", "2", "3", "4", "5"] +DEFAULT_BINARY_THRESHOLD = 3 + + +def load_adele_judge_config(repo_id_or_path: str) -> dict[str, Any]: + path = Path(repo_id_or_path) / "adele_judge_config.json" + if path.exists(): + return json.loads(path.read_text(encoding="utf-8")) + + from huggingface_hub import hf_hub_download + + downloaded = hf_hub_download(repo_id_or_path, "adele_judge_config.json") + return json.loads(Path(downloaded).read_text(encoding="utf-8")) + + +def load_adele_judge_config_or_default(model: Any, tokenizer: Any) -> dict[str, Any]: + candidates = [ + getattr(model, "name_or_path", None), + getattr(getattr(model, "config", None), "_name_or_path", None), + getattr(tokenizer, "name_or_path", None), + getattr(tokenizer, "_name_or_path", None), + ] + for candidate in candidates: + if not candidate: + continue + try: + return load_adele_judge_config(str(candidate)) + except Exception: + continue + return {} + + +def adele_judge_settings(config: dict[str, Any] | None) -> dict[str, Any]: + config = config or {} + prompt_config = config.get("prompt", {}) if isinstance(config.get("prompt"), dict) else {} + inference_config = ( + config.get("inference", {}) if isinstance(config.get("inference"), dict) else {} + ) + model_config = config.get("model", {}) if isinstance(config.get("model"), dict) else {} + return { + "system_prompt": prompt_config.get("system_prompt") or DEFAULT_SYSTEM_PROMPT, + "allowed_scores": [ + str(score) + for score in inference_config.get("allowed_scores", DEFAULT_ALLOWED_SCORES) + ], + "binary_threshold": int( + inference_config.get("binary_threshold", DEFAULT_BINARY_THRESHOLD) + ), + "thinking_mode": model_config.get("thinking_mode") or {}, + } + + +def clean_value(value: Any, fallback: str = "N/A") -> str: + if value is None: + return fallback + text = str(value) + if not text or text.lower() == "nan": + return fallback + return text + + +def validate_example(inputs: Any) -> dict[str, Any]: + if not isinstance(inputs, dict): + raise ValueError("ADeLe judge input must be a mapping") + + missing = [] + if inputs.get("question") is None: + missing.append("question") + if inputs.get("model_response") is None: + missing.append("model_response") + reference_answer = inputs.get("reference_answer") + if reference_answer is None: + reference_answer = inputs.get("ground_truth") + if reference_answer is None: + missing.append("reference_answer or ground_truth") + if missing: + raise ValueError(f"Missing required field(s): {', '.join(missing)}") + + return { + "question": inputs["question"], + "reference_answer": reference_answer, + "model_response": inputs["model_response"], + } + + +def build_user_message(example: dict[str, Any]) -> str: + return "\n\n".join( + [ + f"### QUESTION\n{clean_value(example.get('question'))}", + f"### REFERENCE ANSWER\n{clean_value(example.get('reference_answer'))}", + f"### MODEL RESPONSE\n{clean_value(example.get('model_response'), fallback='')}", + "### SCORE\n", + ] + ) + + +def build_messages(example: dict[str, Any], system_prompt: str) -> list[dict[str, str]]: + return [ + {"role": "system", "content": system_prompt.strip()}, + {"role": "user", "content": build_user_message(example)}, + ] + + +def chat_template_supports_thinking(tokenizer: Any) -> bool: + apply_chat_template = getattr(tokenizer, "apply_chat_template", None) + if apply_chat_template is None: + return False + try: + signature = inspect.signature(apply_chat_template) + except (TypeError, ValueError): + return False + accepts_kwarg = any( + parameter.kind == inspect.Parameter.VAR_KEYWORD or name == THINKING_KWARG + for name, parameter in signature.parameters.items() + ) + if not accepts_kwarg: + return False + + template = getattr(tokenizer, "chat_template", None) + if isinstance(template, str) and THINKING_KWARG in template: + return True + + candidates = [ + getattr(tokenizer, "name_or_path", None), + getattr(tokenizer, "_name_or_path", None), + getattr(tokenizer, "model_name", None), + ] + init_kwargs = getattr(tokenizer, "init_kwargs", None) + if isinstance(init_kwargs, dict): + candidates.extend([init_kwargs.get("name_or_path"), init_kwargs.get("tokenizer_file")]) + return any("qwen3" in str(candidate).lower() for candidate in candidates if candidate) + + +def apply_chat_template_safe( + tokenizer: Any, + messages: list[dict[str, str]], + *, + add_generation_prompt: bool, + thinking_mode: dict[str, Any], +) -> str: + if hasattr(tokenizer, "apply_chat_template") and getattr(tokenizer, "chat_template", None): + template_kwargs = {} + enabled = thinking_mode.get("enabled") + if ( + enabled is not None + and bool(thinking_mode.get("apply_if_supported", True)) + and chat_template_supports_thinking(tokenizer) + ): + template_kwargs[THINKING_KWARG] = bool(enabled) + return tokenizer.apply_chat_template( + messages, + tokenize=False, + add_generation_prompt=add_generation_prompt, + **template_kwargs, + ) + + rendered = [f"<|{message['role']}|>\n{message['content']}" for message in messages] + if add_generation_prompt: + rendered.append("<|assistant|>\n") + return "\n".join(rendered) + + +def encode_text(tokenizer: Any, text: str) -> list[int]: + return tokenizer(text, add_special_tokens=False, truncation=False)["input_ids"] + + +def single_score_token_ids(tokenizer: Any, allowed_scores: list[str]) -> list[int]: + token_ids = [encode_text(tokenizer, score) for score in allowed_scores] + multi_token_scores = [ + score for score, ids in zip(allowed_scores, token_ids, strict=True) if len(ids) != 1 + ] + if multi_token_scores: + raise ValueError( + "ADeLeJudgePipeline requires score continuations to be single tokens; " + f"multi-token scores: {multi_token_scores}" + ) + return [ids[0] for ids in token_ids] + + +class ADeLeJudgePipeline(Pipeline): + """HF-native custom pipeline for restricted ADeLe judge scoring.""" + + def __init__( + self, + *args: Any, + adele_config: dict[str, Any] | None = None, + **kwargs: Any, + ) -> None: + super().__init__(*args, **kwargs) + if self.tokenizer is None: + raise ValueError("ADeLeJudgePipeline requires a tokenizer") + if getattr(self.tokenizer, "pad_token", None) is None: + self.tokenizer.pad_token = getattr(self.tokenizer, "eos_token", None) + + settings = adele_judge_settings( + adele_config + if adele_config is not None + else load_adele_judge_config_or_default(self.model, self.tokenizer) + ) + self.system_prompt = settings["system_prompt"] + self.allowed_scores = settings["allowed_scores"] + self.binary_threshold = settings["binary_threshold"] + self.thinking_mode = settings["thinking_mode"] + self.score_token_ids = single_score_token_ids(self.tokenizer, self.allowed_scores) + + if hasattr(self.model, "eval"): + self.model.eval() + + def _sanitize_parameters( + self, + **kwargs: Any, + ) -> tuple[dict[str, Any], dict[str, Any], dict[str, Any]]: + return {}, {}, {} + + def preprocess(self, inputs: Any) -> dict[str, Any]: + import torch + + example = validate_example(inputs) + prompt = apply_chat_template_safe( + self.tokenizer, + build_messages(example, self.system_prompt), + add_generation_prompt=True, + thinking_mode=self.thinking_mode, + ) + encoded = self.tokenizer( + prompt, + add_special_tokens=False, + truncation=False, + return_tensors="pt", + ) + if "attention_mask" not in encoded: + encoded["attention_mask"] = torch.ones_like(encoded["input_ids"]) + return {"input_ids": encoded["input_ids"], "attention_mask": encoded["attention_mask"]} + + def _forward(self, model_inputs: dict[str, Any]) -> dict[str, Any]: + import torch + import torch.nn.functional as F + + input_ids = model_inputs["input_ids"] + attention_mask = model_inputs["attention_mask"] + with torch.no_grad(): + outputs = self.model(input_ids=input_ids, attention_mask=attention_mask) + + token_positions = torch.arange(input_ids.shape[1], device=input_ids.device).unsqueeze(0) + positions = (attention_mask * token_positions).max(dim=1).values.to(dtype=torch.long) + batch_indices = torch.arange(input_ids.shape[0], device=input_ids.device) + final_logits = outputs.logits[batch_indices, positions] + + score_ids = torch.tensor(self.score_token_ids, dtype=torch.long, device=final_logits.device) + score_logits = final_logits[:, score_ids] + logprobs = F.log_softmax(score_logits, dim=-1) + return { + "score_indices": torch.argmax(logprobs, dim=-1), + "probs": torch.exp(logprobs), + "logprobs": logprobs, + } + + def postprocess(self, model_outputs: dict[str, Any]) -> dict[str, Any]: + import torch + + score_index = int(model_outputs["score_indices"].reshape(-1)[0]) + probs_tensor = model_outputs["probs"].reshape(-1, len(self.allowed_scores))[0] + logprobs_tensor = model_outputs["logprobs"].reshape(-1, len(self.allowed_scores))[0] + + probs = { + score: float(prob) + for score, prob in zip(self.allowed_scores, probs_tensor.tolist(), strict=True) + } + logprobs = { + score: float(logprob) + for score, logprob in zip(self.allowed_scores, logprobs_tensor.tolist(), strict=True) + } + + score = int(self.allowed_scores[score_index]) + sorted_logprobs = torch.sort(logprobs_tensor).values + margin = ( + float(sorted_logprobs[-1] - sorted_logprobs[-2]) + if len(sorted_logprobs) > 1 + else 0.0 + ) + entropy = float( + -(probs_tensor * torch.log(torch.clamp(probs_tensor, min=1e-12))).sum() + ) + return { + "score": score, + "label": "CORRECT" if score >= self.binary_threshold else "INCORRECT", + "probs": probs, + "logprobs": logprobs, + "confidence": max(probs.values()), + "margin": margin, + "entropy": entropy, + } diff --git a/chat_template.jinja b/chat_template.jinja new file mode 100644 index 0000000..01be9b3 --- /dev/null +++ b/chat_template.jinja @@ -0,0 +1,89 @@ +{%- if tools %} + {{- '<|im_start|>system\n' }} + {%- if messages[0].role == 'system' %} + {{- messages[0].content + '\n\n' }} + {%- endif %} + {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within XML tags:\n" }} + {%- for tool in tools %} + {{- "\n" }} + {{- tool | tojson }} + {%- endfor %} + {{- "\n\n\nFor each function call, return a json object with function name and arguments within XML tags:\n\n{\"name\": , \"arguments\": }\n<|im_end|>\n" }} +{%- else %} + {%- if messages[0].role == 'system' %} + {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }} + {%- endif %} +{%- endif %} +{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %} +{%- for message in messages[::-1] %} + {%- set index = (messages|length - 1) - loop.index0 %} + {%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('') and message.content.endswith('')) %} + {%- set ns.multi_step_tool = false %} + {%- set ns.last_query_index = index %} + {%- endif %} +{%- endfor %} +{%- for message in messages %} + {%- if message.content is string %} + {%- set content = message.content %} + {%- else %} + {%- set content = '' %} + {%- endif %} + {%- if (message.role == "user") or (message.role == "system" and not loop.first) %} + {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }} + {%- elif message.role == "assistant" %} + {%- set reasoning_content = '' %} + {%- if message.reasoning_content is string %} + {%- set reasoning_content = message.reasoning_content %} + {%- else %} + {%- if '' in content %} + {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %} + {%- set content = content.split('')[-1].lstrip('\n') %} + {%- endif %} + {%- endif %} + {%- if loop.index0 > ns.last_query_index %} + {%- if loop.last or (not loop.last and reasoning_content) %} + {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content.strip('\n') + '\n\n\n' + content.lstrip('\n') }} + {%- else %} + {{- '<|im_start|>' + message.role + '\n' + content }} + {%- endif %} + {%- else %} + {{- '<|im_start|>' + message.role + '\n' + content }} + {%- endif %} + {%- if message.tool_calls %} + {%- for tool_call in message.tool_calls %} + {%- if (loop.first and content) or (not loop.first) %} + {{- '\n' }} + {%- endif %} + {%- if tool_call.function %} + {%- set tool_call = tool_call.function %} + {%- endif %} + {{- '\n{"name": "' }} + {{- tool_call.name }} + {{- '", "arguments": ' }} + {%- if tool_call.arguments is string %} + {{- tool_call.arguments }} + {%- else %} + {{- tool_call.arguments | tojson }} + {%- endif %} + {{- '}\n' }} + {%- endfor %} + {%- endif %} + {{- '<|im_end|>\n' }} + {%- elif message.role == "tool" %} + {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %} + {{- '<|im_start|>user' }} + {%- endif %} + {{- '\n\n' }} + {{- content }} + {{- '\n' }} + {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %} + {{- '<|im_end|>\n' }} + {%- endif %} + {%- endif %} +{%- endfor %} +{%- if add_generation_prompt %} + {{- '<|im_start|>assistant\n' }} + {%- if enable_thinking is defined and enable_thinking is false %} + {{- '\n\n\n\n' }} + {%- endif %} +{%- endif %} \ No newline at end of file diff --git a/config.json b/config.json new file mode 100644 index 0000000..f413f90 --- /dev/null +++ b/config.json @@ -0,0 +1,85 @@ +{ + "architectures": [ + "Qwen3ForCausalLM" + ], + "attention_bias": false, + "attention_dropout": 0.0, + "bos_token_id": 151643, + "custom_pipelines": { + "adele-judge": { + "impl": "adele_judge_pipeline.ADeLeJudgePipeline", + "pt": [ + "AutoModelForCausalLM" + ], + "tf": [], + "type": "text" + } + }, + "dtype": "bfloat16", + "eos_token_id": 151645, + "head_dim": 128, + "hidden_act": "silu", + "hidden_size": 5120, + "initializer_range": 0.02, + "intermediate_size": 17408, + "layer_types": [ + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention", + "full_attention" + ], + "max_position_embeddings": 40960, + "max_window_layers": 40, + "model_type": "qwen3", + "num_attention_heads": 40, + "num_hidden_layers": 40, + "num_key_value_heads": 8, + "pad_token_id": null, + "rms_norm_eps": 1e-06, + "rope_parameters": { + "rope_theta": 1000000, + "rope_type": "default" + }, + "sliding_window": null, + "tie_word_embeddings": false, + "transformers_version": "5.5.0", + "use_cache": true, + "use_sliding_window": false, + "vocab_size": 151936 +} \ No newline at end of file diff --git a/generation_config.json b/generation_config.json new file mode 100644 index 0000000..efab52f --- /dev/null +++ b/generation_config.json @@ -0,0 +1,5 @@ +{ + "do_sample": false, + "max_new_tokens": 1, + "num_beams": 1 +} \ No newline at end of file diff --git a/model-00001-of-00006.safetensors b/model-00001-of-00006.safetensors new file mode 100644 index 0000000..b0c81ba --- /dev/null +++ b/model-00001-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5739c7705041195574706abc2af5e12c66c2281e1a93dd1419dd2439c0fd1a41 +size 4978180848 diff --git a/model-00002-of-00006.safetensors b/model-00002-of-00006.safetensors new file mode 100644 index 0000000..59c4216 --- /dev/null +++ b/model-00002-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:418628dcaf69131e39dff0aa22339698bf3111fd402a0c677ff8b1fb08a146de +size 4917988408 diff --git a/model-00003-of-00006.safetensors b/model-00003-of-00006.safetensors new file mode 100644 index 0000000..9f0f555 --- /dev/null +++ b/model-00003-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7e63ccd5fd8e91297f66be63117190fbe6c5ad636c669d33263f490c42270e3a +size 4991388720 diff --git a/model-00004-of-00006.safetensors b/model-00004-of-00006.safetensors new file mode 100644 index 0000000..8f5bc9d --- /dev/null +++ b/model-00004-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:483606985498d68d803f7761870f76b0df2965034149f2d4af7f7c93bf707c3b +size 4917988496 diff --git a/model-00005-of-00006.safetensors b/model-00005-of-00006.safetensors new file mode 100644 index 0000000..3d40832 --- /dev/null +++ b/model-00005-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:acd0b2c97fb9d8fda9360f6293ccf899914f43d54ab8bcba3557ac62412451bd +size 4991388720 diff --git a/model-00006-of-00006.safetensors b/model-00006-of-00006.safetensors new file mode 100644 index 0000000..b8830fc --- /dev/null +++ b/model-00006-of-00006.safetensors @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2e2e7fd55cd67b4370f657b6f1d3ee9673f8a8657e4a90eaf873bed6661744ad +size 4739730440 diff --git a/model.safetensors.index.json b/model.safetensors.index.json new file mode 100644 index 0000000..8d53005 --- /dev/null +++ b/model.safetensors.index.json @@ -0,0 +1,451 @@ +{ + "metadata": { + "total_parameters": 14768307200, + "total_size": 29536614400 + }, + "weight_map": { + "lm_head.weight": "model-00001-of-00006.safetensors", + "model.embed_tokens.weight": "model-00001-of-00006.safetensors", + "model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.k_norm.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.q_norm.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.k_norm.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.q_norm.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.10.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.10.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.10.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.10.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.10.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.11.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.11.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.12.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.12.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.13.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.13.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.14.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.14.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.q_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.17.input_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors", + "model.layers.17.self_attn.k_norm.weight": "model-00003-of-00006.safetensors", + "model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors", + "model.layers.17.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.17.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.17.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.17.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.18.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.18.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.19.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.19.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors", + "model.layers.2.self_attn.k_norm.weight": "model-00001-of-00006.safetensors", + "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors", + "model.layers.2.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.2.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.2.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.2.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.20.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.20.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.20.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.21.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.21.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.22.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.22.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.23.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.23.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.k_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.k_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.o_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.q_norm.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.q_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.24.self_attn.v_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.25.input_layernorm.weight": "model-00004-of-00006.safetensors", + "model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors", + "model.layers.25.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.25.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.25.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.25.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.26.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.26.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.27.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.27.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.28.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.28.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.29.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.29.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.3.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.3.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.3.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.30.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.30.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.30.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.31.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.o_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.q_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.q_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.31.self_attn.v_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.32.input_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.32.mlp.down_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.32.mlp.gate_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.32.mlp.up_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.32.post_attention_layernorm.weight": "model-00005-of-00006.safetensors", + "model.layers.32.self_attn.k_norm.weight": "model-00005-of-00006.safetensors", + "model.layers.32.self_attn.k_proj.weight": "model-00005-of-00006.safetensors", + "model.layers.32.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.32.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.32.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.32.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.33.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.33.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.34.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.34.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.35.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.35.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.36.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.36.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.37.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.37.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.38.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.38.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.input_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.39.mlp.down_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.mlp.gate_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.mlp.up_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.post_attention_layernorm.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.k_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.k_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.o_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.q_norm.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.q_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.39.self_attn.v_proj.weight": "model-00006-of-00006.safetensors", + "model.layers.4.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.4.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.4.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.5.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.5.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.6.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.6.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.k_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.q_norm.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors", + "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors", + "model.norm.weight": "model-00006-of-00006.safetensors" + } +} diff --git a/tokenizer.json b/tokenizer.json new file mode 100644 index 0000000..c7afbed --- /dev/null +++ b/tokenizer.json @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 +size 11422650 diff --git a/tokenizer_config.json b/tokenizer_config.json new file mode 100644 index 0000000..9fd0fb4 --- /dev/null +++ b/tokenizer_config.json @@ -0,0 +1,29 @@ +{ + "add_prefix_space": false, + "backend": "tokenizers", + "bos_token": null, + "clean_up_tokenization_spaces": false, + "eos_token": "<|im_end|>", + "errors": "replace", + "extra_special_tokens": [ + "<|im_start|>", + "<|im_end|>", + "<|object_ref_start|>", + "<|object_ref_end|>", + "<|box_start|>", + "<|box_end|>", + "<|quad_start|>", + "<|quad_end|>", + "<|vision_start|>", + "<|vision_end|>", + "<|vision_pad|>", + "<|image_pad|>", + "<|video_pad|>" + ], + "is_local": true, + "model_max_length": 131072, + "pad_token": "<|endoftext|>", + "split_special_tokens": false, + "tokenizer_class": "Qwen2Tokenizer", + "unk_token": null +} diff --git a/training_config.yaml b/training_config.yaml new file mode 100644 index 0000000..7bc69c3 --- /dev/null +++ b/training_config.yaml @@ -0,0 +1,168 @@ +project: + run_name: qwen3_14b_restricted_score_ce + output_dir: runs/qwen3_14b_restricted_score_ce + seed: 42 +data: + path: data/processed/response_scores.parquet + prepared_dir: null + columns: + question: question + reference_answer: ground_truth + response: response + judge_1_score: score_gpt4o + judge_2_score: score_sonnet + model_id: model_id + benchmark: benchmark + task: task + example_id: instance_id + source: source + filters: + max_disagreement: 1 + max_response_tokens: 4096 + on_sequence_overflow: skip + preprocessing_num_workers: 40 + tokenizers_parallelism: true + token_length_batch_size: 2048 +model: + model_name_or_path: Qwen/Qwen3-14B + revision: null + attn_implementation: sdpa + adapter_path: null + trust_remote_code: true + thinking_mode: + enabled: false + apply_if_supported: true +prompt: + system_prompt: 'You are an answer quality evaluator. You will receive a question, + a reference + + answer, and a model response, each delimited by ### headers. + + + Rate the model response on the following scale: + + 1 - Surely incorrect: the response is wrong or completely misses the point. + + 2 - Likely incorrect: the response has significant errors or critical omissions. + + 3 - Minimally correct / sufficient: the response is partially correct or borderline. + + 4 - Likely correct: the response is mostly correct with minor issues. + + 5 - Surely correct: the response is fully correct and complete. + + + Rules: + + - Base your rating strictly on correctness relative to the reference answer. + + - Ignore writing style, verbosity, or formatting unless they affect correctness. + + - If the final answer is correct but the reasoning contains substantive errors, + lower the rating according to the severity of those errors. + + - Output only a single integer from 1 to 5. No explanation. No punctuation. + + ' +split: + mode: fixed_by_model + validation_models: + - gemini-3-flash + - DK-R1-Dist-Qwen-14B + - llama3d2-3b + train_models: auto_except_val_test + held_out_model: null + lomo_validation_fraction: 0.05 + lomo_validation_max_examples: 30000 + lomo_validation_seed: 42 +training: + max_seq_length: 4096 + load_in_4bit: true + dtype: bfloat16 + objective: restricted_score_ce + loss: + type: ce_5way + lambda_binary: 0.5 + class_weights: null + class_weighting: null + score_class_weights: null + lora_r: 32 + lora_alpha: 64 + lora_dropout: 0.0 + target_modules: auto + learning_rate: 3.0e-05 + num_train_epochs: 1 + per_device_train_batch_size: 2 + per_device_eval_batch_size: 2 + gradient_accumulation_steps: 4 + warmup_ratio: 0.03 + lr_scheduler_type: cosine + weight_decay: 0.0 + optim: adamw_8bit + packing: false + cache_tokenized_datasets: true + eval_subset_size: null + eval_subset_strategy: stratified + eval_subset_stratify_columns: + - model_id + - target_score + train_sampling_strategy: random + length_column_name: length + logging_steps: 10 + eval_steps: 500 + save_steps: 500 + save_total_limit: 10 + seed: 42 + resume_from_checkpoint: null + eval_subset_seed: 42 + max_grad_norm: 1.0 +distributed: + enabled: true + strategy: ddp + backend: nccl + mixed_precision: bf16 + gradient_checkpointing: false + find_unused_parameters: false + fsdp: + sharding_strategy: full_shard + transformer_layer_cls_to_wrap: null + activation_checkpointing: true + use_orig_params: true + deepspeed: + zero_stage: 2 + offload_optimizer_device: none + offload_param_device: none + stage3_gather_16bit_weights_on_model_save: true + gradient_clipping: auto + config_overrides: {} +inference: + allowed_scores: + - '1' + - '2' + - '3' + - '4' + - '5' + binary_threshold: 3 + method: restricted_continuation_logprobs_fast + generation_fallback: false + batch_size: 64 + require_adapter: false + allow_base_model: true +evaluation: + length_buckets: + - 0 + - 256 + - 512 + - 1024 + - 2048 + - 3072 + - 4096 + - 1000000000 +hub: + repo_id: null + private: false + commit_message: Upload ADeLe distilled judge + local_checkpoint_dir: null + output_staging_dir: null + create_pr: false + max_shard_size: 5GB