初始化项目,由ModelHub XC社区提供模型
Model: adgomant/adele-judge-qwen3-14B-cre Source: Original Platform
This commit is contained in:
37
.gitattributes
vendored
Normal file
37
.gitattributes
vendored
Normal file
@@ -0,0 +1,37 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
adapter/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
151
README.md
Normal file
151
README.md
Normal file
@@ -0,0 +1,151 @@
|
||||
---
|
||||
library_name: transformers
|
||||
tags:
|
||||
- text-generation
|
||||
- peft
|
||||
- adele
|
||||
- judge
|
||||
base_model: Qwen/Qwen3-14B
|
||||
datasets:
|
||||
- CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0
|
||||
---
|
||||
|
||||
# ADeLe Distilled Judge
|
||||
|
||||
This repository contains an ADeLe-suite-specific distilled judge. It scores a model response against a question and reference answer with an ordinal score from 1 to 5, then derives binary correctness with the ADeLe threshold.
|
||||
|
||||
The repository root contains a merged Transformers model for standard loading. The original LoRA adapter is also included under `adapter/` for provenance and reuse.
|
||||
|
||||
## Intended Use
|
||||
|
||||
Use this model to score ADeLe-style examples where a question, reference answer, and model response are available. It is intended for out-of-model evaluation within the ADeLe benchmark suite, not as a general-purpose evaluator.
|
||||
|
||||
## Input Format
|
||||
|
||||
The recommended helper accepts:
|
||||
|
||||
- `question`
|
||||
- `reference_answer` or `ground_truth`
|
||||
- `model_response`
|
||||
|
||||
## Score Rubric
|
||||
|
||||
Allowed scores: 1, 2, 3, 4, 5
|
||||
|
||||
- 1: surely incorrect
|
||||
- 2: likely incorrect
|
||||
- 3: minimally correct or sufficient
|
||||
- 4: likely correct
|
||||
- 5: surely correct
|
||||
|
||||
Binary label: scores greater than or equal to 3 are `CORRECT`; lower scores are `INCORRECT`.
|
||||
|
||||
## Training And Validation Data
|
||||
|
||||
| Split | Examples | Models |
|
||||
| --- | --- | --- |
|
||||
| train | 239,420 | 16 |
|
||||
| validation | 45,738 | 3 |
|
||||
|
||||
- `train` models: `DK-R1-Dist-Qwen-1.5B`, `DK-R1-Dist-Qwen-32B`, `DK-R1-Dist-Qwen-7B`, `gemini-2.5-flash`, `gemini-3.1-pro`, `gpt-35-turbo`, `gpt-5.2`, `gpt4o`, `llama3d1-405b`, `llama3d2-11b`, `llama3d2-1b`, `llama3d2-90b`, `llama4-17B-128E`, `o1-mini`, `o1_re=low`, `o3-mini`
|
||||
- `validation` models: `DK-R1-Dist-Qwen-14B`, `gemini-3-flash`, `llama3d2-3b`
|
||||
|
||||
## Data Quality And Label Construction
|
||||
|
||||
Training labels are distilled from two proprietary judge scores used by the ADeLe evaluation pipeline to derive the official correctness signal. The configured source columns are `score_gpt4o` and `score_sonnet`.
|
||||
|
||||
- Ordinal target: `floor(mean(score_gpt4o, score_sonnet))`.
|
||||
- Binary target: `CORRECT` when the ordinal target is >= `3`.
|
||||
- Judge-agreement filter: keep examples with `abs(score_gpt4o - score_sonnet) <= 1`.
|
||||
- Response-length filter: keep responses with at most `4096` base-tokenizer tokens before prompt formatting.
|
||||
- Sequence-length filter: keep full chat-formatted examples within `max_seq_length=8192`.
|
||||
|
||||
## Validation Results
|
||||
|
||||
Source artifact: `validation_trainer_metrics.json`.
|
||||
|
||||
| Metric | Value |
|
||||
| --- | --- |
|
||||
| Epoch | 1.0000 |
|
||||
| Binary accuracy | 0.9894 |
|
||||
| Binary macro F1 | 0.9880 |
|
||||
| Precision, CORRECT | 0.9932 |
|
||||
| Recall, CORRECT | 0.9909 |
|
||||
| Precision, INCORRECT | 0.9817 |
|
||||
| Recall, INCORRECT | 0.9863 |
|
||||
| False negative rate, CORRECT | 0.0091 |
|
||||
| False positive rate, CORRECT | 0.0137 |
|
||||
| Ordinal accuracy | 0.9639 |
|
||||
| Ordinal macro F1 | 0.7351 |
|
||||
| Mean confidence | 0.9604 |
|
||||
|
||||
## Recommended Inference
|
||||
|
||||
Do not use free-form generation as the primary prediction method. The recommended path scores the restricted continuations `"1"`, `"2"`, `"3"`, `"4"`, and `"5"`.
|
||||
|
||||
```python
|
||||
from transformers import pipeline
|
||||
|
||||
judge = pipeline(
|
||||
"adele-judge",
|
||||
model="adgomant/adele-judge-qwen3-14-cre",
|
||||
trust_remote_code=True,
|
||||
device_map="auto",
|
||||
)
|
||||
result = judge(
|
||||
{"question": "...", "reference_answer": "...", "model_response": "..."}
|
||||
)
|
||||
print(result)
|
||||
|
||||
results = judge([
|
||||
{"question": "...", "reference_answer": "...", "model_response": "..."},
|
||||
{"question": "...", "ground_truth": "...", "model_response": "..."},
|
||||
], batch_size=8)
|
||||
```
|
||||
|
||||
The result has this shape:
|
||||
|
||||
```python
|
||||
{
|
||||
"score": 4,
|
||||
"label": "CORRECT",
|
||||
"probs": {"1": 0.01, "2": 0.02, "3": 0.08, "4": 0.70, "5": 0.19},
|
||||
"logprobs": {"1": -5.0, "2": -4.2, "3": -2.9, "4": -0.8, "5": -2.1},
|
||||
"confidence": 0.70,
|
||||
"margin": 1.3,
|
||||
"entropy": 0.82,
|
||||
}
|
||||
```
|
||||
|
||||
## Standard Transformers Loading
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained("adgomant/adele-judge-qwen3-14-cre", trust_remote_code=True)
|
||||
model = AutoModelForCausalLM.from_pretrained("adgomant/adele-judge-qwen3-14-cre", trust_remote_code=True)
|
||||
```
|
||||
|
||||
`generation_config.json` uses safe one-token defaults for debugging, but `generate()` is not the recommended scoring method.
|
||||
|
||||
## Metadata
|
||||
|
||||
Training, filtering, split, tokenization, and metric artifacts available at packaging time are stored in `adele_judge_metadata.json`.
|
||||
|
||||
The model is trained on distilled judge targets. These targets are useful for reproducing the ADeLe paper-style correctness signal at lower inference cost, but they should not be interpreted as independent human annotations.
|
||||
|
||||
## References
|
||||
|
||||
- ADeLe project page: [ADeLe v1.0](https://kinds-of-intelligence-cfi.github.io/ADELE/).
|
||||
- ADeLe paper and official correctness definition: [General scales unlock AI evaluation with explanatory and predictive power](https://www.nature.com/articles/s41586-026-10303-2).
|
||||
- Official ADeLe dataset: [CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0](https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0).
|
||||
- Official instance-level model-response data used for distillation: [https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation/tree/main/ADeLe_battery_data/subject_specific_instance_level_data](https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation/tree/main/ADeLe_battery_data/subject_specific_instance_level_data).
|
||||
- Training and Hub packaging implementation: [https://github.com/adgomant/adele-judge](https://github.com/adgomant/adele-judge).
|
||||
|
||||
## Limitations
|
||||
|
||||
- ADeLe-specific judge; not a general-purpose evaluator.
|
||||
- Distilled from proprietary judge labels and inherits their noise, calibration, and biases.
|
||||
- Intended for scoring responses against a reference answer.
|
||||
- It should not produce explanations; the expected output is a single score.
|
||||
- Validation is out-of-model within the ADeLe suite, so transfer outside that suite should be measured before relying on it.
|
||||
207
adapter/README.md
Normal file
207
adapter/README.md
Normal file
@@ -0,0 +1,207 @@
|
||||
---
|
||||
base_model: Qwen/Qwen3-14B
|
||||
library_name: peft
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- base_model:adapter:Qwen/Qwen3-14B
|
||||
- lora
|
||||
- transformers
|
||||
---
|
||||
|
||||
# Model Card for Model ID
|
||||
|
||||
<!-- Provide a quick summary of what the model is/does. -->
|
||||
|
||||
|
||||
|
||||
## Model Details
|
||||
|
||||
### Model Description
|
||||
|
||||
<!-- Provide a longer summary of what this model is. -->
|
||||
|
||||
|
||||
|
||||
- **Developed by:** [More Information Needed]
|
||||
- **Funded by [optional]:** [More Information Needed]
|
||||
- **Shared by [optional]:** [More Information Needed]
|
||||
- **Model type:** [More Information Needed]
|
||||
- **Language(s) (NLP):** [More Information Needed]
|
||||
- **License:** [More Information Needed]
|
||||
- **Finetuned from model [optional]:** [More Information Needed]
|
||||
|
||||
### Model Sources [optional]
|
||||
|
||||
<!-- Provide the basic links for the model. -->
|
||||
|
||||
- **Repository:** [More Information Needed]
|
||||
- **Paper [optional]:** [More Information Needed]
|
||||
- **Demo [optional]:** [More Information Needed]
|
||||
|
||||
## Uses
|
||||
|
||||
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
||||
|
||||
### Direct Use
|
||||
|
||||
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Downstream Use [optional]
|
||||
|
||||
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Out-of-Scope Use
|
||||
|
||||
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Bias, Risks, and Limitations
|
||||
|
||||
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Recommendations
|
||||
|
||||
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
||||
|
||||
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
||||
|
||||
## How to Get Started with the Model
|
||||
|
||||
Use the code below to get started with the model.
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Training Details
|
||||
|
||||
### Training Data
|
||||
|
||||
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Training Procedure
|
||||
|
||||
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
||||
|
||||
#### Preprocessing [optional]
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
|
||||
#### Training Hyperparameters
|
||||
|
||||
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
||||
|
||||
#### Speeds, Sizes, Times [optional]
|
||||
|
||||
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Evaluation
|
||||
|
||||
<!-- This section describes the evaluation protocols and provides the results. -->
|
||||
|
||||
### Testing Data, Factors & Metrics
|
||||
|
||||
#### Testing Data
|
||||
|
||||
<!-- This should link to a Dataset Card if possible. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
#### Factors
|
||||
|
||||
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
#### Metrics
|
||||
|
||||
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Results
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
#### Summary
|
||||
|
||||
|
||||
|
||||
## Model Examination [optional]
|
||||
|
||||
<!-- Relevant interpretability work for the model goes here -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Environmental Impact
|
||||
|
||||
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
||||
|
||||
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
||||
|
||||
- **Hardware Type:** [More Information Needed]
|
||||
- **Hours used:** [More Information Needed]
|
||||
- **Cloud Provider:** [More Information Needed]
|
||||
- **Compute Region:** [More Information Needed]
|
||||
- **Carbon Emitted:** [More Information Needed]
|
||||
|
||||
## Technical Specifications [optional]
|
||||
|
||||
### Model Architecture and Objective
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
### Compute Infrastructure
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
#### Hardware
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
#### Software
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Citation [optional]
|
||||
|
||||
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
||||
|
||||
**BibTeX:**
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
**APA:**
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Glossary [optional]
|
||||
|
||||
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## More Information [optional]
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Model Card Authors [optional]
|
||||
|
||||
[More Information Needed]
|
||||
|
||||
## Model Card Contact
|
||||
|
||||
[More Information Needed]
|
||||
### Framework versions
|
||||
|
||||
- PEFT 0.19.1
|
||||
48
adapter/adapter_config.json
Normal file
48
adapter/adapter_config.json
Normal file
@@ -0,0 +1,48 @@
|
||||
{
|
||||
"alora_invocation_tokens": null,
|
||||
"alpha_pattern": {},
|
||||
"arrow_config": null,
|
||||
"auto_mapping": null,
|
||||
"base_model_name_or_path": "Qwen/Qwen3-14B",
|
||||
"bias": "none",
|
||||
"corda_config": null,
|
||||
"ensure_weight_tying": false,
|
||||
"eva_config": null,
|
||||
"exclude_modules": null,
|
||||
"fan_in_fan_out": false,
|
||||
"inference_mode": true,
|
||||
"init_lora_weights": true,
|
||||
"layer_replication": null,
|
||||
"layers_pattern": null,
|
||||
"layers_to_transform": null,
|
||||
"loftq_config": {},
|
||||
"lora_alpha": 64,
|
||||
"lora_bias": false,
|
||||
"lora_dropout": 0.0,
|
||||
"lora_ga_config": null,
|
||||
"megatron_config": null,
|
||||
"megatron_core": "megatron.core",
|
||||
"modules_to_save": null,
|
||||
"peft_type": "LORA",
|
||||
"peft_version": "0.19.1",
|
||||
"qalora_group_size": 16,
|
||||
"r": 32,
|
||||
"rank_pattern": {},
|
||||
"revision": null,
|
||||
"target_modules": [
|
||||
"gate_proj",
|
||||
"v_proj",
|
||||
"k_proj",
|
||||
"up_proj",
|
||||
"down_proj",
|
||||
"o_proj",
|
||||
"q_proj"
|
||||
],
|
||||
"target_parameters": null,
|
||||
"task_type": "CAUSAL_LM",
|
||||
"trainable_token_indices": null,
|
||||
"use_bdlora": null,
|
||||
"use_dora": false,
|
||||
"use_qalora": false,
|
||||
"use_rslora": false
|
||||
}
|
||||
3
adapter/adapter_model.safetensors
Normal file
3
adapter/adapter_model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:68ee829cf0d836f834e05f7fd8048ed8728a005134279148191bcc08ac9559bb
|
||||
size 513877864
|
||||
89
adapter/chat_template.jinja
Normal file
89
adapter/chat_template.jinja
Normal file
@@ -0,0 +1,89 @@
|
||||
{%- if tools %}
|
||||
{{- '<|im_start|>system\n' }}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- messages[0].content + '\n\n' }}
|
||||
{%- endif %}
|
||||
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
||||
{%- for tool in tools %}
|
||||
{{- "\n" }}
|
||||
{{- tool | tojson }}
|
||||
{%- endfor %}
|
||||
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
||||
{%- else %}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
||||
{%- for message in messages[::-1] %}
|
||||
{%- set index = (messages|length - 1) - loop.index0 %}
|
||||
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
||||
{%- set ns.multi_step_tool = false %}
|
||||
{%- set ns.last_query_index = index %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- for message in messages %}
|
||||
{%- if message.content is string %}
|
||||
{%- set content = message.content %}
|
||||
{%- else %}
|
||||
{%- set content = '' %}
|
||||
{%- endif %}
|
||||
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
||||
{%- elif message.role == "assistant" %}
|
||||
{%- set reasoning_content = '' %}
|
||||
{%- if message.reasoning_content is string %}
|
||||
{%- set reasoning_content = message.reasoning_content %}
|
||||
{%- else %}
|
||||
{%- if '</think>' in content %}
|
||||
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
||||
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- if loop.index0 > ns.last_query_index %}
|
||||
{%- if loop.last or (not loop.last and reasoning_content) %}
|
||||
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- if message.tool_calls %}
|
||||
{%- for tool_call in message.tool_calls %}
|
||||
{%- if (loop.first and content) or (not loop.first) %}
|
||||
{{- '\n' }}
|
||||
{%- endif %}
|
||||
{%- if tool_call.function %}
|
||||
{%- set tool_call = tool_call.function %}
|
||||
{%- endif %}
|
||||
{{- '<tool_call>\n{"name": "' }}
|
||||
{{- tool_call.name }}
|
||||
{{- '", "arguments": ' }}
|
||||
{%- if tool_call.arguments is string %}
|
||||
{{- tool_call.arguments }}
|
||||
{%- else %}
|
||||
{{- tool_call.arguments | tojson }}
|
||||
{%- endif %}
|
||||
{{- '}\n</tool_call>' }}
|
||||
{%- endfor %}
|
||||
{%- endif %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- elif message.role == "tool" %}
|
||||
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
||||
{{- '<|im_start|>user' }}
|
||||
{%- endif %}
|
||||
{{- '\n<tool_response>\n' }}
|
||||
{{- content }}
|
||||
{{- '\n</tool_response>' }}
|
||||
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- if add_generation_prompt %}
|
||||
{{- '<|im_start|>assistant\n' }}
|
||||
{%- if enable_thinking is defined and enable_thinking is false %}
|
||||
{{- '<think>\n\n</think>\n\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
3
adapter/tokenizer.json
Normal file
3
adapter/tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506
|
||||
size 11422650
|
||||
29
adapter/tokenizer_config.json
Normal file
29
adapter/tokenizer_config.json
Normal file
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>",
|
||||
"<|object_ref_start|>",
|
||||
"<|object_ref_end|>",
|
||||
"<|box_start|>",
|
||||
"<|box_end|>",
|
||||
"<|quad_start|>",
|
||||
"<|quad_end|>",
|
||||
"<|vision_start|>",
|
||||
"<|vision_end|>",
|
||||
"<|vision_pad|>",
|
||||
"<|image_pad|>",
|
||||
"<|video_pad|>"
|
||||
],
|
||||
"is_local": false,
|
||||
"model_max_length": 131072,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null
|
||||
}
|
||||
3
adapter/training_args.bin
Normal file
3
adapter/training_args.bin
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:f7b8e0e2fb5df1b2a243361e988314c091bebb39a1d6987c307d909b5412bd56
|
||||
size 5201
|
||||
28
adele_judge_config.json
Normal file
28
adele_judge_config.json
Normal file
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"inference": {
|
||||
"allowed_scores": [
|
||||
"1",
|
||||
"2",
|
||||
"3",
|
||||
"4",
|
||||
"5"
|
||||
],
|
||||
"binary_threshold": 3,
|
||||
"method": "restricted_continuation_logprobs_fast"
|
||||
},
|
||||
"model": {
|
||||
"base_model": "Qwen/Qwen3-14B",
|
||||
"revision": null,
|
||||
"thinking_mode": {
|
||||
"apply_if_supported": true,
|
||||
"enabled": false
|
||||
}
|
||||
},
|
||||
"prompt": {
|
||||
"system_prompt": "You are an answer quality evaluator. You will receive a question, a reference\nanswer, and a model response, each delimited by ### headers.\n\nRate the model response on the following scale:\n1 - Surely incorrect: the response is wrong or completely misses the point.\n2 - Likely incorrect: the response has significant errors or critical omissions.\n3 - Minimally correct / sufficient: the response is partially correct or borderline.\n4 - Likely correct: the response is mostly correct with minor issues.\n5 - Surely correct: the response is fully correct and complete.\n\nRules:\n- Base your rating strictly on correctness relative to the reference answer.\n- Ignore writing style, verbosity, or formatting unless they affect correctness.\n- If the final answer is correct but the reasoning contains substantive errors, lower the rating according to the severity of those errors.\n- Output only a single integer from 1 to 5. No explanation. No punctuation.\n"
|
||||
},
|
||||
"training": {
|
||||
"max_seq_length": 4096,
|
||||
"objective": "restricted_score_ce"
|
||||
}
|
||||
}
|
||||
1059
adele_judge_metadata.json
Normal file
1059
adele_judge_metadata.json
Normal file
File diff suppressed because it is too large
Load Diff
304
adele_judge_pipeline.py
Normal file
304
adele_judge_pipeline.py
Normal file
@@ -0,0 +1,304 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import inspect
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from transformers import Pipeline
|
||||
|
||||
|
||||
THINKING_KWARG = "enable_thinking"
|
||||
DEFAULT_SYSTEM_PROMPT = "Return only one score from 1 to 5. Do not explain."
|
||||
DEFAULT_ALLOWED_SCORES = ["1", "2", "3", "4", "5"]
|
||||
DEFAULT_BINARY_THRESHOLD = 3
|
||||
|
||||
|
||||
def load_adele_judge_config(repo_id_or_path: str) -> dict[str, Any]:
|
||||
path = Path(repo_id_or_path) / "adele_judge_config.json"
|
||||
if path.exists():
|
||||
return json.loads(path.read_text(encoding="utf-8"))
|
||||
|
||||
from huggingface_hub import hf_hub_download
|
||||
|
||||
downloaded = hf_hub_download(repo_id_or_path, "adele_judge_config.json")
|
||||
return json.loads(Path(downloaded).read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def load_adele_judge_config_or_default(model: Any, tokenizer: Any) -> dict[str, Any]:
|
||||
candidates = [
|
||||
getattr(model, "name_or_path", None),
|
||||
getattr(getattr(model, "config", None), "_name_or_path", None),
|
||||
getattr(tokenizer, "name_or_path", None),
|
||||
getattr(tokenizer, "_name_or_path", None),
|
||||
]
|
||||
for candidate in candidates:
|
||||
if not candidate:
|
||||
continue
|
||||
try:
|
||||
return load_adele_judge_config(str(candidate))
|
||||
except Exception:
|
||||
continue
|
||||
return {}
|
||||
|
||||
|
||||
def adele_judge_settings(config: dict[str, Any] | None) -> dict[str, Any]:
|
||||
config = config or {}
|
||||
prompt_config = config.get("prompt", {}) if isinstance(config.get("prompt"), dict) else {}
|
||||
inference_config = (
|
||||
config.get("inference", {}) if isinstance(config.get("inference"), dict) else {}
|
||||
)
|
||||
model_config = config.get("model", {}) if isinstance(config.get("model"), dict) else {}
|
||||
return {
|
||||
"system_prompt": prompt_config.get("system_prompt") or DEFAULT_SYSTEM_PROMPT,
|
||||
"allowed_scores": [
|
||||
str(score)
|
||||
for score in inference_config.get("allowed_scores", DEFAULT_ALLOWED_SCORES)
|
||||
],
|
||||
"binary_threshold": int(
|
||||
inference_config.get("binary_threshold", DEFAULT_BINARY_THRESHOLD)
|
||||
),
|
||||
"thinking_mode": model_config.get("thinking_mode") or {},
|
||||
}
|
||||
|
||||
|
||||
def clean_value(value: Any, fallback: str = "N/A") -> str:
|
||||
if value is None:
|
||||
return fallback
|
||||
text = str(value)
|
||||
if not text or text.lower() == "nan":
|
||||
return fallback
|
||||
return text
|
||||
|
||||
|
||||
def validate_example(inputs: Any) -> dict[str, Any]:
|
||||
if not isinstance(inputs, dict):
|
||||
raise ValueError("ADeLe judge input must be a mapping")
|
||||
|
||||
missing = []
|
||||
if inputs.get("question") is None:
|
||||
missing.append("question")
|
||||
if inputs.get("model_response") is None:
|
||||
missing.append("model_response")
|
||||
reference_answer = inputs.get("reference_answer")
|
||||
if reference_answer is None:
|
||||
reference_answer = inputs.get("ground_truth")
|
||||
if reference_answer is None:
|
||||
missing.append("reference_answer or ground_truth")
|
||||
if missing:
|
||||
raise ValueError(f"Missing required field(s): {', '.join(missing)}")
|
||||
|
||||
return {
|
||||
"question": inputs["question"],
|
||||
"reference_answer": reference_answer,
|
||||
"model_response": inputs["model_response"],
|
||||
}
|
||||
|
||||
|
||||
def build_user_message(example: dict[str, Any]) -> str:
|
||||
return "\n\n".join(
|
||||
[
|
||||
f"### QUESTION\n{clean_value(example.get('question'))}",
|
||||
f"### REFERENCE ANSWER\n{clean_value(example.get('reference_answer'))}",
|
||||
f"### MODEL RESPONSE\n{clean_value(example.get('model_response'), fallback='')}",
|
||||
"### SCORE\n",
|
||||
]
|
||||
)
|
||||
|
||||
|
||||
def build_messages(example: dict[str, Any], system_prompt: str) -> list[dict[str, str]]:
|
||||
return [
|
||||
{"role": "system", "content": system_prompt.strip()},
|
||||
{"role": "user", "content": build_user_message(example)},
|
||||
]
|
||||
|
||||
|
||||
def chat_template_supports_thinking(tokenizer: Any) -> bool:
|
||||
apply_chat_template = getattr(tokenizer, "apply_chat_template", None)
|
||||
if apply_chat_template is None:
|
||||
return False
|
||||
try:
|
||||
signature = inspect.signature(apply_chat_template)
|
||||
except (TypeError, ValueError):
|
||||
return False
|
||||
accepts_kwarg = any(
|
||||
parameter.kind == inspect.Parameter.VAR_KEYWORD or name == THINKING_KWARG
|
||||
for name, parameter in signature.parameters.items()
|
||||
)
|
||||
if not accepts_kwarg:
|
||||
return False
|
||||
|
||||
template = getattr(tokenizer, "chat_template", None)
|
||||
if isinstance(template, str) and THINKING_KWARG in template:
|
||||
return True
|
||||
|
||||
candidates = [
|
||||
getattr(tokenizer, "name_or_path", None),
|
||||
getattr(tokenizer, "_name_or_path", None),
|
||||
getattr(tokenizer, "model_name", None),
|
||||
]
|
||||
init_kwargs = getattr(tokenizer, "init_kwargs", None)
|
||||
if isinstance(init_kwargs, dict):
|
||||
candidates.extend([init_kwargs.get("name_or_path"), init_kwargs.get("tokenizer_file")])
|
||||
return any("qwen3" in str(candidate).lower() for candidate in candidates if candidate)
|
||||
|
||||
|
||||
def apply_chat_template_safe(
|
||||
tokenizer: Any,
|
||||
messages: list[dict[str, str]],
|
||||
*,
|
||||
add_generation_prompt: bool,
|
||||
thinking_mode: dict[str, Any],
|
||||
) -> str:
|
||||
if hasattr(tokenizer, "apply_chat_template") and getattr(tokenizer, "chat_template", None):
|
||||
template_kwargs = {}
|
||||
enabled = thinking_mode.get("enabled")
|
||||
if (
|
||||
enabled is not None
|
||||
and bool(thinking_mode.get("apply_if_supported", True))
|
||||
and chat_template_supports_thinking(tokenizer)
|
||||
):
|
||||
template_kwargs[THINKING_KWARG] = bool(enabled)
|
||||
return tokenizer.apply_chat_template(
|
||||
messages,
|
||||
tokenize=False,
|
||||
add_generation_prompt=add_generation_prompt,
|
||||
**template_kwargs,
|
||||
)
|
||||
|
||||
rendered = [f"<|{message['role']}|>\n{message['content']}" for message in messages]
|
||||
if add_generation_prompt:
|
||||
rendered.append("<|assistant|>\n")
|
||||
return "\n".join(rendered)
|
||||
|
||||
|
||||
def encode_text(tokenizer: Any, text: str) -> list[int]:
|
||||
return tokenizer(text, add_special_tokens=False, truncation=False)["input_ids"]
|
||||
|
||||
|
||||
def single_score_token_ids(tokenizer: Any, allowed_scores: list[str]) -> list[int]:
|
||||
token_ids = [encode_text(tokenizer, score) for score in allowed_scores]
|
||||
multi_token_scores = [
|
||||
score for score, ids in zip(allowed_scores, token_ids, strict=True) if len(ids) != 1
|
||||
]
|
||||
if multi_token_scores:
|
||||
raise ValueError(
|
||||
"ADeLeJudgePipeline requires score continuations to be single tokens; "
|
||||
f"multi-token scores: {multi_token_scores}"
|
||||
)
|
||||
return [ids[0] for ids in token_ids]
|
||||
|
||||
|
||||
class ADeLeJudgePipeline(Pipeline):
|
||||
"""HF-native custom pipeline for restricted ADeLe judge scoring."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*args: Any,
|
||||
adele_config: dict[str, Any] | None = None,
|
||||
**kwargs: Any,
|
||||
) -> None:
|
||||
super().__init__(*args, **kwargs)
|
||||
if self.tokenizer is None:
|
||||
raise ValueError("ADeLeJudgePipeline requires a tokenizer")
|
||||
if getattr(self.tokenizer, "pad_token", None) is None:
|
||||
self.tokenizer.pad_token = getattr(self.tokenizer, "eos_token", None)
|
||||
|
||||
settings = adele_judge_settings(
|
||||
adele_config
|
||||
if adele_config is not None
|
||||
else load_adele_judge_config_or_default(self.model, self.tokenizer)
|
||||
)
|
||||
self.system_prompt = settings["system_prompt"]
|
||||
self.allowed_scores = settings["allowed_scores"]
|
||||
self.binary_threshold = settings["binary_threshold"]
|
||||
self.thinking_mode = settings["thinking_mode"]
|
||||
self.score_token_ids = single_score_token_ids(self.tokenizer, self.allowed_scores)
|
||||
|
||||
if hasattr(self.model, "eval"):
|
||||
self.model.eval()
|
||||
|
||||
def _sanitize_parameters(
|
||||
self,
|
||||
**kwargs: Any,
|
||||
) -> tuple[dict[str, Any], dict[str, Any], dict[str, Any]]:
|
||||
return {}, {}, {}
|
||||
|
||||
def preprocess(self, inputs: Any) -> dict[str, Any]:
|
||||
import torch
|
||||
|
||||
example = validate_example(inputs)
|
||||
prompt = apply_chat_template_safe(
|
||||
self.tokenizer,
|
||||
build_messages(example, self.system_prompt),
|
||||
add_generation_prompt=True,
|
||||
thinking_mode=self.thinking_mode,
|
||||
)
|
||||
encoded = self.tokenizer(
|
||||
prompt,
|
||||
add_special_tokens=False,
|
||||
truncation=False,
|
||||
return_tensors="pt",
|
||||
)
|
||||
if "attention_mask" not in encoded:
|
||||
encoded["attention_mask"] = torch.ones_like(encoded["input_ids"])
|
||||
return {"input_ids": encoded["input_ids"], "attention_mask": encoded["attention_mask"]}
|
||||
|
||||
def _forward(self, model_inputs: dict[str, Any]) -> dict[str, Any]:
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
|
||||
input_ids = model_inputs["input_ids"]
|
||||
attention_mask = model_inputs["attention_mask"]
|
||||
with torch.no_grad():
|
||||
outputs = self.model(input_ids=input_ids, attention_mask=attention_mask)
|
||||
|
||||
token_positions = torch.arange(input_ids.shape[1], device=input_ids.device).unsqueeze(0)
|
||||
positions = (attention_mask * token_positions).max(dim=1).values.to(dtype=torch.long)
|
||||
batch_indices = torch.arange(input_ids.shape[0], device=input_ids.device)
|
||||
final_logits = outputs.logits[batch_indices, positions]
|
||||
|
||||
score_ids = torch.tensor(self.score_token_ids, dtype=torch.long, device=final_logits.device)
|
||||
score_logits = final_logits[:, score_ids]
|
||||
logprobs = F.log_softmax(score_logits, dim=-1)
|
||||
return {
|
||||
"score_indices": torch.argmax(logprobs, dim=-1),
|
||||
"probs": torch.exp(logprobs),
|
||||
"logprobs": logprobs,
|
||||
}
|
||||
|
||||
def postprocess(self, model_outputs: dict[str, Any]) -> dict[str, Any]:
|
||||
import torch
|
||||
|
||||
score_index = int(model_outputs["score_indices"].reshape(-1)[0])
|
||||
probs_tensor = model_outputs["probs"].reshape(-1, len(self.allowed_scores))[0]
|
||||
logprobs_tensor = model_outputs["logprobs"].reshape(-1, len(self.allowed_scores))[0]
|
||||
|
||||
probs = {
|
||||
score: float(prob)
|
||||
for score, prob in zip(self.allowed_scores, probs_tensor.tolist(), strict=True)
|
||||
}
|
||||
logprobs = {
|
||||
score: float(logprob)
|
||||
for score, logprob in zip(self.allowed_scores, logprobs_tensor.tolist(), strict=True)
|
||||
}
|
||||
|
||||
score = int(self.allowed_scores[score_index])
|
||||
sorted_logprobs = torch.sort(logprobs_tensor).values
|
||||
margin = (
|
||||
float(sorted_logprobs[-1] - sorted_logprobs[-2])
|
||||
if len(sorted_logprobs) > 1
|
||||
else 0.0
|
||||
)
|
||||
entropy = float(
|
||||
-(probs_tensor * torch.log(torch.clamp(probs_tensor, min=1e-12))).sum()
|
||||
)
|
||||
return {
|
||||
"score": score,
|
||||
"label": "CORRECT" if score >= self.binary_threshold else "INCORRECT",
|
||||
"probs": probs,
|
||||
"logprobs": logprobs,
|
||||
"confidence": max(probs.values()),
|
||||
"margin": margin,
|
||||
"entropy": entropy,
|
||||
}
|
||||
89
chat_template.jinja
Normal file
89
chat_template.jinja
Normal file
@@ -0,0 +1,89 @@
|
||||
{%- if tools %}
|
||||
{{- '<|im_start|>system\n' }}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- messages[0].content + '\n\n' }}
|
||||
{%- endif %}
|
||||
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
||||
{%- for tool in tools %}
|
||||
{{- "\n" }}
|
||||
{{- tool | tojson }}
|
||||
{%- endfor %}
|
||||
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
||||
{%- else %}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
||||
{%- for message in messages[::-1] %}
|
||||
{%- set index = (messages|length - 1) - loop.index0 %}
|
||||
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
||||
{%- set ns.multi_step_tool = false %}
|
||||
{%- set ns.last_query_index = index %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- for message in messages %}
|
||||
{%- if message.content is string %}
|
||||
{%- set content = message.content %}
|
||||
{%- else %}
|
||||
{%- set content = '' %}
|
||||
{%- endif %}
|
||||
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
||||
{%- elif message.role == "assistant" %}
|
||||
{%- set reasoning_content = '' %}
|
||||
{%- if message.reasoning_content is string %}
|
||||
{%- set reasoning_content = message.reasoning_content %}
|
||||
{%- else %}
|
||||
{%- if '</think>' in content %}
|
||||
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
||||
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- if loop.index0 > ns.last_query_index %}
|
||||
{%- if loop.last or (not loop.last and reasoning_content) %}
|
||||
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- if message.tool_calls %}
|
||||
{%- for tool_call in message.tool_calls %}
|
||||
{%- if (loop.first and content) or (not loop.first) %}
|
||||
{{- '\n' }}
|
||||
{%- endif %}
|
||||
{%- if tool_call.function %}
|
||||
{%- set tool_call = tool_call.function %}
|
||||
{%- endif %}
|
||||
{{- '<tool_call>\n{"name": "' }}
|
||||
{{- tool_call.name }}
|
||||
{{- '", "arguments": ' }}
|
||||
{%- if tool_call.arguments is string %}
|
||||
{{- tool_call.arguments }}
|
||||
{%- else %}
|
||||
{{- tool_call.arguments | tojson }}
|
||||
{%- endif %}
|
||||
{{- '}\n</tool_call>' }}
|
||||
{%- endfor %}
|
||||
{%- endif %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- elif message.role == "tool" %}
|
||||
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
||||
{{- '<|im_start|>user' }}
|
||||
{%- endif %}
|
||||
{{- '\n<tool_response>\n' }}
|
||||
{{- content }}
|
||||
{{- '\n</tool_response>' }}
|
||||
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- if add_generation_prompt %}
|
||||
{{- '<|im_start|>assistant\n' }}
|
||||
{%- if enable_thinking is defined and enable_thinking is false %}
|
||||
{{- '<think>\n\n</think>\n\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
85
config.json
Normal file
85
config.json
Normal file
@@ -0,0 +1,85 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen3ForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 151643,
|
||||
"custom_pipelines": {
|
||||
"adele-judge": {
|
||||
"impl": "adele_judge_pipeline.ADeLeJudgePipeline",
|
||||
"pt": [
|
||||
"AutoModelForCausalLM"
|
||||
],
|
||||
"tf": [],
|
||||
"type": "text"
|
||||
}
|
||||
},
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 151645,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 5120,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 17408,
|
||||
"layer_types": [
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention"
|
||||
],
|
||||
"max_position_embeddings": 40960,
|
||||
"max_window_layers": 40,
|
||||
"model_type": "qwen3",
|
||||
"num_attention_heads": 40,
|
||||
"num_hidden_layers": 40,
|
||||
"num_key_value_heads": 8,
|
||||
"pad_token_id": null,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 1000000,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": false,
|
||||
"transformers_version": "5.5.0",
|
||||
"use_cache": true,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 151936
|
||||
}
|
||||
5
generation_config.json
Normal file
5
generation_config.json
Normal file
@@ -0,0 +1,5 @@
|
||||
{
|
||||
"do_sample": false,
|
||||
"max_new_tokens": 1,
|
||||
"num_beams": 1
|
||||
}
|
||||
3
model-00001-of-00006.safetensors
Normal file
3
model-00001-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:5739c7705041195574706abc2af5e12c66c2281e1a93dd1419dd2439c0fd1a41
|
||||
size 4978180848
|
||||
3
model-00002-of-00006.safetensors
Normal file
3
model-00002-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:418628dcaf69131e39dff0aa22339698bf3111fd402a0c677ff8b1fb08a146de
|
||||
size 4917988408
|
||||
3
model-00003-of-00006.safetensors
Normal file
3
model-00003-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:7e63ccd5fd8e91297f66be63117190fbe6c5ad636c669d33263f490c42270e3a
|
||||
size 4991388720
|
||||
3
model-00004-of-00006.safetensors
Normal file
3
model-00004-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:483606985498d68d803f7761870f76b0df2965034149f2d4af7f7c93bf707c3b
|
||||
size 4917988496
|
||||
3
model-00005-of-00006.safetensors
Normal file
3
model-00005-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:acd0b2c97fb9d8fda9360f6293ccf899914f43d54ab8bcba3557ac62412451bd
|
||||
size 4991388720
|
||||
3
model-00006-of-00006.safetensors
Normal file
3
model-00006-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:2e2e7fd55cd67b4370f657b6f1d3ee9673f8a8657e4a90eaf873bed6661744ad
|
||||
size 4739730440
|
||||
451
model.safetensors.index.json
Normal file
451
model.safetensors.index.json
Normal file
@@ -0,0 +1,451 @@
|
||||
{
|
||||
"metadata": {
|
||||
"total_parameters": 14768307200,
|
||||
"total_size": 29536614400
|
||||
},
|
||||
"weight_map": {
|
||||
"lm_head.weight": "model-00001-of-00006.safetensors",
|
||||
"model.embed_tokens.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.10.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.20.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.3.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.30.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.4.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.norm.weight": "model-00006-of-00006.safetensors"
|
||||
}
|
||||
}
|
||||
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506
|
||||
size 11422650
|
||||
29
tokenizer_config.json
Normal file
29
tokenizer_config.json
Normal file
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>",
|
||||
"<|object_ref_start|>",
|
||||
"<|object_ref_end|>",
|
||||
"<|box_start|>",
|
||||
"<|box_end|>",
|
||||
"<|quad_start|>",
|
||||
"<|quad_end|>",
|
||||
"<|vision_start|>",
|
||||
"<|vision_end|>",
|
||||
"<|vision_pad|>",
|
||||
"<|image_pad|>",
|
||||
"<|video_pad|>"
|
||||
],
|
||||
"is_local": true,
|
||||
"model_max_length": 131072,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null
|
||||
}
|
||||
168
training_config.yaml
Normal file
168
training_config.yaml
Normal file
@@ -0,0 +1,168 @@
|
||||
project:
|
||||
run_name: qwen3_14b_restricted_score_ce
|
||||
output_dir: runs/qwen3_14b_restricted_score_ce
|
||||
seed: 42
|
||||
data:
|
||||
path: data/processed/response_scores.parquet
|
||||
prepared_dir: null
|
||||
columns:
|
||||
question: question
|
||||
reference_answer: ground_truth
|
||||
response: response
|
||||
judge_1_score: score_gpt4o
|
||||
judge_2_score: score_sonnet
|
||||
model_id: model_id
|
||||
benchmark: benchmark
|
||||
task: task
|
||||
example_id: instance_id
|
||||
source: source
|
||||
filters:
|
||||
max_disagreement: 1
|
||||
max_response_tokens: 4096
|
||||
on_sequence_overflow: skip
|
||||
preprocessing_num_workers: 40
|
||||
tokenizers_parallelism: true
|
||||
token_length_batch_size: 2048
|
||||
model:
|
||||
model_name_or_path: Qwen/Qwen3-14B
|
||||
revision: null
|
||||
attn_implementation: sdpa
|
||||
adapter_path: null
|
||||
trust_remote_code: true
|
||||
thinking_mode:
|
||||
enabled: false
|
||||
apply_if_supported: true
|
||||
prompt:
|
||||
system_prompt: 'You are an answer quality evaluator. You will receive a question,
|
||||
a reference
|
||||
|
||||
answer, and a model response, each delimited by ### headers.
|
||||
|
||||
|
||||
Rate the model response on the following scale:
|
||||
|
||||
1 - Surely incorrect: the response is wrong or completely misses the point.
|
||||
|
||||
2 - Likely incorrect: the response has significant errors or critical omissions.
|
||||
|
||||
3 - Minimally correct / sufficient: the response is partially correct or borderline.
|
||||
|
||||
4 - Likely correct: the response is mostly correct with minor issues.
|
||||
|
||||
5 - Surely correct: the response is fully correct and complete.
|
||||
|
||||
|
||||
Rules:
|
||||
|
||||
- Base your rating strictly on correctness relative to the reference answer.
|
||||
|
||||
- Ignore writing style, verbosity, or formatting unless they affect correctness.
|
||||
|
||||
- If the final answer is correct but the reasoning contains substantive errors,
|
||||
lower the rating according to the severity of those errors.
|
||||
|
||||
- Output only a single integer from 1 to 5. No explanation. No punctuation.
|
||||
|
||||
'
|
||||
split:
|
||||
mode: fixed_by_model
|
||||
validation_models:
|
||||
- gemini-3-flash
|
||||
- DK-R1-Dist-Qwen-14B
|
||||
- llama3d2-3b
|
||||
train_models: auto_except_val_test
|
||||
held_out_model: null
|
||||
lomo_validation_fraction: 0.05
|
||||
lomo_validation_max_examples: 30000
|
||||
lomo_validation_seed: 42
|
||||
training:
|
||||
max_seq_length: 4096
|
||||
load_in_4bit: true
|
||||
dtype: bfloat16
|
||||
objective: restricted_score_ce
|
||||
loss:
|
||||
type: ce_5way
|
||||
lambda_binary: 0.5
|
||||
class_weights: null
|
||||
class_weighting: null
|
||||
score_class_weights: null
|
||||
lora_r: 32
|
||||
lora_alpha: 64
|
||||
lora_dropout: 0.0
|
||||
target_modules: auto
|
||||
learning_rate: 3.0e-05
|
||||
num_train_epochs: 1
|
||||
per_device_train_batch_size: 2
|
||||
per_device_eval_batch_size: 2
|
||||
gradient_accumulation_steps: 4
|
||||
warmup_ratio: 0.03
|
||||
lr_scheduler_type: cosine
|
||||
weight_decay: 0.0
|
||||
optim: adamw_8bit
|
||||
packing: false
|
||||
cache_tokenized_datasets: true
|
||||
eval_subset_size: null
|
||||
eval_subset_strategy: stratified
|
||||
eval_subset_stratify_columns:
|
||||
- model_id
|
||||
- target_score
|
||||
train_sampling_strategy: random
|
||||
length_column_name: length
|
||||
logging_steps: 10
|
||||
eval_steps: 500
|
||||
save_steps: 500
|
||||
save_total_limit: 10
|
||||
seed: 42
|
||||
resume_from_checkpoint: null
|
||||
eval_subset_seed: 42
|
||||
max_grad_norm: 1.0
|
||||
distributed:
|
||||
enabled: true
|
||||
strategy: ddp
|
||||
backend: nccl
|
||||
mixed_precision: bf16
|
||||
gradient_checkpointing: false
|
||||
find_unused_parameters: false
|
||||
fsdp:
|
||||
sharding_strategy: full_shard
|
||||
transformer_layer_cls_to_wrap: null
|
||||
activation_checkpointing: true
|
||||
use_orig_params: true
|
||||
deepspeed:
|
||||
zero_stage: 2
|
||||
offload_optimizer_device: none
|
||||
offload_param_device: none
|
||||
stage3_gather_16bit_weights_on_model_save: true
|
||||
gradient_clipping: auto
|
||||
config_overrides: {}
|
||||
inference:
|
||||
allowed_scores:
|
||||
- '1'
|
||||
- '2'
|
||||
- '3'
|
||||
- '4'
|
||||
- '5'
|
||||
binary_threshold: 3
|
||||
method: restricted_continuation_logprobs_fast
|
||||
generation_fallback: false
|
||||
batch_size: 64
|
||||
require_adapter: false
|
||||
allow_base_model: true
|
||||
evaluation:
|
||||
length_buckets:
|
||||
- 0
|
||||
- 256
|
||||
- 512
|
||||
- 1024
|
||||
- 2048
|
||||
- 3072
|
||||
- 4096
|
||||
- 1000000000
|
||||
hub:
|
||||
repo_id: null
|
||||
private: false
|
||||
commit_message: Upload ADeLe distilled judge
|
||||
local_checkpoint_dir: null
|
||||
output_staging_dir: null
|
||||
create_pr: false
|
||||
max_shard_size: 5GB
|
||||
Reference in New Issue
Block a user