初始化项目,由ModelHub XC社区提供模型
Model: yuanmodel/Qwen3-0.6B-GRPO-GSM8K-Think Source: Original Platform
This commit is contained in:
49
.gitattributes
vendored
Normal file
49
.gitattributes
vendored
Normal file
@@ -0,0 +1,49 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zstandard filter=lfs diff=lfs merge=lfs -text
|
||||
*.tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
*.db* filter=lfs diff=lfs merge=lfs -text
|
||||
*.ark* filter=lfs diff=lfs merge=lfs -text
|
||||
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
|
||||
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
|
||||
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.gguf* filter=lfs diff=lfs merge=lfs -text
|
||||
*.ggml filter=lfs diff=lfs merge=lfs -text
|
||||
*.llamafile* filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
201
README.md
Normal file
201
README.md
Normal file
@@ -0,0 +1,201 @@
|
||||
---
|
||||
base_model: Qwen/Qwen3-0.6B
|
||||
library_name: transformers
|
||||
model_name: Qwen3-0.6B-GRPO-GSM8K-Think-Liger
|
||||
tags:
|
||||
- generated_from_trainer
|
||||
- grpo
|
||||
- trl
|
||||
- Liger
|
||||
licence: license
|
||||
license: Apache License 2.0
|
||||
language:
|
||||
- en
|
||||
frameworks: PyTorch
|
||||
base_model_relation: finetune
|
||||
metrics:
|
||||
- accuracy
|
||||
datasets:
|
||||
- modelscope/gsm8k
|
||||
---
|
||||
# Model Card for Qwen3-0.6B-GRPO-GSM8K-Think-Liger
|
||||
|
||||
This model is a fine-tuned version of [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B).
|
||||
It has been trained using [TRL](https://github.com/huggingface/trl).
|
||||
|
||||
## Quick start
|
||||
|
||||
```python
|
||||
from modelscope import pipeline
|
||||
from modelscope.msdatasets import MsDataset
|
||||
ds = MsDataset.load('modelscope/gsm8k', subset_name='main', split='test',trust_remote_code=True)
|
||||
sample = ds[0]
|
||||
question = sample['question']
|
||||
true_answer = sample['answer']
|
||||
print(question,true_answer)
|
||||
|
||||
generator = pipeline(
|
||||
task="text-generation",
|
||||
model="yuanmodel/Qwen3-0.6B-GRPO-GSM8K-Think",
|
||||
device="cuda",trust_remote_code=True
|
||||
)
|
||||
|
||||
output = generator([{"role": "user", "content": question}], max_new_tokens=512)
|
||||
print("Raw output:", output)
|
||||
print("answer", output["message"]["content"])
|
||||
```
|
||||
|
||||
```markdown
|
||||
Janet’s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market? Janet sells 16 - 3 - 4 = <<16-3-4=9>>9 duck eggs a day.
|
||||
She makes 9 * 2 = $<<9*2=18>>18 every day at the farmer’s market.
|
||||
#### 18
|
||||
|
||||
|
||||
Raw output: {'message': {'role': 'assistant', 'content': "<think>\nOkay, let me try to figure out how much Janet makes every day at the farmers' market. Hmm, so first, I need to break down her daily activities and calculate each part. Let me start by listing out what she does each day.\n\nJanet's ducks lay 16 eggs per day. So, if she does that every day, the total eggs laid would be 16 eggs. Then, she eats three eggs for breakfast. So, that's 3 eggs. Then, she bakes muffins for her friends every day with four eggs per muffin. So, 4 eggs per muffin times the number of days she bakes. Since it's a daily activity, that's 4 times 1, so 4 eggs total for muffins. \n\nAfter that, the remainder of the eggs is sold at the farmers' market for $2 per egg. So, the eggs remaining after breakfast and muffins would be 16 eggs minus 3 eggs (breakfast) minus 4 eggs (muffins) equals 9 eggs. Then, 9 eggs multiplied by $2 per egg gives the total amount for the market.\n\nLet me check the math step by step to make sure I didn't make a mistake. 16 eggs total. Subtract 3 for breakfast: 16 - 3 = 13 eggs left. Then 13 - 4 = 9 eggs. Multiply 9 by 2: 18 dollars. So, $18 should be the amount she makes at the market. That seems right.\n\nWait, just to confirm, 16 - 3 is 13, 13 - 4 is 9, 9*2 is 18. Yep, that adds up. I don't think I missed anything here. So the answer should be $18.\n</think>\n\nJanet's ducks lay 16 eggs per day. She eats 3 eggs for breakfast, and bakes 4 eggs for muffins daily. The remainder is sold at the farmers' market for $2 per egg.\n\n1. Total eggs: $16$\n2. Eggs eaten: $16 - 3 = 13$\n3. Eggs remaining: $13 - 4 = 9$\n4. Revenue from market: $9 \\times 2 = 18$\n\n**Answer:** Janet makes $\\boxed{18}$ dollars every day at the farmers' market."}}
|
||||
answer <think>
|
||||
Okay, let me try to figure out how much Janet makes every day at the farmers' market. Hmm, so first, I need to break down her daily activities and calculate each part. Let me start by listing out what she does each day.
|
||||
|
||||
Janet's ducks lay 16 eggs per day. So, if she does that every day, the total eggs laid would be 16 eggs. Then, she eats three eggs for breakfast. So, that's 3 eggs. Then, she bakes muffins for her friends every day with four eggs per muffin. So, 4 eggs per muffin times the number of days she bakes. Since it's a daily activity, that's 4 times 1, so 4 eggs total for muffins.
|
||||
|
||||
After that, the remainder of the eggs is sold at the farmers' market for $2 per egg. So, the eggs remaining after breakfast and muffins would be 16 eggs minus 3 eggs (breakfast) minus 4 eggs (muffins) equals 9 eggs. Then, 9 eggs multiplied by $2 per egg gives the total amount for the market.
|
||||
|
||||
Let me check the math step by step to make sure I didn't make a mistake. 16 eggs total. Subtract 3 for breakfast: 16 - 3 = 13 eggs left. Then 13 - 4 = 9 eggs. Multiply 9 by 2: 18 dollars. So, $18 should be the amount she makes at the market. That seems right.
|
||||
|
||||
Wait, just to confirm, 16 - 3 is 13, 13 - 4 is 9, 9*2 is 18. Yep, that adds up. I don't think I missed anything here. So the answer should be $18.
|
||||
</think>
|
||||
|
||||
Janet's ducks lay 16 eggs per day. She eats 3 eggs for breakfast, and bakes 4 eggs for muffins daily. The remainder is sold at the farmers' market for $2 per egg.
|
||||
|
||||
1. Total eggs: $16$
|
||||
2. Eggs eaten: $16 - 3 = 13$
|
||||
3. Eggs remaining: $13 - 4 = 9$
|
||||
4. Revenue from market: $9 \times 2 = 18$
|
||||
|
||||
**Answer:** Janet makes $\boxed{18}$ dollars every day at the farmers' market.
|
||||
```
|
||||
|
||||
## Quick start with chat template
|
||||
|
||||
```python
|
||||
from modelscope import pipeline
|
||||
from modelscope.msdatasets import MsDataset
|
||||
|
||||
SYSTEM_PROMPT = """
|
||||
Please solve the math problem step by step. Use <think> tags to show your reasoning process, then provide the final numerical answer.
|
||||
|
||||
IMPORTANT: Your final answer must be a pure number only, following these rules:
|
||||
- No currency symbols: use 42, not $42 or 42 dollars
|
||||
- No thousand separators: use 16000, not 16,000 or 16 000
|
||||
- No decimal if integer: use 42, not 42.0 or 42.00
|
||||
- No units or text: use 144, not 144 gallons or 144 total
|
||||
|
||||
Examples:
|
||||
- Good: 42, 16000, 144, 999000
|
||||
- Bad: $42, 16,000, 42.0, 144 gallons
|
||||
|
||||
Format:
|
||||
<think>
|
||||
Your step-by-step reasoning here...
|
||||
</think>
|
||||
Final answer: [pure number only]
|
||||
"""
|
||||
|
||||
|
||||
ds = MsDataset.load('modelscope/gsm8k', subset_name='main', split='test', trust_remote_code=True)
|
||||
sample = ds[0]
|
||||
question = sample['question']
|
||||
true_answer = sample['answer']
|
||||
print("Question:", question)
|
||||
print("True answer:", true_answer)
|
||||
|
||||
messages = [
|
||||
{"role": "system", "content": SYSTEM_PROMPT},
|
||||
{"role": "user", "content": question}
|
||||
]
|
||||
|
||||
generator = pipeline(
|
||||
task="text-generation",
|
||||
model="yuanmodel/Qwen3-0.6B-GRPO-GSM8K-Think",
|
||||
device="cuda",
|
||||
trust_remote_code=True
|
||||
)
|
||||
|
||||
output = generator(messages, max_new_tokens=512)
|
||||
print("Raw output:", output)
|
||||
print("Model answer:", output["message"]["content"])
|
||||
```
|
||||
|
||||
|
||||
```markdown
|
||||
Raw output: {'message': {'role': 'assistant', 'content': "<think>\nOkay, let's see. Janet's ducks lay 16 eggs per day. She eats three eggs for breakfast and four eggs for baking muffins each day. Then she sells the rest at $2 per egg. So, I need to calculate how many eggs she has left each day, multiply that by $2, and that should give the total amount she makes.\n\nFirst, let's find out the total eggs per day. She lays 16 eggs, eats 3 + 4 = 7 eggs. So total eggs left = 16 - 7 = 9 eggs per day. Then multiply by $2: 9 * 2 = $18. So, the answer should be $18.\n</think>\n\nFinal answer: 18"}}
|
||||
Model answer: <think>
|
||||
Okay, let's see. Janet's ducks lay 16 eggs per day. She eats three eggs for breakfast and four eggs for baking muffins each day. Then she sells the rest at $2 per egg. So, I need to calculate how many eggs she has left each day, multiply that by $2, and that should give the total amount she makes.
|
||||
|
||||
First, let's find out the total eggs per day. She lays 16 eggs, eats 3 + 4 = 7 eggs. So total eggs left = 16 - 7 = 9 eggs per day. Then multiply by $2: 9 * 2 = $18. So, the answer should be $18.
|
||||
</think>
|
||||
|
||||
Final answer: 18
|
||||
|
||||
```
|
||||
|
||||
## Training procedure
|
||||
|
||||
|
||||
|
||||
|
||||
This model was trained with GRPO, a method introduced in [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://huggingface.co/papers/2402.03300).
|
||||
|
||||
|
||||
|
||||
|
||||
# GSM8K Validation and Test Set Evaluation Results
|
||||
|
||||
| Model | GSM8K-Val | GSM8K-Test | Improvement |
|
||||
|-------------------------------------|-----------|------------|-------------|
|
||||
| Qwen/Qwen3-0.6B (official) | - | 59.59% | - |
|
||||
| Qwen/Qwen3-0.6B (local reproduction)| 53.25% | 56.33% | - |
|
||||
| Qwen/Qwen3-0.6B GRPO-checkpoint-160| - | 59.67% | +5.92% |
|
||||
| Qwen/Qwen3-0.6B GRPO-checkpoint-180| 62.8% | 60.65% | +7.66% |
|
||||
|
||||
- "official" refers to results published by ModelScope or HuggingFace.
|
||||
- "local reproduction" refers to results reproduced by the local pipeline.
|
||||
- "GRPO-checkpoint-160/180" are intermediate/final checkpoints trained in this project.
|
||||
- "Improvement" is relative to the local reproduction Qwen3-0.6B baseline.
|
||||
|
||||
|
||||
|
||||
|
||||
### Framework versions
|
||||
|
||||
- TRL: 0.19.0
|
||||
- Transformers: 4.52.4
|
||||
- Pytorch: 2.7.1
|
||||
- Datasets: 3.6.0
|
||||
- Tokenizers: 0.21.2
|
||||
|
||||
## Citations
|
||||
|
||||
Cite GRPO as:
|
||||
|
||||
```bibtex
|
||||
@article{zhihong2024deepseekmath,
|
||||
title = {{DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models}},
|
||||
author = {Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo},
|
||||
year = 2024,
|
||||
eprint = {arXiv:2402.03300},
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
Cite TRL as:
|
||||
|
||||
```bibtex
|
||||
@misc{vonwerra2022trl,
|
||||
title = {{TRL: Transformer Reinforcement Learning}},
|
||||
author = {Leandro von Werra and Younes Belkada and Lewis Tunstall and Edward Beeching and Tristan Thrush and Nathan Lambert and Shengyi Huang and Kashif Rasul and Quentin Gallou{\'e}dec},
|
||||
year = 2020,
|
||||
journal = {GitHub repository},
|
||||
publisher = {GitHub},
|
||||
howpublished = {\url{https://github.com/huggingface/trl}}
|
||||
}
|
||||
```
|
||||
28
added_tokens.json
Normal file
28
added_tokens.json
Normal file
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"</think>": 151668,
|
||||
"</tool_call>": 151658,
|
||||
"</tool_response>": 151666,
|
||||
"<think>": 151667,
|
||||
"<tool_call>": 151657,
|
||||
"<tool_response>": 151665,
|
||||
"<|box_end|>": 151649,
|
||||
"<|box_start|>": 151648,
|
||||
"<|endoftext|>": 151643,
|
||||
"<|file_sep|>": 151664,
|
||||
"<|fim_middle|>": 151660,
|
||||
"<|fim_pad|>": 151662,
|
||||
"<|fim_prefix|>": 151659,
|
||||
"<|fim_suffix|>": 151661,
|
||||
"<|im_end|>": 151645,
|
||||
"<|im_start|>": 151644,
|
||||
"<|image_pad|>": 151655,
|
||||
"<|object_ref_end|>": 151647,
|
||||
"<|object_ref_start|>": 151646,
|
||||
"<|quad_end|>": 151651,
|
||||
"<|quad_start|>": 151650,
|
||||
"<|repo_name|>": 151663,
|
||||
"<|video_pad|>": 151656,
|
||||
"<|vision_end|>": 151653,
|
||||
"<|vision_pad|>": 151654,
|
||||
"<|vision_start|>": 151652
|
||||
}
|
||||
89
chat_template.jinja
Normal file
89
chat_template.jinja
Normal file
@@ -0,0 +1,89 @@
|
||||
{%- if tools %}
|
||||
{{- '<|im_start|>system\n' }}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- messages[0].content + '\n\n' }}
|
||||
{%- endif %}
|
||||
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
||||
{%- for tool in tools %}
|
||||
{{- "\n" }}
|
||||
{{- tool | tojson }}
|
||||
{%- endfor %}
|
||||
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
||||
{%- else %}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
||||
{%- for message in messages[::-1] %}
|
||||
{%- set index = (messages|length - 1) - loop.index0 %}
|
||||
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
||||
{%- set ns.multi_step_tool = false %}
|
||||
{%- set ns.last_query_index = index %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- for message in messages %}
|
||||
{%- if message.content is string %}
|
||||
{%- set content = message.content %}
|
||||
{%- else %}
|
||||
{%- set content = '' %}
|
||||
{%- endif %}
|
||||
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
||||
{%- elif message.role == "assistant" %}
|
||||
{%- set reasoning_content = '' %}
|
||||
{%- if message.reasoning_content is string %}
|
||||
{%- set reasoning_content = message.reasoning_content %}
|
||||
{%- else %}
|
||||
{%- if '</think>' in content %}
|
||||
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
||||
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- if loop.index0 > ns.last_query_index %}
|
||||
{%- if loop.last or (not loop.last and reasoning_content) %}
|
||||
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- if message.tool_calls %}
|
||||
{%- for tool_call in message.tool_calls %}
|
||||
{%- if (loop.first and content) or (not loop.first) %}
|
||||
{{- '\n' }}
|
||||
{%- endif %}
|
||||
{%- if tool_call.function %}
|
||||
{%- set tool_call = tool_call.function %}
|
||||
{%- endif %}
|
||||
{{- '<tool_call>\n{"name": "' }}
|
||||
{{- tool_call.name }}
|
||||
{{- '", "arguments": ' }}
|
||||
{%- if tool_call.arguments is string %}
|
||||
{{- tool_call.arguments }}
|
||||
{%- else %}
|
||||
{{- tool_call.arguments | tojson }}
|
||||
{%- endif %}
|
||||
{{- '}\n</tool_call>' }}
|
||||
{%- endfor %}
|
||||
{%- endif %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- elif message.role == "tool" %}
|
||||
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
||||
{{- '<|im_start|>user' }}
|
||||
{%- endif %}
|
||||
{{- '\n<tool_response>\n' }}
|
||||
{{- content }}
|
||||
{{- '\n</tool_response>' }}
|
||||
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- if add_generation_prompt %}
|
||||
{{- '<|im_start|>assistant\n' }}
|
||||
{%- if enable_thinking is defined and enable_thinking is false %}
|
||||
{{- '<think>\n\n</think>\n\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
33
config.json
Normal file
33
config.json
Normal file
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen3ForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 151643,
|
||||
"eos_token_id": 151645,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 1024,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 3072,
|
||||
"max_position_embeddings": 40960,
|
||||
"max_window_layers": 28,
|
||||
"model_type": "qwen3",
|
||||
"num_attention_heads": 16,
|
||||
"num_hidden_layers": 28,
|
||||
"num_key_value_heads": 8,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_scaling": null,
|
||||
"rope_theta": 1000000,
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": true,
|
||||
"torch_dtype": "float32",
|
||||
"transformers_version": "4.52.4",
|
||||
"use_cache": false,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 151936,
|
||||
"tp_parallel": false,
|
||||
"tp_size": 1,
|
||||
"_tp_plan": {}
|
||||
}
|
||||
1
configuration.json
Normal file
1
configuration.json
Normal file
@@ -0,0 +1 @@
|
||||
{"framework":"Pytorch","task":"text-generation"}
|
||||
13
generation_config.json
Normal file
13
generation_config.json
Normal file
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"bos_token_id": 151643,
|
||||
"do_sample": true,
|
||||
"eos_token_id": [
|
||||
151645,
|
||||
151643
|
||||
],
|
||||
"pad_token_id": 151643,
|
||||
"temperature": 0.6,
|
||||
"top_k": 20,
|
||||
"top_p": 0.95,
|
||||
"transformers_version": "4.52.4"
|
||||
}
|
||||
151388
merges.txt
Normal file
151388
merges.txt
Normal file
File diff suppressed because it is too large
Load Diff
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:18f1d4922ea674cb93d872314bb2ba0d6f909ae00d153a35f424adc71807ed32
|
||||
size 2384234968
|
||||
3
optimizer.pt
Normal file
3
optimizer.pt
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:56a0f4f9c40c9ec72f35bc26ec25cf635209e875870b40811301a7ce277ec1f0
|
||||
size 4768663315
|
||||
3
rng_state.pth
Normal file
3
rng_state.pth
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:03a4c151da9d29c0f8fc90d74a54cbaf6905cb0d2d63d592a4340464691d231f
|
||||
size 14645
|
||||
3
scheduler.pt
Normal file
3
scheduler.pt
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:1d294b432655589c7578fe4d1043bdcc928c911d3943d27df13ee43d94f1d95d
|
||||
size 1465
|
||||
31
special_tokens_map.json
Normal file
31
special_tokens_map.json
Normal file
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>",
|
||||
"<|object_ref_start|>",
|
||||
"<|object_ref_end|>",
|
||||
"<|box_start|>",
|
||||
"<|box_end|>",
|
||||
"<|quad_start|>",
|
||||
"<|quad_end|>",
|
||||
"<|vision_start|>",
|
||||
"<|vision_end|>",
|
||||
"<|vision_pad|>",
|
||||
"<|image_pad|>",
|
||||
"<|video_pad|>"
|
||||
],
|
||||
"eos_token": {
|
||||
"content": "<|im_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"pad_token": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
}
|
||||
}
|
||||
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:67cc0080ffd7555f723f423c27cfef314e1ad9d335c8b79f465c5faba1ed478b
|
||||
size 11422821
|
||||
240
tokenizer_config.json
Normal file
240
tokenizer_config.json
Normal file
@@ -0,0 +1,240 @@
|
||||
{
|
||||
"add_bos_token": false,
|
||||
"add_prefix_space": false,
|
||||
"added_tokens_decoder": {
|
||||
"151643": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151644": {
|
||||
"content": "<|im_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151645": {
|
||||
"content": "<|im_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151646": {
|
||||
"content": "<|object_ref_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151647": {
|
||||
"content": "<|object_ref_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151648": {
|
||||
"content": "<|box_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151649": {
|
||||
"content": "<|box_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151650": {
|
||||
"content": "<|quad_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151651": {
|
||||
"content": "<|quad_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151652": {
|
||||
"content": "<|vision_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151653": {
|
||||
"content": "<|vision_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151654": {
|
||||
"content": "<|vision_pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151655": {
|
||||
"content": "<|image_pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151656": {
|
||||
"content": "<|video_pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"151657": {
|
||||
"content": "<tool_call>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151658": {
|
||||
"content": "</tool_call>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151659": {
|
||||
"content": "<|fim_prefix|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151660": {
|
||||
"content": "<|fim_middle|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151661": {
|
||||
"content": "<|fim_suffix|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151662": {
|
||||
"content": "<|fim_pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151663": {
|
||||
"content": "<|repo_name|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151664": {
|
||||
"content": "<|file_sep|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151665": {
|
||||
"content": "<tool_response>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151666": {
|
||||
"content": "</tool_response>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151667": {
|
||||
"content": "<think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
},
|
||||
"151668": {
|
||||
"content": "</think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": false
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>",
|
||||
"<|object_ref_start|>",
|
||||
"<|object_ref_end|>",
|
||||
"<|box_start|>",
|
||||
"<|box_end|>",
|
||||
"<|quad_start|>",
|
||||
"<|quad_end|>",
|
||||
"<|vision_start|>",
|
||||
"<|vision_end|>",
|
||||
"<|vision_pad|>",
|
||||
"<|image_pad|>",
|
||||
"<|video_pad|>"
|
||||
],
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": {},
|
||||
"model_max_length": 131072,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"padding_side": "left",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null
|
||||
}
|
||||
1564
trainer_state.json
Normal file
1564
trainer_state.json
Normal file
File diff suppressed because it is too large
Load Diff
3
training_args.bin
Normal file
3
training_args.bin
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:9e9886e4e07ab43f0706e4efefa54a94c4176dc568cbcd9903e797347697a364
|
||||
size 6865
|
||||
1
vocab.json
Normal file
1
vocab.json
Normal file
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user