初始化项目,由ModelHub XC社区提供模型

Model: DavidWen2025/Qwen3-Embedding-0.6B-GPTQ-Int8
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-11 07:01:13 +08:00
commit 358041fdbc
16 changed files with 152490 additions and 0 deletions

50
.gitattributes vendored Normal file
View File

@@ -0,0 +1,50 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
model.safetensors filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text

349
README.md Normal file
View File

@@ -0,0 +1,349 @@
---
license: apache-2.0
base_model:
- Qwen/Qwen3-Embedding-0.6B
tags:
- transformers
- sentence-transformers
- sentence-similarity
- feature-extraction
- text-embeddings-inference
base_model_relation: quantized
tasks:
- text-generation
frameworks: PyTorch
---
# Qwen3-Embedding-0.6B-GPTQ-Int8
## Introduction
**Produced by AXERA-TECH(From Huggingface)**
https://huggingface.co/AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8
| Model Weight | Quant_method |
| ------------ | -------------------------------- |
| 742 MiB | GPTQ (vllm gptq_marlin support) |
> Notice: GPTQ_Marlin is not supported on NVIDIA GPU architectures older than Ampere (such as Turing or Volta).
## vllm example
Hardware: RTX 4090D 24G
vllm version: 0.11.2
```bash
vllm serve "/home/model/Qwen3-Embedding-0.6B-GPTQ-Int8" \
--host 0.0.0.0 \
--port 3001 \
--gpu-memory-utilization 0.1 \
--served-model-name "Qwen3-Embedding-0.6B-GPTQ-Int8" \
--quantization "gptq_marlin" \
--dtype "float16" \
--max-model-len 3000 \
--task embed
```
**Double the effective KV cache capacity by adding the code below.**
```bash
--kv-cache-dtype fp8
```
> Notice: Compatible with NVIDIA Ampere GPUs or later (e.g., RTX 30 Series, Ada Lovelace, Hopper, and Blackwell).
<p align="center">
<img src="https://qianwen-res.oss-accelerate-overseas.aliyuncs.com/logo_qwen3.png" width="400"/>
<p>
## Highlights
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.
**Exceptional Versatility**: The embedding model has achieved state-of-the-art performance across a wide range of downstream application evaluations. The 8B size embedding model ranks **No.1** in the MTEB multilingual leaderboard (as of June 5, 2025, score **70.58**), while the reranking model excels in various text retrieval scenarios.
**Comprehensive Flexibility**: The Qwen3 Embedding series offers a full spectrum of sizes (from 0.6B to 8B) for both embedding and reranking models, catering to diverse use cases that prioritize efficiency and effectiveness. Developers can seamlessly combine these two modules. Additionally, the embedding model allows for flexible vector definitions across all dimensions, and both embedding and reranking models support user-defined instructions to enhance performance for specific tasks, languages, or scenarios.
**Multilingual Capability**: The Qwen3 Embedding series offer support for over 100 languages, thanks to the multilingual capabilites of Qwen3 models. This includes various programming languages, and provides robust multilingual, cross-lingual, and code retrieval capabilities.
## Model Overview
**Qwen3-Embedding-0.6B-GPTQ-Int8** has the following features:
- Model Type: Text Embedding
- Supported Languages: 100+ Languages
- Number of Paramaters: 0.6B
- Context Length: 32k
- Embedding Dimension: Up to 1024, supports user-defined output dimensions ranging from 32 to 1024
For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our [blog](https://qwenlm.github.io/blog/qwen3-embedding/), [GitHub](https://github.com/QwenLM/Qwen3-Embedding).
## Qwen3 Embedding Series Model list
| Model Type | Models | Size | Layers | Sequence Length | Embedding Dimension | MRL Support | Instruction Aware |
|------------------|----------------------|------|--------|-----------------|---------------------|-------------|----------------|
| Text Embedding | [Qwen3-Embedding-0.6B-GPTQ-Int8](https://huggingface.co/AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8) | 0.6B | 28 | 32K | 1024 | Yes | Yes |
| Text Embedding | [Qwen3-Embedding-4B](https://huggingface.co/Qwen/Qwen3-Embedding-4B) | 4B | 36 | 32K | 2560 | Yes | Yes |
| Text Embedding | [Qwen3-Embedding-8B](https://huggingface.co/Qwen/Qwen3-Embedding-8B) | 8B | 36 | 32K | 4096 | Yes | Yes |
| Text Reranking | [Qwen3-Reranker-0.6B](https://huggingface.co/Qwen/Qwen3-Reranker-0.6B) | 0.6B | 28 | 32K | - | - | Yes |
| Text Reranking | [Qwen3-Reranker-4B](https://huggingface.co/Qwen/Qwen3-Reranker-4B) | 4B | 36 | 32K | - | - | Yes |
| Text Reranking | [Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) | 8B | 36 | 32K | - | - | Yes |
> **Note**:
> - `MRL Support` indicates whether the embedding model supports custom dimensions for the final embedding.
> - `Instruction Aware` notes whether the embedding or reranking model supports customizing the input instruction according to different tasks.
> - Our evaluation indicates that, for most downstream tasks, using instructions (instruct) typically yields an improvement of 1% to 5% compared to not using them. Therefore, we recommend that developers create tailored instructions specific to their tasks and scenarios. In multilingual contexts, we also advise users to write their instructions in English, as most instructions utilized during the model training process were originally written in English.
## Usage
With Transformers versions earlier than 4.51.0, you may encounter the following error:
```
KeyError: 'qwen3'
```
### Sentence Transformers Usage
```python
# Requires transformers>=4.51.0
# Requires sentence-transformers>=2.7.0
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer("AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8")
# We recommend enabling flash_attention_2 for better acceleration and memory saving,
# together with setting `padding_side` to "left":
# model = SentenceTransformer(
# "AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8",
# model_kwargs={"attn_implementation": "flash_attention_2", "device_map": "auto"},
# tokenizer_kwargs={"padding_side": "left"},
# )
# The queries and documents to embed
queries = [
"What is the capital of China?",
"Explain gravity",
]
documents = [
"The capital of China is Beijing.",
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.",
]
# Encode the queries and documents. Note that queries benefit from using a prompt
# Here we use the prompt called "query" stored under `model.prompts`, but you can
# also pass your own prompt via the `prompt` argument
query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents)
# Compute the (cosine) similarity between the query and document embeddings
similarity = model.similarity(query_embeddings, document_embeddings)
print(similarity)
# tensor([[0.7646, 0.1414],
# [0.1355, 0.6000]])
```
### Transformers Usage
```python
# Requires transformers>=4.51.0
import torch
import torch.nn.functional as F
from torch import Tensor
from transformers import AutoTokenizer, AutoModel
def last_token_pool(last_hidden_states: Tensor,
attention_mask: Tensor) -> Tensor:
left_padding = (attention_mask[:, -1].sum() == attention_mask.shape[0])
if left_padding:
return last_hidden_states[:, -1]
else:
sequence_lengths = attention_mask.sum(dim=1) - 1
batch_size = last_hidden_states.shape[0]
return last_hidden_states[torch.arange(batch_size, device=last_hidden_states.device), sequence_lengths]
def get_detailed_instruct(task_description: str, query: str) -> str:
return f'Instruct: {task_description}\nQuery:{query}'
# Each query must come with a one-sentence instruction that describes the task
task = 'Given a web search query, retrieve relevant passages that answer the query'
queries = [
get_detailed_instruct(task, 'What is the capital of China?'),
get_detailed_instruct(task, 'Explain gravity')
]
# No need to add instruction for retrieval documents
documents = [
"The capital of China is Beijing.",
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun."
]
input_texts = queries + documents
tokenizer = AutoTokenizer.from_pretrained('AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8', padding_side='left')
model = AutoModel.from_pretrained('AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8')
# We recommend enabling flash_attention_2 for better acceleration and memory saving.
# model = AutoModel.from_pretrained('AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8', attn_implementation="flash_attention_2", torch_dtype=torch.float16).cuda()
max_length = 8192
# Tokenize the input texts
batch_dict = tokenizer(
input_texts,
padding=True,
truncation=True,
max_length=max_length,
return_tensors="pt",
)
batch_dict.to(model.device)
outputs = model(**batch_dict)
embeddings = last_token_pool(outputs.last_hidden_state, batch_dict['attention_mask'])
# normalize embeddings
embeddings = F.normalize(embeddings, p=2, dim=1)
scores = (embeddings[:2] @ embeddings[2:].T)
print(scores.tolist())
# [[0.7645568251609802, 0.14142508804798126], [0.13549736142158508, 0.5999549627304077]] # 非 GPTQ
# [[0.76171875, 0.1409912109375], [0.1334228515625, 0.59765625]] # GPTQ-Int8
```
### vLLM Usage
```python
# Requires vllm>=0.8.5
import torch
import vllm
from vllm import LLM
def get_detailed_instruct(task_description: str, query: str) -> str:
return f'Instruct: {task_description}\nQuery:{query}'
# Each query must come with a one-sentence instruction that describes the task
task = 'Given a web search query, retrieve relevant passages that answer the query'
queries = [
get_detailed_instruct(task, 'What is the capital of China?'),
get_detailed_instruct(task, 'Explain gravity')
]
# No need to add instruction for retrieval documents
documents = [
"The capital of China is Beijing.",
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun."
]
input_texts = queries + documents
model = LLM(model="AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8", task="embed")
outputs = model.embed(input_texts)
embeddings = torch.tensor([o.outputs.embedding for o in outputs])
scores = (embeddings[:2] @ embeddings[2:].T)
print(scores.tolist())
# [[0.7620252966880798, 0.14078938961029053], [0.1358368694782257, 0.6013815999031067]]
```
📌 **Tip**: We recommend that developers customize the `instruct` according to their specific scenarios, tasks, and languages. Our tests have shown that in most retrieval scenarios, not using an `instruct` on the query side can lead to a drop in retrieval performance by approximately 1% to 5%.
### Text Embeddings Inference (TEI) Usage
You can either run / deploy TEI on NVIDIA GPUs as:
```bash
docker run --gpus all -p 8080:80 -v hf_cache:/data --pull always ghcr.io/huggingface/text-embeddings-inference:cpu-1.7.2 --model-id AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8 --dtype float16
```
Or on CPU devices as:
```bash
docker run -p 8080:80 -v hf_cache:/data --pull always ghcr.io/huggingface/text-embeddings-inference:1.7.2 --model-id AXERA-TECH/Qwen3-Embedding-0.6B-GPTQ-Int8
```
And then, generate the embeddings sending a HTTP POST request as:
```bash
curl http://localhost:8080/embed \
-X POST \
-d '{"inputs": ["Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: What is the capital of China?", "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: Explain gravity"]}' \
-H "Content-Type: application/json"
```
## Evaluation
### MTEB (Multilingual)
| Model | Size | Mean (Task) | Mean (Type) | Bitxt Mining | Class. | Clust. | Inst. Retri. | Multi. Class. | Pair. Class. | Rerank | Retri. | STS |
|----------------------------------|:-------:|:-------------:|:-------------:|:--------------:|:--------:|:--------:|:--------------:|:---------------:|:--------------:|:--------:|:--------:|:------:|
| NV-Embed-v2 | 7B | 56.29 | 49.58 | 57.84 | 57.29 | 40.80 | 1.04 | 18.63 | 78.94 | 63.82 | 56.72 | 71.10|
| GritLM-7B | 7B | 60.92 | 53.74 | 70.53 | 61.83 | 49.75 | 3.45 | 22.77 | 79.94 | 63.78 | 58.31 | 73.33|
| BGE-M3 | 0.6B | 59.56 | 52.18 | 79.11 | 60.35 | 40.88 | -3.11 | 20.1 | 80.76 | 62.79 | 54.60 | 74.12|
| multilingual-e5-large-instruct | 0.6B | 63.22 | 55.08 | 80.13 | 64.94 | 50.75 | -0.40 | 22.91 | 80.86 | 62.61 | 57.12 | 76.81|
| gte-Qwen2-1.5B-instruct | 1.5B | 59.45 | 52.69 | 62.51 | 58.32 | 52.05 | 0.74 | 24.02 | 81.58 | 62.58 | 60.78 | 71.61|
| gte-Qwen2-7b-Instruct | 7B | 62.51 | 55.93 | 73.92 | 61.55 | 52.77 | 4.94 | 25.48 | 85.13 | 65.55 | 60.08 | 73.98|
| text-embedding-3-large | - | 58.93 | 51.41 | 62.17 | 60.27 | 46.89 | -2.68 | 22.03 | 79.17 | 63.89 | 59.27 | 71.68|
| Cohere-embed-multilingual-v3.0 | - | 61.12 | 53.23 | 70.50 | 62.95 | 46.89 | -1.89 | 22.74 | 79.88 | 64.07 | 59.16 | 74.80|
| Gemini Embedding | - | 68.37 | 59.59 | 79.28 | 71.82 | 54.59 | 5.18 | **29.16** | 83.63 | 65.58 | 67.71 | 79.40|
| **Qwen3-Embedding-0.6B-GPTQ-Int8** | 0.6B | 64.33 | 56.00 | 72.22 | 66.83 | 52.33 | 5.09 | 24.59 | 80.83 | 61.41 | 64.64 | 76.17|
| **Qwen3-Embedding-4B** | 4B | 69.45 | 60.86 | 79.36 | 72.33 | 57.15 | **11.56** | 26.77 | 85.05 | 65.08 | 69.60 | 80.86|
| **Qwen3-Embedding-8B** | 8B | **70.58** | **61.69** | **80.89** | **74.00** | **57.65** | 10.06 | 28.66 | **86.40** | **65.63** | **70.88** | **81.08** |
> **Note**: For compared models, the scores are retrieved from MTEB online [leaderboard](https://huggingface.co/spaces/mteb/leaderboard) on May 24th, 2025.
### MTEB (Eng v2)
| MTEB English / Models | Param. | Mean(Task) | Mean(Type) | Class. | Clust. | Pair Class. | Rerank. | Retri. | STS | Summ. |
|--------------------------------|:--------:|:------------:|:------------:|:--------:|:--------:|:-------------:|:---------:|:--------:|:-------:|:-------:|
| multilingual-e5-large-instruct | 0.6B | 65.53 | 61.21 | 75.54 | 49.89 | 86.24 | 48.74 | 53.47 | 84.72 | 29.89 |
| NV-Embed-v2 | 7.8B | 69.81 | 65.00 | 87.19 | 47.66 | 88.69 | 49.61 | 62.84 | 83.82 | 35.21 |
| GritLM-7B | 7.2B | 67.07 | 63.22 | 81.25 | 50.82 | 87.29 | 49.59 | 54.95 | 83.03 | 35.65 |
| gte-Qwen2-1.5B-instruct | 1.5B | 67.20 | 63.26 | 85.84 | 53.54 | 87.52 | 49.25 | 50.25 | 82.51 | 33.94 |
| stella_en_1.5B_v5 | 1.5B | 69.43 | 65.32 | 89.38 | 57.06 | 88.02 | 50.19 | 52.42 | 83.27 | 36.91 |
| gte-Qwen2-7B-instruct | 7.6B | 70.72 | 65.77 | 88.52 | 58.97 | 85.9 | 50.47 | 58.09 | 82.69 | 35.74 |
| gemini-embedding-exp-03-07 | - | 73.3 | 67.67 | 90.05 | 59.39 | 87.7 | 48.59 | 64.35 | 85.29 | 38.28 |
| **Qwen3-Embedding-0.6B-GPTQ-Int8** | 0.6B | 70.70 | 64.88 | 85.76 | 54.05 | 84.37 | 48.18 | 61.83 | 86.57 | 33.43 |
| **Qwen3-Embedding-4B** | 4B | 74.60 | 68.10 | 89.84 | 57.51 | 87.01 | 50.76 | 68.46 | 88.72 | 34.39 |
| **Qwen3-Embedding-8B** | 8B | 75.22 | 68.71 | 90.43 | 58.57 | 87.52 | 51.56 | 69.44 | 88.58 | 34.83 |
### C-MTEB (MTEB Chinese)
| C-MTEB | Param. | Mean(Task) | Mean(Type) | Class. | Clust. | Pair Class. | Rerank. | Retr. | STS |
|------------------|--------|------------|------------|--------|--------|-------------|---------|-------|-------|
| multilingual-e5-large-instruct | 0.6B | 58.08 | 58.24 | 69.80 | 48.23 | 64.52 | 57.45 | 63.65 | 45.81 |
| bge-multilingual-gemma2 | 9B | 67.64 | 75.31 | 59.30 | 86.67 | 68.28 | 73.73 | 55.19 | - |
| gte-Qwen2-1.5B-instruct | 1.5B | 67.12 | 67.79 | 72.53 | 54.61 | 79.5 | 68.21 | 71.86 | 60.05 |
| gte-Qwen2-7B-instruct | 7.6B | 71.62 | 72.19 | 75.77 | 66.06 | 81.16 | 69.24 | 75.70 | 65.20 |
| ritrieve_zh_v1 | 0.3B | 72.71 | 73.85 | 76.88 | 66.5 | 85.98 | 72.86 | 76.97 | 63.92 |
| **Qwen3-Embedding-0.6B-GPTQ-Int8** | 0.6B | 66.33 | 67.45 | 71.40 | 68.74 | 76.42 | 62.58 | 71.03 | 54.52 |
| **Qwen3-Embedding-4B** | 4B | 72.27 | 73.51 | 75.46 | 77.89 | 83.34 | 66.05 | 77.03 | 61.26 |
| **Qwen3-Embedding-8B** | 8B | 73.84 | 75.00 | 76.97 | 80.08 | 84.23 | 66.99 | 78.21 | 63.53 |
## Citation
If you find our work helpful, feel free to give us a cite.
```
@article{qwen3embedding,
title={Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models},
author={Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren},
journal={arXiv preprint arXiv:2506.05176},
year={2025}
}
```

28
added_tokens.json Normal file
View File

@@ -0,0 +1,28 @@
{
"</think>": 151668,
"</tool_call>": 151658,
"</tool_response>": 151666,
"<think>": 151667,
"<tool_call>": 151657,
"<tool_response>": 151665,
"<|box_end|>": 151649,
"<|box_start|>": 151648,
"<|endoftext|>": 151643,
"<|file_sep|>": 151664,
"<|fim_middle|>": 151660,
"<|fim_pad|>": 151662,
"<|fim_prefix|>": 151659,
"<|fim_suffix|>": 151661,
"<|im_end|>": 151645,
"<|im_start|>": 151644,
"<|image_pad|>": 151655,
"<|object_ref_end|>": 151647,
"<|object_ref_start|>": 151646,
"<|quad_end|>": 151651,
"<|quad_start|>": 151650,
"<|repo_name|>": 151663,
"<|video_pad|>": 151656,
"<|vision_end|>": 151653,
"<|vision_pad|>": 151654,
"<|vision_start|>": 151652
}

85
chat_template.jinja Normal file
View File

@@ -0,0 +1,85 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0].role == 'system' %}
{{- messages[0].content + '\n\n' }}
{%- endif %}
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0].role == 'system' %}
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endfor %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{%- set content = message.content %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
{%- set reasoning_content = message.reasoning_content %}
{%- else %}
{%- if '</think>' in message.content %}
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
{%- endif %}
{%- endif %}
{%- if loop.index0 > ns.last_query_index %}
{%- if loop.last or (not loop.last and reasoning_content) %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- if message.tool_calls %}
{%- for tool_call in message.tool_calls %}
{%- if (loop.first and content) or (not loop.first) %}
{{- '\n' }}
{%- endif %}
{%- if tool_call.function %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{%- if tool_call.arguments is string %}
{{- tool_call.arguments }}
{%- else %}
{{- tool_call.arguments | tojson }}
{%- endif %}
{{- '}\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if enable_thinking is defined and enable_thinking is false %}
{{- '<think>\n\n</think>\n\n' }}
{%- endif %}
{%- endif %}

84
config.json Normal file
View File

@@ -0,0 +1,84 @@
{
"architectures": [
"Qwen3ForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "bfloat16",
"eos_token_id": 151643,
"head_dim": 128,
"hidden_act": "silu",
"hidden_size": 1024,
"initializer_range": 0.02,
"intermediate_size": 3072,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 28,
"model_type": "qwen3",
"num_attention_heads": 16,
"num_hidden_layers": 28,
"num_key_value_heads": 8,
"quantization_config": {
"bits": 8,
"checkpoint_format": "gptq",
"desc_act": false,
"group_size": 128,
"lm_head": false,
"meta": {
"act_group_aware": false,
"damp_auto_increment": 0.01,
"damp_percent": 0.01,
"mse": 0.0,
"quantizer": [
"gptqmodel:5.0.0-dev0"
],
"static_groups": false,
"true_sequential": true,
"uri": "https://github.com/modelcloud/gptqmodel",
"v2": false,
"v2_alpha": 0.25
},
"pack_dtype": "int32",
"quant_method": "gptq",
"sym": true
},
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000,
"sliding_window": null,
"tie_word_embeddings": true,
"transformers_version": "4.56.2",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 151669
}

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework":"Pytorch","task":"text-generation"}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token_id": 151643,
"eos_token_id": 151643,
"max_new_tokens": 2048,
"transformers_version": "4.56.2"
}

151388
merges.txt Normal file

File diff suppressed because it is too large Load Diff

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ef6e7ba596d1d4ff7aa0e119eae4c602c1ef81017dcb12febce66cfb2e54708
size 762718928

6
quant_eval_results.json Normal file
View File

@@ -0,0 +1,6 @@
{
"cosine_mean": 1.0,
"cosine_std": 0.0005459785461425781,
"cosine_min": 0.99853515625,
"cosine_max": 1.0009765625
}

197
quant_log.csv Normal file
View File

@@ -0,0 +1,197 @@
layer,module,loss,samples,damp,time
0,self_attn.q_proj,0.0000000157,0.01000,0.453
0,self_attn.k_proj,0.0000000070,0.01000,0.218
0,self_attn.v_proj,0.0000000055,0.01000,0.220
0,self_attn.o_proj,0.0000000039,0.01000,0.433
0,mlp.gate_proj,0.0000001457,0.01000,0.216
0,mlp.up_proj,0.0000000745,0.01000,0.215
0,mlp.down_proj,0.0000000050,0.01000,0.652
1,self_attn.q_proj,0.0000000056,0.01000,0.286
1,self_attn.k_proj,0.0000000025,0.01000,0.217
1,self_attn.v_proj,0.0000000024,0.01000,0.217
1,self_attn.o_proj,0.0000000010,0.01000,0.434
1,mlp.gate_proj,0.0000008918,0.01000,0.216
1,mlp.up_proj,0.0000002667,0.01000,0.216
1,mlp.down_proj,0.0000000066,0.01000,0.659
2,self_attn.q_proj,0.0000000110,0.01000,0.255
2,self_attn.k_proj,0.0000000047,0.01000,0.219
2,self_attn.v_proj,0.0000000046,0.01000,0.218
2,self_attn.o_proj,0.0000000016,0.01000,0.435
2,mlp.gate_proj,0.0000006115,0.01000,0.219
2,mlp.up_proj,0.0000002231,0.01000,0.218
2,mlp.down_proj,0.0000078325,0.01000,0.661
3,self_attn.q_proj,0.0000000868,0.01000,0.260
3,self_attn.k_proj,0.0000000421,0.01000,0.220
3,self_attn.v_proj,0.0000000429,0.01000,0.219
3,self_attn.o_proj,0.0000000023,0.01000,0.442
3,mlp.gate_proj,0.0000008147,0.01000,0.218
3,mlp.up_proj,0.0000003186,0.01000,0.218
3,mlp.down_proj,0.0000000175,0.01000,0.701
4,self_attn.q_proj,0.0000000812,0.01000,0.263
4,self_attn.k_proj,0.0000000383,0.01000,0.217
4,self_attn.v_proj,0.0000000410,0.01000,0.219
4,self_attn.o_proj,0.0000000053,0.01000,0.437
4,mlp.gate_proj,0.0000007532,0.01000,0.219
4,mlp.up_proj,0.0000003315,0.01000,0.218
4,mlp.down_proj,0.0000000219,0.01000,0.659
5,self_attn.q_proj,0.0000001455,0.01000,0.262
5,self_attn.k_proj,0.0000000597,0.01000,0.223
5,self_attn.v_proj,0.0000000636,0.01000,0.226
5,self_attn.o_proj,0.0000000076,0.01000,0.452
5,mlp.gate_proj,0.0000005187,0.01000,0.223
5,mlp.up_proj,0.0000003102,0.01000,0.218
5,mlp.down_proj,0.0000000246,0.01000,0.664
6,self_attn.q_proj,0.0000000990,0.01000,0.257
6,self_attn.k_proj,0.0000000439,0.01000,0.222
6,self_attn.v_proj,0.0000000422,0.01000,0.218
6,self_attn.o_proj,0.0000000059,0.01000,0.438
6,mlp.gate_proj,0.0000006165,0.01000,0.216
6,mlp.up_proj,0.0000003886,0.01000,0.216
6,mlp.down_proj,0.0000000316,0.01000,0.661
7,self_attn.q_proj,0.0000001937,0.01000,0.282
7,self_attn.k_proj,0.0000000798,0.01000,0.249
7,self_attn.v_proj,0.0000000900,0.01000,0.246
7,self_attn.o_proj,0.0000000101,0.01000,0.490
7,mlp.gate_proj,0.0000007165,0.01000,0.243
7,mlp.up_proj,0.0000004434,0.01000,0.243
7,mlp.down_proj,0.0000000395,0.01000,0.700
8,self_attn.q_proj,0.0000002418,0.01000,0.257
8,self_attn.k_proj,0.0000001092,0.01000,0.221
8,self_attn.v_proj,0.0000001022,0.01000,0.217
8,self_attn.o_proj,0.0000000103,0.01000,0.435
8,mlp.gate_proj,0.0000007137,0.01000,0.218
8,mlp.up_proj,0.0000004578,0.01000,0.216
8,mlp.down_proj,0.0000000417,0.01000,0.659
9,self_attn.q_proj,0.0000004561,0.01000,0.291
9,self_attn.k_proj,0.0000001857,0.01000,0.219
9,self_attn.v_proj,0.0000001929,0.01000,0.221
9,self_attn.o_proj,0.0000000167,0.01000,0.439
9,mlp.gate_proj,0.0000007954,0.01000,0.220
9,mlp.up_proj,0.0000004992,0.01000,0.218
9,mlp.down_proj,0.0000000556,0.01000,0.661
10,self_attn.q_proj,0.0000003840,0.01000,0.289
10,self_attn.k_proj,0.0000001600,0.01000,0.220
10,self_attn.v_proj,0.0000001638,0.01000,0.220
10,self_attn.o_proj,0.0000000157,0.01000,0.446
10,mlp.gate_proj,0.0000008065,0.01000,0.221
10,mlp.up_proj,0.0000005084,0.01000,0.222
10,mlp.down_proj,0.0000000788,0.01000,0.671
11,self_attn.q_proj,0.0000008024,0.01000,0.283
11,self_attn.k_proj,0.0000003071,0.01000,0.239
11,self_attn.v_proj,0.0000002804,0.01000,0.220
11,self_attn.o_proj,0.0000000493,0.01000,0.443
11,mlp.gate_proj,0.0000006262,0.01000,0.220
11,mlp.up_proj,0.0000004808,0.01000,0.220
11,mlp.down_proj,0.0000000880,0.01000,0.668
12,self_attn.q_proj,0.0000007135,0.01000,0.266
12,self_attn.k_proj,0.0000002560,0.01000,0.218
12,self_attn.v_proj,0.0000002706,0.01000,0.215
12,self_attn.o_proj,0.0000000163,0.01000,0.443
12,mlp.gate_proj,0.0000005636,0.01000,0.215
12,mlp.up_proj,0.0000004659,0.01000,0.216
12,mlp.down_proj,0.0000000890,0.01000,0.659
13,self_attn.q_proj,0.0000007450,0.01000,0.255
13,self_attn.k_proj,0.0000002548,0.01000,0.215
13,self_attn.v_proj,0.0000002986,0.01000,0.215
13,self_attn.o_proj,0.0000000196,0.01000,0.432
13,mlp.gate_proj,0.0000006345,0.01000,0.214
13,mlp.up_proj,0.0000005136,0.01000,0.218
13,mlp.down_proj,0.0000000943,0.01000,0.656
14,self_attn.q_proj,0.0000010099,0.01000,0.258
14,self_attn.k_proj,0.0000003666,0.01000,0.215
14,self_attn.v_proj,0.0000003883,0.01000,0.215
14,self_attn.o_proj,0.0000000272,0.01000,0.435
14,mlp.gate_proj,0.0000006731,0.01000,0.215
14,mlp.up_proj,0.0000005531,0.01000,0.217
14,mlp.down_proj,0.0000001202,0.01000,0.653
15,self_attn.q_proj,0.0000019066,0.01000,0.286
15,self_attn.k_proj,0.0000005995,0.01000,0.218
15,self_attn.v_proj,0.0000007671,0.01000,0.215
15,self_attn.o_proj,0.0000000275,0.01000,0.434
15,mlp.gate_proj,0.0000007228,0.01000,0.215
15,mlp.up_proj,0.0000006042,0.01000,0.215
15,mlp.down_proj,0.0000001389,0.01000,0.656
16,self_attn.q_proj,0.0000023607,0.01000,0.288
16,self_attn.k_proj,0.0000008254,0.01000,0.216
16,self_attn.v_proj,0.0000007666,0.01000,0.215
16,self_attn.o_proj,0.0000000499,0.01000,0.436
16,mlp.gate_proj,0.0000007419,0.01000,0.213
16,mlp.up_proj,0.0000006750,0.01000,0.215
16,mlp.down_proj,0.0000002760,0.01000,0.653
17,self_attn.q_proj,0.0000051417,0.01000,0.286
17,self_attn.k_proj,0.0000016262,0.01000,0.216
17,self_attn.v_proj,0.0000019666,0.01000,0.215
17,self_attn.o_proj,0.0000001240,0.01000,0.433
17,mlp.gate_proj,0.0000010348,0.01000,0.216
17,mlp.up_proj,0.0000009428,0.01000,0.213
17,mlp.down_proj,0.0000003055,0.01000,0.656
18,self_attn.q_proj,0.0000046798,0.01000,0.266
18,self_attn.k_proj,0.0000014684,0.01000,0.216
18,self_attn.v_proj,0.0000017435,0.01000,0.216
18,self_attn.o_proj,0.0000000528,0.01000,0.432
18,mlp.gate_proj,0.0000011800,0.01000,0.214
18,mlp.up_proj,0.0000010920,0.01000,0.213
18,mlp.down_proj,0.0000005320,0.01000,0.651
19,self_attn.q_proj,0.0000081104,0.01000,0.262
19,self_attn.k_proj,0.0000023922,0.01000,0.216
19,self_attn.v_proj,0.0000029877,0.01000,0.215
19,self_attn.o_proj,0.0000001101,0.01000,0.433
19,mlp.gate_proj,0.0000012617,0.01000,0.214
19,mlp.up_proj,0.0000013705,0.01000,0.214
19,mlp.down_proj,0.0000010956,0.01000,0.654
20,self_attn.q_proj,0.0000104670,0.01000,0.278
20,self_attn.k_proj,0.0000034022,0.01000,0.243
20,self_attn.v_proj,0.0000042421,0.01000,0.243
20,self_attn.o_proj,0.0000001907,0.01000,0.493
20,mlp.gate_proj,0.0000013812,0.01000,0.242
20,mlp.up_proj,0.0000015767,0.01000,0.242
20,mlp.down_proj,0.0000016353,0.01000,0.748
21,self_attn.q_proj,0.0000185475,0.01000,0.276
21,self_attn.k_proj,0.0000059908,0.01000,0.242
21,self_attn.v_proj,0.0000075683,0.01000,0.242
21,self_attn.o_proj,0.0000003463,0.01000,0.487
21,mlp.gate_proj,0.0000014603,0.01000,0.239
21,mlp.up_proj,0.0000018684,0.01000,0.239
21,mlp.down_proj,0.0000023298,0.01000,0.744
22,self_attn.q_proj,0.0000192655,0.01000,0.280
22,self_attn.k_proj,0.0000066621,0.01000,0.248
22,self_attn.v_proj,0.0000090980,0.01000,0.250
22,self_attn.o_proj,0.0000002642,0.01000,0.503
22,mlp.gate_proj,0.0000016697,0.01000,0.243
22,mlp.up_proj,0.0000021728,0.01000,0.240
22,mlp.down_proj,0.0000026377,0.01000,0.750
23,self_attn.q_proj,0.0000223355,0.01000,0.282
23,self_attn.k_proj,0.0000091246,0.01000,0.251
23,self_attn.v_proj,0.0000111566,0.01000,0.250
23,self_attn.o_proj,0.0000002412,0.01000,0.487
23,mlp.gate_proj,0.0000018961,0.01000,0.243
23,mlp.up_proj,0.0000025225,0.01000,0.239
23,mlp.down_proj,0.0000023748,0.01000,0.733
24,self_attn.q_proj,0.0000478594,0.01000,0.277
24,self_attn.k_proj,0.0000162195,0.01000,0.244
24,self_attn.v_proj,0.0000187105,0.01000,0.244
24,self_attn.o_proj,0.0000003140,0.01000,0.490
24,mlp.gate_proj,0.0000018491,0.01000,0.242
24,mlp.up_proj,0.0000025646,0.01000,0.241
24,mlp.down_proj,0.0000022875,0.01000,0.743
25,self_attn.q_proj,0.0000736474,0.01000,0.261
25,self_attn.k_proj,0.0000219495,0.01000,0.218
25,self_attn.v_proj,0.0000348672,0.01000,0.220
25,self_attn.o_proj,0.0000005398,0.01000,0.441
25,mlp.gate_proj,0.0000018458,0.01000,0.217
25,mlp.up_proj,0.0000026985,0.01000,0.217
25,mlp.down_proj,0.0000027734,0.01000,0.647
26,self_attn.q_proj,0.0000973767,0.01000,0.251
26,self_attn.k_proj,0.0000243580,0.01000,0.241
26,self_attn.v_proj,0.0000364561,0.01000,0.243
26,self_attn.o_proj,0.0000044389,0.01000,0.485
26,mlp.gate_proj,0.0000018377,0.01000,0.240
26,mlp.up_proj,0.0000026512,0.01000,0.246
26,mlp.down_proj,0.0000062773,0.01000,0.744
27,self_attn.q_proj,0.0000334389,0.01000,0.290
27,self_attn.k_proj,0.0000149378,0.01000,0.243
27,self_attn.v_proj,0.0000184760,0.01000,0.243
27,self_attn.o_proj,0.0000022290,0.01000,0.489
27,mlp.gate_proj,0.0000068754,0.01000,0.242
27,mlp.up_proj,0.0000081715,0.01000,0.240
27,mlp.down_proj,0.0000067090,0.01000,0.735
1 layer,module,loss,samples,damp,time
2 0,self_attn.q_proj,0.0000000157,0.01000,0.453
3 0,self_attn.k_proj,0.0000000070,0.01000,0.218
4 0,self_attn.v_proj,0.0000000055,0.01000,0.220
5 0,self_attn.o_proj,0.0000000039,0.01000,0.433
6 0,mlp.gate_proj,0.0000001457,0.01000,0.216
7 0,mlp.up_proj,0.0000000745,0.01000,0.215
8 0,mlp.down_proj,0.0000000050,0.01000,0.652
9 1,self_attn.q_proj,0.0000000056,0.01000,0.286
10 1,self_attn.k_proj,0.0000000025,0.01000,0.217
11 1,self_attn.v_proj,0.0000000024,0.01000,0.217
12 1,self_attn.o_proj,0.0000000010,0.01000,0.434
13 1,mlp.gate_proj,0.0000008918,0.01000,0.216
14 1,mlp.up_proj,0.0000002667,0.01000,0.216
15 1,mlp.down_proj,0.0000000066,0.01000,0.659
16 2,self_attn.q_proj,0.0000000110,0.01000,0.255
17 2,self_attn.k_proj,0.0000000047,0.01000,0.219
18 2,self_attn.v_proj,0.0000000046,0.01000,0.218
19 2,self_attn.o_proj,0.0000000016,0.01000,0.435
20 2,mlp.gate_proj,0.0000006115,0.01000,0.219
21 2,mlp.up_proj,0.0000002231,0.01000,0.218
22 2,mlp.down_proj,0.0000078325,0.01000,0.661
23 3,self_attn.q_proj,0.0000000868,0.01000,0.260
24 3,self_attn.k_proj,0.0000000421,0.01000,0.220
25 3,self_attn.v_proj,0.0000000429,0.01000,0.219
26 3,self_attn.o_proj,0.0000000023,0.01000,0.442
27 3,mlp.gate_proj,0.0000008147,0.01000,0.218
28 3,mlp.up_proj,0.0000003186,0.01000,0.218
29 3,mlp.down_proj,0.0000000175,0.01000,0.701
30 4,self_attn.q_proj,0.0000000812,0.01000,0.263
31 4,self_attn.k_proj,0.0000000383,0.01000,0.217
32 4,self_attn.v_proj,0.0000000410,0.01000,0.219
33 4,self_attn.o_proj,0.0000000053,0.01000,0.437
34 4,mlp.gate_proj,0.0000007532,0.01000,0.219
35 4,mlp.up_proj,0.0000003315,0.01000,0.218
36 4,mlp.down_proj,0.0000000219,0.01000,0.659
37 5,self_attn.q_proj,0.0000001455,0.01000,0.262
38 5,self_attn.k_proj,0.0000000597,0.01000,0.223
39 5,self_attn.v_proj,0.0000000636,0.01000,0.226
40 5,self_attn.o_proj,0.0000000076,0.01000,0.452
41 5,mlp.gate_proj,0.0000005187,0.01000,0.223
42 5,mlp.up_proj,0.0000003102,0.01000,0.218
43 5,mlp.down_proj,0.0000000246,0.01000,0.664
44 6,self_attn.q_proj,0.0000000990,0.01000,0.257
45 6,self_attn.k_proj,0.0000000439,0.01000,0.222
46 6,self_attn.v_proj,0.0000000422,0.01000,0.218
47 6,self_attn.o_proj,0.0000000059,0.01000,0.438
48 6,mlp.gate_proj,0.0000006165,0.01000,0.216
49 6,mlp.up_proj,0.0000003886,0.01000,0.216
50 6,mlp.down_proj,0.0000000316,0.01000,0.661
51 7,self_attn.q_proj,0.0000001937,0.01000,0.282
52 7,self_attn.k_proj,0.0000000798,0.01000,0.249
53 7,self_attn.v_proj,0.0000000900,0.01000,0.246
54 7,self_attn.o_proj,0.0000000101,0.01000,0.490
55 7,mlp.gate_proj,0.0000007165,0.01000,0.243
56 7,mlp.up_proj,0.0000004434,0.01000,0.243
57 7,mlp.down_proj,0.0000000395,0.01000,0.700
58 8,self_attn.q_proj,0.0000002418,0.01000,0.257
59 8,self_attn.k_proj,0.0000001092,0.01000,0.221
60 8,self_attn.v_proj,0.0000001022,0.01000,0.217
61 8,self_attn.o_proj,0.0000000103,0.01000,0.435
62 8,mlp.gate_proj,0.0000007137,0.01000,0.218
63 8,mlp.up_proj,0.0000004578,0.01000,0.216
64 8,mlp.down_proj,0.0000000417,0.01000,0.659
65 9,self_attn.q_proj,0.0000004561,0.01000,0.291
66 9,self_attn.k_proj,0.0000001857,0.01000,0.219
67 9,self_attn.v_proj,0.0000001929,0.01000,0.221
68 9,self_attn.o_proj,0.0000000167,0.01000,0.439
69 9,mlp.gate_proj,0.0000007954,0.01000,0.220
70 9,mlp.up_proj,0.0000004992,0.01000,0.218
71 9,mlp.down_proj,0.0000000556,0.01000,0.661
72 10,self_attn.q_proj,0.0000003840,0.01000,0.289
73 10,self_attn.k_proj,0.0000001600,0.01000,0.220
74 10,self_attn.v_proj,0.0000001638,0.01000,0.220
75 10,self_attn.o_proj,0.0000000157,0.01000,0.446
76 10,mlp.gate_proj,0.0000008065,0.01000,0.221
77 10,mlp.up_proj,0.0000005084,0.01000,0.222
78 10,mlp.down_proj,0.0000000788,0.01000,0.671
79 11,self_attn.q_proj,0.0000008024,0.01000,0.283
80 11,self_attn.k_proj,0.0000003071,0.01000,0.239
81 11,self_attn.v_proj,0.0000002804,0.01000,0.220
82 11,self_attn.o_proj,0.0000000493,0.01000,0.443
83 11,mlp.gate_proj,0.0000006262,0.01000,0.220
84 11,mlp.up_proj,0.0000004808,0.01000,0.220
85 11,mlp.down_proj,0.0000000880,0.01000,0.668
86 12,self_attn.q_proj,0.0000007135,0.01000,0.266
87 12,self_attn.k_proj,0.0000002560,0.01000,0.218
88 12,self_attn.v_proj,0.0000002706,0.01000,0.215
89 12,self_attn.o_proj,0.0000000163,0.01000,0.443
90 12,mlp.gate_proj,0.0000005636,0.01000,0.215
91 12,mlp.up_proj,0.0000004659,0.01000,0.216
92 12,mlp.down_proj,0.0000000890,0.01000,0.659
93 13,self_attn.q_proj,0.0000007450,0.01000,0.255
94 13,self_attn.k_proj,0.0000002548,0.01000,0.215
95 13,self_attn.v_proj,0.0000002986,0.01000,0.215
96 13,self_attn.o_proj,0.0000000196,0.01000,0.432
97 13,mlp.gate_proj,0.0000006345,0.01000,0.214
98 13,mlp.up_proj,0.0000005136,0.01000,0.218
99 13,mlp.down_proj,0.0000000943,0.01000,0.656
100 14,self_attn.q_proj,0.0000010099,0.01000,0.258
101 14,self_attn.k_proj,0.0000003666,0.01000,0.215
102 14,self_attn.v_proj,0.0000003883,0.01000,0.215
103 14,self_attn.o_proj,0.0000000272,0.01000,0.435
104 14,mlp.gate_proj,0.0000006731,0.01000,0.215
105 14,mlp.up_proj,0.0000005531,0.01000,0.217
106 14,mlp.down_proj,0.0000001202,0.01000,0.653
107 15,self_attn.q_proj,0.0000019066,0.01000,0.286
108 15,self_attn.k_proj,0.0000005995,0.01000,0.218
109 15,self_attn.v_proj,0.0000007671,0.01000,0.215
110 15,self_attn.o_proj,0.0000000275,0.01000,0.434
111 15,mlp.gate_proj,0.0000007228,0.01000,0.215
112 15,mlp.up_proj,0.0000006042,0.01000,0.215
113 15,mlp.down_proj,0.0000001389,0.01000,0.656
114 16,self_attn.q_proj,0.0000023607,0.01000,0.288
115 16,self_attn.k_proj,0.0000008254,0.01000,0.216
116 16,self_attn.v_proj,0.0000007666,0.01000,0.215
117 16,self_attn.o_proj,0.0000000499,0.01000,0.436
118 16,mlp.gate_proj,0.0000007419,0.01000,0.213
119 16,mlp.up_proj,0.0000006750,0.01000,0.215
120 16,mlp.down_proj,0.0000002760,0.01000,0.653
121 17,self_attn.q_proj,0.0000051417,0.01000,0.286
122 17,self_attn.k_proj,0.0000016262,0.01000,0.216
123 17,self_attn.v_proj,0.0000019666,0.01000,0.215
124 17,self_attn.o_proj,0.0000001240,0.01000,0.433
125 17,mlp.gate_proj,0.0000010348,0.01000,0.216
126 17,mlp.up_proj,0.0000009428,0.01000,0.213
127 17,mlp.down_proj,0.0000003055,0.01000,0.656
128 18,self_attn.q_proj,0.0000046798,0.01000,0.266
129 18,self_attn.k_proj,0.0000014684,0.01000,0.216
130 18,self_attn.v_proj,0.0000017435,0.01000,0.216
131 18,self_attn.o_proj,0.0000000528,0.01000,0.432
132 18,mlp.gate_proj,0.0000011800,0.01000,0.214
133 18,mlp.up_proj,0.0000010920,0.01000,0.213
134 18,mlp.down_proj,0.0000005320,0.01000,0.651
135 19,self_attn.q_proj,0.0000081104,0.01000,0.262
136 19,self_attn.k_proj,0.0000023922,0.01000,0.216
137 19,self_attn.v_proj,0.0000029877,0.01000,0.215
138 19,self_attn.o_proj,0.0000001101,0.01000,0.433
139 19,mlp.gate_proj,0.0000012617,0.01000,0.214
140 19,mlp.up_proj,0.0000013705,0.01000,0.214
141 19,mlp.down_proj,0.0000010956,0.01000,0.654
142 20,self_attn.q_proj,0.0000104670,0.01000,0.278
143 20,self_attn.k_proj,0.0000034022,0.01000,0.243
144 20,self_attn.v_proj,0.0000042421,0.01000,0.243
145 20,self_attn.o_proj,0.0000001907,0.01000,0.493
146 20,mlp.gate_proj,0.0000013812,0.01000,0.242
147 20,mlp.up_proj,0.0000015767,0.01000,0.242
148 20,mlp.down_proj,0.0000016353,0.01000,0.748
149 21,self_attn.q_proj,0.0000185475,0.01000,0.276
150 21,self_attn.k_proj,0.0000059908,0.01000,0.242
151 21,self_attn.v_proj,0.0000075683,0.01000,0.242
152 21,self_attn.o_proj,0.0000003463,0.01000,0.487
153 21,mlp.gate_proj,0.0000014603,0.01000,0.239
154 21,mlp.up_proj,0.0000018684,0.01000,0.239
155 21,mlp.down_proj,0.0000023298,0.01000,0.744
156 22,self_attn.q_proj,0.0000192655,0.01000,0.280
157 22,self_attn.k_proj,0.0000066621,0.01000,0.248
158 22,self_attn.v_proj,0.0000090980,0.01000,0.250
159 22,self_attn.o_proj,0.0000002642,0.01000,0.503
160 22,mlp.gate_proj,0.0000016697,0.01000,0.243
161 22,mlp.up_proj,0.0000021728,0.01000,0.240
162 22,mlp.down_proj,0.0000026377,0.01000,0.750
163 23,self_attn.q_proj,0.0000223355,0.01000,0.282
164 23,self_attn.k_proj,0.0000091246,0.01000,0.251
165 23,self_attn.v_proj,0.0000111566,0.01000,0.250
166 23,self_attn.o_proj,0.0000002412,0.01000,0.487
167 23,mlp.gate_proj,0.0000018961,0.01000,0.243
168 23,mlp.up_proj,0.0000025225,0.01000,0.239
169 23,mlp.down_proj,0.0000023748,0.01000,0.733
170 24,self_attn.q_proj,0.0000478594,0.01000,0.277
171 24,self_attn.k_proj,0.0000162195,0.01000,0.244
172 24,self_attn.v_proj,0.0000187105,0.01000,0.244
173 24,self_attn.o_proj,0.0000003140,0.01000,0.490
174 24,mlp.gate_proj,0.0000018491,0.01000,0.242
175 24,mlp.up_proj,0.0000025646,0.01000,0.241
176 24,mlp.down_proj,0.0000022875,0.01000,0.743
177 25,self_attn.q_proj,0.0000736474,0.01000,0.261
178 25,self_attn.k_proj,0.0000219495,0.01000,0.218
179 25,self_attn.v_proj,0.0000348672,0.01000,0.220
180 25,self_attn.o_proj,0.0000005398,0.01000,0.441
181 25,mlp.gate_proj,0.0000018458,0.01000,0.217
182 25,mlp.up_proj,0.0000026985,0.01000,0.217
183 25,mlp.down_proj,0.0000027734,0.01000,0.647
184 26,self_attn.q_proj,0.0000973767,0.01000,0.251
185 26,self_attn.k_proj,0.0000243580,0.01000,0.241
186 26,self_attn.v_proj,0.0000364561,0.01000,0.243
187 26,self_attn.o_proj,0.0000044389,0.01000,0.485
188 26,mlp.gate_proj,0.0000018377,0.01000,0.240
189 26,mlp.up_proj,0.0000026512,0.01000,0.246
190 26,mlp.down_proj,0.0000062773,0.01000,0.744
191 27,self_attn.q_proj,0.0000334389,0.01000,0.290
192 27,self_attn.k_proj,0.0000149378,0.01000,0.243
193 27,self_attn.v_proj,0.0000184760,0.01000,0.243
194 27,self_attn.o_proj,0.0000022290,0.01000,0.489
195 27,mlp.gate_proj,0.0000068754,0.01000,0.242
196 27,mlp.up_proj,0.0000081715,0.01000,0.240
197 27,mlp.down_proj,0.0000067090,0.01000,0.735

24
quantize_config.json Normal file
View File

@@ -0,0 +1,24 @@
{
"bits": 8,
"group_size": 128,
"desc_act": false,
"sym": true,
"lm_head": false,
"quant_method": "gptq",
"checkpoint_format": "gptq",
"pack_dtype": "int32",
"meta": {
"quantizer": [
"gptqmodel:5.0.0-dev0"
],
"uri": "https://github.com/modelcloud/gptqmodel",
"damp_percent": 0.01,
"damp_auto_increment": 0.01,
"static_groups": false,
"true_sequential": true,
"mse": 0.0,
"v2": false,
"v2_alpha": 0.25,
"act_group_aware": false
}
}

25
special_tokens_map.json Normal file
View File

@@ -0,0 +1,25 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"eos_token": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": "<|endoftext|>"
}

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:def76fb086971c7867b829c23a26261e38d9d74e02139253b38aeb9df8b4b50a
size 11423705

240
tokenizer_config.json Normal file
View File

@@ -0,0 +1,240 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151665": {
"content": "<tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151666": {
"content": "</tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151667": {
"content": "<think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151668": {
"content": "</think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2TokenizerFast",
"unk_token": null,
"_commit_hash": null
}

1
vocab.json Normal file

File diff suppressed because one or more lines are too long