初始化项目,由ModelHub XC社区提供模型

Model: robertspumiaca1975/Qwen2.5-Coder-14B-n8n-Workflow-Generator
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-12 18:46:20 +08:00
commit fca4cc8081
46 changed files with 458160 additions and 0 deletions

45
.gitattributes vendored Normal file
View File

@@ -0,0 +1,45 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
checkpoint-200/tokenizer.json filter=lfs diff=lfs merge=lfs -text
checkpoint-300/tokenizer.json filter=lfs diff=lfs merge=lfs -text
checkpoint-400/tokenizer.json filter=lfs diff=lfs merge=lfs -text
checkpoint-432/tokenizer.json filter=lfs diff=lfs merge=lfs -text
img-assets/header.jpg filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text
gguf/qwen25-coder-14b-n8n-f16.gguf filter=lfs diff=lfs merge=lfs -text
gguf/qwen25-coder-14b-n8n-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
merged-hf/tokenizer.json filter=lfs diff=lfs merge=lfs -text
mlx-q4/tokenizer.json filter=lfs diff=lfs merge=lfs -text

125
README.md Normal file
View File

@@ -0,0 +1,125 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-Coder-14B-Instruct
tags:
- n8n
- workflow
- automation
- fine-tuned
- code-generation
- qlora
datasets:
- mbakgun/n8nbuilder-n8n-workflows-dataset
pipeline_tag: text-generation
language:
- en
library_name: mlx
---
# Qwen2.5-Coder-14B-n8n-Workflow-Generator
![n8nbuilder.dev](./img-assets/header.jpg)
Fine-tuned Qwen2.5-Coder-14B-Instruct model specialized for generating n8n workflow JSONs from natural language descriptions.
## Model Description
This model is a QLoRA fine-tuned version of [Qwen/Qwen2.5-Coder-14B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-14B-Instruct) on the [n8nbuilder-n8n-workflows-dataset](https://huggingface.co/datasets/mbakgun/n8nbuilder-n8n-workflows-dataset), containing +2.5K n8n workflow templates.
**Training Details:**
- **Base Model**: Qwen/Qwen2.5-Coder-14B-Instruct
- **Method**: QLoRA (4-bit quantization)
- **LoRA Rank**: 32
- **LoRA Alpha**: 64
- **Training Steps**: 432 (3 epochs)
- **Sequence Length**: 8192 tokens
- **Learning Rate**: 2e-4
## Usage
### Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "mbakgun/Qwen2.5-Coder-14B-n8n-Workflow-Generator"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
system_prompt = "You are an expert n8n workflow generation assistant. Your goal is to create valid, efficient, and functional n8n workflow configurations."
user_input = "Create a workflow that monitors a RSS feed and sends new items to Discord."
prompt = f"{system_prompt}\n\n{user_input}"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=4096,
temperature=0.7,
do_sample=True
)
workflow_json = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(workflow_json)
```
### MLX (Apple Silicon)
```bash
# Download MLX Q4 model
mlx_lm.generate \
--model mbakgun/Qwen2.5-Coder-14B-n8n-Workflow-Generator/mlx-q4 \
--prompt "You are an expert n8n workflow generation assistant...\n\nCreate a workflow that sends Slack notifications when GitHub issues are created." \
--max-tokens 4096
```
## Training Data
This model was fine-tuned on the [n8nbuilder-n8n-workflows-dataset](https://huggingface.co/datasets/mbakgun/n8nbuilder-n8n-workflows-dataset), which contains:
- **+2,304 workflow templates** (after filtering sequences >8192 tokens)
- Format: Alpaca (instruction/input/output)
- Source: n8n.io public template gallery
- [n8nbuilder.dev - Create n8n Workflows in Seconds with AI](https://n8nbuilder.dev)
## Performance
- **Training Speed**: ~33.85s/step on H100 PCIe
- **VRAM Usage**: ~30GB (4-bit QLoRA)
- **Inference**: ~25-40 tok/s on Mac Mini M4 64GB (MLX)
## Limitations
- Generated workflows may require manual validation
- Long workflows (>8192 tokens) may be truncated
- Model trained on public templates only
## Citation
```bibtex
@model{qwen25_coder_n8n_2025,
title={Qwen2.5-Coder-14B-n8n-Workflow-Generator},
author={mbakgun},
year={2025},
base_model={Qwen/Qwen2.5-Coder-14B-Instruct},
dataset={mbakgun/n8nbuilder-n8n-workflows-dataset},
url={https://huggingface.co/mbakgun/Qwen2.5-Coder-14B-n8n-Workflow-Generator}
}
```
## Acknowledgments
- [Qwen Team](https://huggingface.co/Qwen) for the base model
- [n8n](https://n8n.io) for the workflow automation platform
- [n8n-mcp](https://github.com/czlonkowski/n8n-mcp) for template indexing
## License
Apache 2.0

42
adapter_config.json Normal file
View File

@@ -0,0 +1,42 @@
{
"alpha_pattern": {},
"auto_mapping": null,
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-14B-Instruct",
"bias": "none",
"corda_config": null,
"eva_config": null,
"exclude_modules": null,
"fan_in_fan_out": null,
"inference_mode": true,
"init_lora_weights": true,
"layer_replication": null,
"layers_pattern": null,
"layers_to_transform": null,
"loftq_config": {},
"lora_alpha": 64,
"lora_bias": false,
"lora_dropout": 0.05,
"megatron_config": null,
"megatron_core": "megatron.core",
"modules_to_save": null,
"peft_type": "LORA",
"qalora_group_size": 16,
"r": 32,
"rank_pattern": {},
"revision": null,
"target_modules": [
"up_proj",
"q_proj",
"down_proj",
"gate_proj",
"v_proj",
"o_proj",
"k_proj"
],
"target_parameters": [],
"task_type": "CAUSAL_LM",
"trainable_token_indices": null,
"use_dora": false,
"use_qalora": false,
"use_rslora": false
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8b1d9d4ffcee0be992300b65795d78890d05f2efab9415fbf6c3b0d766246125
size 550593184

24
added_tokens.json Normal file
View File

@@ -0,0 +1,24 @@
{
"</tool_call>": 151658,
"<tool_call>": 151657,
"<|box_end|>": 151649,
"<|box_start|>": 151648,
"<|endoftext|>": 151643,
"<|file_sep|>": 151664,
"<|fim_middle|>": 151660,
"<|fim_pad|>": 151662,
"<|fim_prefix|>": 151659,
"<|fim_suffix|>": 151661,
"<|im_end|>": 151645,
"<|im_start|>": 151644,
"<|image_pad|>": 151655,
"<|object_ref_end|>": 151647,
"<|object_ref_start|>": 151646,
"<|quad_end|>": 151651,
"<|quad_start|>": 151650,
"<|repo_name|>": 151663,
"<|video_pad|>": 151656,
"<|vision_end|>": 151653,
"<|vision_pad|>": 151654,
"<|vision_start|>": 151652
}

54
chat_template.jinja Normal file
View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

92
config.json Normal file
View File

@@ -0,0 +1,92 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"dtype": "bfloat16",
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 48,
"model_type": "qwen2",
"num_attention_heads": 40,
"num_hidden_layers": 48,
"num_key_value_heads": 8,
"quantization_config": {
"_load_in_4bit": true,
"_load_in_8bit": false,
"bnb_4bit_compute_dtype": "bfloat16",
"bnb_4bit_quant_storage": "bfloat16",
"bnb_4bit_quant_type": "nf4",
"bnb_4bit_use_double_quant": true,
"llm_int8_enable_fp32_cpu_offload": false,
"llm_int8_has_fp16_weight": false,
"llm_int8_skip_modules": null,
"llm_int8_threshold": 6.0,
"load_in_4bit": true,
"load_in_8bit": false,
"quant_method": "bitsandbytes"
},
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000.0,
"sliding_window": null,
"tie_word_embeddings": false,
"transformers_version": "4.57.0",
"use_cache": false,
"use_sliding_window": false,
"vocab_size": 152064
}

613
debug.log Normal file
View File

@@ -0,0 +1,613 @@
[2025-12-25 22:27:59,815] [DEBUG] [axolotl.utils.config.log_gpu_memory_usage:127] [PID:1133] baseline 0.000GB ()
[2025-12-25 22:27:59,815] [INFO] [axolotl.cli.config.load_cfg:248] [PID:1133] config:
{
"activation_offloading": false,
"adapter": "qlora",
"axolotl_config_path": "config.yaml",
"base_model": "Qwen/Qwen2.5-Coder-14B-Instruct",
"base_model_config": "Qwen/Qwen2.5-Coder-14B-Instruct",
"batch_size": 16,
"bf16": true,
"capabilities": {
"bf16": true,
"compute_capability": "sm_90",
"fp8": false,
"n_gpu": 1,
"n_node": 1
},
"context_parallel_size": 1,
"dataloader_num_workers": 1,
"dataloader_pin_memory": true,
"dataloader_prefetch_factor": 256,
"dataset_processes": 36,
"datasets": [
{
"message_property_mappings": {
"content": "content",
"role": "role"
},
"path": "mbakgun/n8nbuilder-n8n-workflows-dataset",
"trust_remote_code": false,
"type": "alpaca"
}
],
"ddp": false,
"device": "cuda:0",
"dion_rank_fraction": 1.0,
"dion_rank_multiple_of": 1,
"env_capabilities": {
"torch_version": "2.7.1"
},
"eval_batch_size": 1,
"eval_causal_lm_metrics": [
"sacrebleu",
"comet",
"ter",
"chrf"
],
"eval_max_new_tokens": 128,
"eval_table_size": 0,
"experimental_skip_move_to_device": true,
"flash_attention": true,
"fp16": false,
"gradient_accumulation_steps": 16,
"gradient_checkpointing": true,
"gradient_checkpointing_kwargs": {
"use_reentrant": false
},
"include_tkps": true,
"learning_rate": 0.0002,
"lisa_layers_attribute": "model.layers",
"load_best_model_at_end": false,
"load_in_4bit": true,
"load_in_8bit": false,
"local_rank": 0,
"logging_steps": 1,
"lora_alpha": 64,
"lora_dropout": 0.05,
"lora_r": 32,
"lora_target_modules": [
"q_proj",
"k_proj",
"v_proj",
"o_proj",
"gate_proj",
"up_proj",
"down_proj"
],
"loraplus_lr_embedding": 1e-06,
"lr_scheduler": "cosine",
"mean_resizing_embeddings": false,
"micro_batch_size": 1,
"model_config_type": "qwen2",
"num_epochs": 3.0,
"optimizer": "adamw_bnb_8bit",
"output_dir": "./outputs/qwen25-coder-n8n",
"pad_to_sequence_len": false,
"pretrain_multipack_attn": true,
"profiler_steps_start": 0,
"qlora_sharded_model_loading": false,
"ray_num_workers": 1,
"resources_per_worker": {
"GPU": 1
},
"sample_packing": false,
"sample_packing_bin_size": 200,
"sample_packing_group_size": 100000,
"save_only_model": false,
"save_safetensors": true,
"save_steps": 100,
"save_strategy": "steps",
"sequence_len": 8192,
"shuffle_before_merging_datasets": false,
"shuffle_merged_datasets": true,
"skip_prepare_dataset": false,
"streaming_multipack_buffer_size": 10000,
"strict": false,
"tensor_parallel_size": 1,
"tf32": true,
"tiled_mlp_use_original_mlp": true,
"tokenizer_config": "Qwen/Qwen2.5-Coder-14B-Instruct",
"tokenizer_save_jinja_files": true,
"torch_dtype": "torch.bfloat16",
"train_on_inputs": false,
"trl": {
"log_completions": false,
"mask_truncated_completions": false,
"ref_model_mixup_alpha": 0.9,
"ref_model_sync_steps": 64,
"scale_rewards": true,
"sync_ref_model": false,
"use_vllm": false,
"vllm_server_host": "0.0.0.0",
"vllm_server_port": 8000
},
"use_ray": false,
"val_set_size": 0.0,
"vllm": {
"device": "auto",
"dtype": "auto",
"gpu_memory_utilization": 0.9,
"host": "0.0.0.0",
"port": 8000
},
"warmup_ratio": 0.1,
"weight_decay": 0.01,
"world_size": 1
}
[2025-12-25 22:28:00,313] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:278] [PID:1133] EOS: 151645 / <|im_end|>
[2025-12-25 22:28:00,314] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:279] [PID:1133] BOS: None / None
[2025-12-25 22:28:00,314] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:280] [PID:1133] PAD: 151643 / <|endoftext|>
[2025-12-25 22:28:00,314] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:281] [PID:1133] UNK: None / None
[2025-12-25 22:28:00,314] [INFO] [axolotl.utils.data.shared.load_preprocessed_dataset:476] [PID:1133] Unable to find prepared dataset in last_run_prepared/fd30d23b351de719c91e124efcc5fe43
[2025-12-25 22:28:00,314] [INFO] [axolotl.utils.data.sft._load_raw_datasets:320] [PID:1133] Loading raw datasets...
[2025-12-25 22:28:00,314] [WARNING] [axolotl.utils.data.sft._load_raw_datasets:322] [PID:1133] Processing datasets during training can lead to VRAM instability. Please pre-process your dataset using `axolotl preprocess path/to/config.yml`.
[2025-12-25 22:28:00,955] [INFO] [axolotl.utils.data.wrappers.get_dataset_wrapper:87] [PID:1133] Loading dataset: mbakgun/n8nbuilder-n8n-workflows-dataset with base_type: alpaca and prompt_style: None
[2025-12-25 22:28:01,168] [INFO] [axolotl.utils.data.utils.handle_long_seq_in_dataset:218] [PID:1133] min_input_len: 878
[2025-12-25 22:28:01,168] [INFO] [axolotl.utils.data.utils.handle_long_seq_in_dataset:220] [PID:1133] max_input_len: 12396
Dropping Long Sequences (>8192) (num_proc=36): 0%| | 0/2737 [00:00<?, ? examples/s]
Dropping Long Sequences (>8192) (num_proc=36): 3%|██▎ | 77/2737 [00:00<00:31, 85.60 examples/s]
Dropping Long Sequences (>8192) (num_proc=36): 42%|████████████████████████████████▉ | 1141/2737 [00:00<00:01, 1531.24 examples/s]
Dropping Long Sequences (>8192) (num_proc=36): 100%|███████████████████████████████████████████████████████████████████████████████| 2737/2737 [00:01<00:00, 3888.53 examples/s]
Dropping Long Sequences (>8192) (num_proc=36): 100%|███████████████████████████████████████████████████████████████████████████████| 2737/2737 [00:01<00:00, 2115.72 examples/s]
[2025-12-25 22:28:02,540] [WARNING] [axolotl.utils.data.utils.handle_long_seq_in_dataset:260] [PID:1133] Dropped 433 samples from dataset
Saving the dataset (0/9 shards): 0%| | 0/2304 [00:00<?, ? examples/s]
Saving the dataset (0/9 shards): 11%|██████████▌ | 256/2304 [00:00<00:02, 916.49 examples/s]
Saving the dataset (1/9 shards): 11%|██████████▌ | 256/2304 [00:00<00:02, 916.49 examples/s]
Saving the dataset (2/9 shards): 33%|███████████████████████████████▋ | 768/2304 [00:00<00:01, 916.49 examples/s]
Saving the dataset (3/9 shards): 33%|███████████████████████████████▋ | 768/2304 [00:00<00:01, 916.49 examples/s]
Saving the dataset (4/9 shards): 44%|█████████████████████████████████████████▊ | 1024/2304 [00:00<00:01, 916.49 examples/s]
Saving the dataset (5/9 shards): 56%|████████████████████████████████████████████████████▏ | 1280/2304 [00:00<00:01, 916.49 examples/s]
Saving the dataset (6/9 shards): 67%|██████████████████████████████████████████████████████████████▋ | 1536/2304 [00:00<00:00, 916.49 examples/s]
Saving the dataset (7/9 shards): 78%|█████████████████████████████████████████████████████████████████████████ | 1792/2304 [00:00<00:00, 916.49 examples/s]
Saving the dataset (8/9 shards): 89%|███████████████████████████████████████████████████████████████████████████████████▌ | 2048/2304 [00:00<00:00, 916.49 examples/s]
Saving the dataset (9/9 shards): 100%|██████████████████████████████████████████████████████████████████████████████████████████████| 2304/2304 [00:00<00:00, 916.49 examples/s]
Saving the dataset (9/9 shards): 100%|█████████████████████████████████████████████████████████████████████████████████████████████| 2304/2304 [00:00<00:00, 5976.92 examples/s]
[2025-12-25 22:28:03,194] [DEBUG] [axolotl.utils.trainer.calculate_total_num_steps:404] [PID:1133] total_num_tokens: 9_507_792
[2025-12-25 22:28:03,238] [DEBUG] [axolotl.utils.trainer.calculate_total_num_steps:422] [PID:1133] `total_supervised_tokens: 11_572_652`
[2025-12-25 22:28:03,239] [DEBUG] [axolotl.utils.trainer.calculate_total_num_steps:520] [PID:1133] total_num_steps: 432
[2025-12-25 22:28:03,239] [INFO] [axolotl.utils.data.sft._prepare_standard_dataset:121] [PID:1133] Maximum number of steps set at 432
[2025-12-25 22:28:03,267] [DEBUG] [axolotl.train.setup_model_and_tokenizer:65] [PID:1133] Loading tokenizer... Qwen/Qwen2.5-Coder-14B-Instruct
[2025-12-25 22:28:03,684] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:278] [PID:1133] EOS: 151645 / <|im_end|>
[2025-12-25 22:28:03,684] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:279] [PID:1133] BOS: None / None
[2025-12-25 22:28:03,685] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:280] [PID:1133] PAD: 151643 / <|endoftext|>
[2025-12-25 22:28:03,685] [DEBUG] [axolotl.loaders.tokenizer.load_tokenizer:281] [PID:1133] UNK: None / None
[2025-12-25 22:28:03,685] [DEBUG] [axolotl.train.setup_model_and_tokenizer:74] [PID:1133] Loading model
[2025-12-25 22:28:03,736] [DEBUG] [axolotl.monkeypatch.transformers.trainer_loss_calc.patch_evaluation_loop:87] [PID:1133] Patched Trainer.evaluation_loop with nanmean loss calculation
[2025-12-25 22:28:03,737] [DEBUG] [axolotl.monkeypatch.transformers.trainer_loss_calc.patch_maybe_log_save_evaluate:138] [PID:1133] Patched Trainer._maybe_log_save_evaluate with nanmean loss calculation
Loading checkpoint shards: 0%| | 0/6 [00:00<?, ?it/s]
Loading checkpoint shards: 17%|███████████████████ | 1/6 [00:03<00:19, 3.93s/it]
Loading checkpoint shards: 33%|██████████████████████████████████████ | 2/6 [00:09<00:18, 4.62s/it]
Loading checkpoint shards: 50%|█████████████████████████████████████████████████████████ | 3/6 [00:14<00:14, 4.80s/it]
Loading checkpoint shards: 67%|████████████████████████████████████████████████████████████████████████████ | 4/6 [00:19<00:09, 4.89s/it]
Loading checkpoint shards: 83%|███████████████████████████████████████████████████████████████████████████████████████████████ | 5/6 [00:24<00:04, 4.91s/it]
Loading checkpoint shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:27<00:00, 4.42s/it]
Loading checkpoint shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:27<00:00, 4.58s/it]
[2025-12-25 22:28:32,092] [INFO] [axolotl.loaders.model._prepare_model_for_quantization:863] [PID:1133] converting PEFT model w/ prepare_model_for_kbit_training
[2025-12-25 22:28:32,098] [INFO] [axolotl.loaders.model._configure_embedding_dtypes:345] [PID:1133] Converting modules to torch.bfloat16
[2025-12-25 22:28:32,100] [DEBUG] [axolotl.loaders.model.log_gpu_memory_usage:127] [PID:1133] Memory usage after model load 13.888GB (+13.888GB allocated, +15.756GB reserved)
trainable params: 137,625,600 || all params: 14,907,659,264 || trainable%: 0.9232
[2025-12-25 22:28:33,795] [DEBUG] [axolotl.loaders.model.log_gpu_memory_usage:127] [PID:1133] after adapters 9.900GB (+9.900GB allocated, +16.031GB reserved)
[2025-12-25 22:28:40,857] [INFO] [axolotl.train.save_initial_configs:398] [PID:1133] Pre-saving adapter config to ./outputs/qwen25-coder-n8n...
[2025-12-25 22:28:40,857] [INFO] [axolotl.train.save_initial_configs:402] [PID:1133] Pre-saving tokenizer to ./outputs/qwen25-coder-n8n...
[2025-12-25 22:28:41,011] [INFO] [axolotl.train.save_initial_configs:407] [PID:1133] Pre-saving model config to ./outputs/qwen25-coder-n8n...
[2025-12-25 22:28:41,013] [INFO] [axolotl.train.execute_training:196] [PID:1133] Starting trainer...
0%| | 0/432 [00:00<?, ?it/s]
0%|▎ | 1/432 [00:33<4:03:10, 33.85s/it]
{'loss': 1.0821, 'grad_norm': 0.09848134219646454, 'learning_rate': 0.0, 'memory/max_active (GiB)': 27.69, 'memory/max_allocated (GiB)': 27.69, 'memory/device_reserved (GiB)': 30.8, 'tokens_per_second_per_gpu': 1763.68, 'epoch': 0.01}
0%|▎ | 1/432 [00:33<4:03:10, 33.85s/it]
0%|▋ | 2/432 [01:03<3:44:28, 31.32s/it]
{'loss': 1.2119, 'grad_norm': 0.11234692484140396, 'learning_rate': 4.651162790697674e-06, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 31.05, 'tokens_per_second_per_gpu': 1834.64, 'epoch': 0.01}
0%|▋ | 2/432 [01:03<3:44:28, 31.32s/it]
1%|▉ | 3/432 [01:33<3:40:04, 30.78s/it]
{'loss': 1.2053, 'grad_norm': 0.11071926355361938, 'learning_rate': 9.302325581395349e-06, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 31.05, 'tokens_per_second_per_gpu': 1875.82, 'epoch': 0.02}
1%|▉ | 3/432 [01:33<3:40:04, 30.78s/it]
1%|█▎ | 4/432 [02:07<3:49:36, 32.19s/it]
{'loss': 1.0514, 'grad_norm': 0.10147764533758163, 'learning_rate': 1.3953488372093024e-05, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.27, 'tokens_per_second_per_gpu': 1792.72, 'epoch': 0.03}
1%|█▎ | 4/432 [02:07<3:49:36, 32.19s/it]
1%|█▌ | 5/432 [02:41<3:51:32, 32.54s/it]
{'loss': 1.209, 'grad_norm': 0.10568977892398834, 'learning_rate': 1.8604651162790697e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.27, 'tokens_per_second_per_gpu': 1840.61, 'epoch': 0.03}
1%|█▌ | 5/432 [02:41<3:51:32, 32.54s/it]
1%|█▉ | 6/432 [03:13<3:51:42, 32.63s/it]
{'loss': 1.0817, 'grad_norm': 0.10363873094320297, 'learning_rate': 2.3255813953488374e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.27, 'tokens_per_second_per_gpu': 1784.92, 'epoch': 0.04}
1%|█▉ | 6/432 [03:13<3:51:42, 32.63s/it]
2%|██▏ | 7/432 [03:52<4:04:08, 34.47s/it]
{'loss': 1.1571, 'grad_norm': 0.113986074924469, 'learning_rate': 2.7906976744186048e-05, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1892.63, 'epoch': 0.05}
2%|██▏ | 7/432 [03:52<4:04:08, 34.47s/it]
2%|██▌ | 8/432 [04:26<4:02:41, 34.34s/it]
{'loss': 1.1444, 'grad_norm': 0.1191892921924591, 'learning_rate': 3.2558139534883724e-05, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1863.93, 'epoch': 0.06}
2%|██▌ | 8/432 [04:26<4:02:41, 34.34s/it]
2%|██▊ | 9/432 [04:54<3:49:22, 32.53s/it]
{'loss': 1.1786, 'grad_norm': 0.11628979444503784, 'learning_rate': 3.7209302325581394e-05, 'memory/max_active (GiB)': 24.64, 'memory/max_allocated (GiB)': 24.64, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1764.21, 'epoch': 0.06}
2%|██▊ | 9/432 [04:54<3:49:22, 32.53s/it]
2%|███▏ | 10/432 [05:25<3:44:45, 31.96s/it]
{'loss': 1.0695, 'grad_norm': 0.10155434161424637, 'learning_rate': 4.186046511627907e-05, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1823.2, 'epoch': 0.07}
2%|███▏ | 10/432 [05:25<3:44:45, 31.96s/it]
3%|███▍ | 11/432 [06:01<3:52:59, 33.21s/it]
{'loss': 1.0805, 'grad_norm': 0.08485760539770126, 'learning_rate': 4.651162790697675e-05, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1805.0, 'epoch': 0.08}
3%|███▍ | 11/432 [06:01<3:52:59, 33.21s/it]
3%|███▊ | 12/432 [06:30<3:44:23, 32.05s/it]
{'loss': 1.0824, 'grad_norm': 0.07211048156023026, 'learning_rate': 5.1162790697674425e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1799.74, 'epoch': 0.08}
3%|███▊ | 12/432 [06:30<3:44:23, 32.05s/it]
3%|████ | 13/432 [07:04<3:47:30, 32.58s/it]
{'loss': 1.0264, 'grad_norm': 0.06483420729637146, 'learning_rate': 5.5813953488372095e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1798.61, 'epoch': 0.09}
3%|████ | 13/432 [07:04<3:47:30, 32.58s/it]
3%|████▍ | 14/432 [07:32<3:37:27, 31.21s/it]
{'loss': 1.0967, 'grad_norm': 0.06657296419143677, 'learning_rate': 6.0465116279069765e-05, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1788.44, 'epoch': 0.1}
3%|████▍ | 14/432 [07:32<3:37:27, 31.21s/it]
3%|████▋ | 15/432 [07:59<3:26:38, 29.73s/it]
{'loss': 1.1489, 'grad_norm': 0.195042684674263, 'learning_rate': 6.511627906976745e-05, 'memory/max_active (GiB)': 24.8, 'memory/max_allocated (GiB)': 24.8, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1785.21, 'epoch': 0.1}
3%|████▋ | 15/432 [07:59<3:26:38, 29.73s/it]
4%|█████ | 16/432 [08:26<3:21:43, 29.10s/it]
{'loss': 1.0985, 'grad_norm': 0.07728952169418335, 'learning_rate': 6.976744186046513e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1752.91, 'epoch': 0.11}
4%|█████ | 16/432 [08:26<3:21:43, 29.10s/it]
4%|█████▎ | 17/432 [08:55<3:20:08, 28.93s/it]
{'loss': 1.1112, 'grad_norm': 0.08134876191616058, 'learning_rate': 7.441860465116279e-05, 'memory/max_active (GiB)': 26.06, 'memory/max_allocated (GiB)': 26.06, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1813.48, 'epoch': 0.12}
4%|█████▎ | 17/432 [08:55<3:20:08, 28.93s/it]
4%|█████▋ | 18/432 [09:24<3:20:20, 29.04s/it]
{'loss': 1.0222, 'grad_norm': 0.08289807289838791, 'learning_rate': 7.906976744186047e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1852.47, 'epoch': 0.12}
4%|█████▋ | 18/432 [09:24<3:20:20, 29.04s/it]
4%|█████▉ | 19/432 [09:50<3:12:40, 27.99s/it]
{'loss': 1.1493, 'grad_norm': 0.09635733813047409, 'learning_rate': 8.372093023255814e-05, 'memory/max_active (GiB)': 24.17, 'memory/max_allocated (GiB)': 24.17, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1745.06, 'epoch': 0.13}
4%|█████▉ | 19/432 [09:50<3:12:40, 27.99s/it]
5%|██████▎ | 20/432 [10:16<3:09:20, 27.57s/it]
{'loss': 0.9912, 'grad_norm': 0.08602173626422882, 'learning_rate': 8.837209302325582e-05, 'memory/max_active (GiB)': 26.25, 'memory/max_allocated (GiB)': 26.25, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1758.58, 'epoch': 0.14}
5%|██████▎ | 20/432 [10:16<3:09:20, 27.57s/it]
5%|██████▌ | 21/432 [10:45<3:11:39, 27.98s/it]
{'loss': 1.1637, 'grad_norm': 0.08320974558591843, 'learning_rate': 9.30232558139535e-05, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.33, 'tokens_per_second_per_gpu': 1772.58, 'epoch': 0.15}
5%|██████▌ | 21/432 [10:45<3:11:39, 27.98s/it]
5%|██████▉ | 22/432 [11:22<3:29:56, 30.72s/it]
{'loss': 1.0209, 'grad_norm': 0.0785663053393364, 'learning_rate': 9.767441860465116e-05, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1822.66, 'epoch': 0.15}
5%|██████▉ | 22/432 [11:22<3:29:56, 30.72s/it]
5%|███████▏ | 23/432 [11:55<3:33:42, 31.35s/it]
{'loss': 0.9858, 'grad_norm': 0.07734047621488571, 'learning_rate': 0.00010232558139534885, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1889.88, 'epoch': 0.16}
5%|███████▏ | 23/432 [11:55<3:33:42, 31.35s/it]
6%|███████▌ | 24/432 [12:27<3:35:25, 31.68s/it]
{'loss': 1.0003, 'grad_norm': 0.07255646586418152, 'learning_rate': 0.00010697674418604651, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1833.16, 'epoch': 0.17}
6%|███████▌ | 24/432 [12:27<3:35:25, 31.68s/it]
6%|███████▊ | 25/432 [13:04<3:45:06, 33.19s/it]
{'loss': 1.0143, 'grad_norm': 0.07897679507732391, 'learning_rate': 0.00011162790697674419, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1900.98, 'epoch': 0.17}
6%|███████▊ | 25/432 [13:04<3:45:06, 33.19s/it]
6%|████████▏ | 26/432 [13:29<3:28:01, 30.74s/it]
{'loss': 1.0341, 'grad_norm': 0.09510312229394913, 'learning_rate': 0.00011627906976744187, 'memory/max_active (GiB)': 23.18, 'memory/max_allocated (GiB)': 23.18, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1701.34, 'epoch': 0.18}
6%|████████▏ | 26/432 [13:29<3:28:01, 30.74s/it]
6%|████████▌ | 27/432 [14:03<3:33:11, 31.58s/it]
{'loss': 1.0004, 'grad_norm': 0.07016909122467041, 'learning_rate': 0.00012093023255813953, 'memory/max_active (GiB)': 28.38, 'memory/max_allocated (GiB)': 28.38, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1891.5, 'epoch': 0.19}
6%|████████▌ | 27/432 [14:03<3:33:11, 31.58s/it]
6%|████████▊ | 28/432 [14:40<3:45:02, 33.42s/it]
{'loss': 0.9588, 'grad_norm': 0.07151541113853455, 'learning_rate': 0.0001255813953488372, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1890.08, 'epoch': 0.19}
6%|████████▊ | 28/432 [14:40<3:45:02, 33.42s/it]
7%|█████████▏ | 29/432 [15:12<3:39:52, 32.74s/it]
{'loss': 1.0078, 'grad_norm': 0.07155290246009827, 'learning_rate': 0.0001302325581395349, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1759.11, 'epoch': 0.2}
7%|█████████▏ | 29/432 [15:12<3:39:52, 32.74s/it]
7%|█████████▍ | 30/432 [15:39<3:27:49, 31.02s/it]
{'loss': 1.0343, 'grad_norm': 0.08267220109701157, 'learning_rate': 0.00013488372093023256, 'memory/max_active (GiB)': 26.49, 'memory/max_allocated (GiB)': 26.49, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1794.99, 'epoch': 0.21}
7%|█████████▍ | 30/432 [15:39<3:27:49, 31.02s/it]
7%|█████████▊ | 31/432 [16:13<3:34:58, 32.17s/it]
{'loss': 0.9155, 'grad_norm': 0.06379543989896774, 'learning_rate': 0.00013953488372093025, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1829.94, 'epoch': 0.22}
7%|█████████▊ | 31/432 [16:13<3:34:58, 32.17s/it]
7%|██████████ | 32/432 [16:40<3:23:21, 30.50s/it]
{'loss': 1.0727, 'grad_norm': 0.07846751064062119, 'learning_rate': 0.00014418604651162791, 'memory/max_active (GiB)': 23.79, 'memory/max_allocated (GiB)': 23.79, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1854.31, 'epoch': 0.22}
7%|██████████ | 32/432 [16:40<3:23:21, 30.50s/it]
8%|██████████▍ | 33/432 [17:11<3:23:04, 30.54s/it]
{'loss': 1.0472, 'grad_norm': 0.07601239532232285, 'learning_rate': 0.00014883720930232558, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1806.02, 'epoch': 0.23}
8%|██████████▍ | 33/432 [17:11<3:23:04, 30.54s/it]
8%|██████████▋ | 34/432 [17:38<3:15:19, 29.45s/it]
{'loss': 1.034, 'grad_norm': 0.09074926376342773, 'learning_rate': 0.00015348837209302327, 'memory/max_active (GiB)': 23.79, 'memory/max_allocated (GiB)': 23.79, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1814.81, 'epoch': 0.24}
8%|██████████▋ | 34/432 [17:38<3:15:19, 29.45s/it]
8%|███████████ | 35/432 [18:09<3:19:04, 30.09s/it]
{'loss': 0.9786, 'grad_norm': 0.07441543787717819, 'learning_rate': 0.00015813953488372093, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1865.18, 'epoch': 0.24}
8%|███████████ | 35/432 [18:09<3:19:04, 30.09s/it]
8%|███████████▎ | 36/432 [18:43<3:26:48, 31.33s/it]
{'loss': 0.9218, 'grad_norm': 0.08436308056116104, 'learning_rate': 0.00016279069767441862, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1860.2, 'epoch': 0.25}
8%|███████████▎ | 36/432 [18:43<3:26:48, 31.33s/it]
9%|███████████▋ | 37/432 [19:13<3:22:58, 30.83s/it]
{'loss': 1.0504, 'grad_norm': 0.07554468512535095, 'learning_rate': 0.00016744186046511629, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1816.31, 'epoch': 0.26}
9%|███████████▋ | 37/432 [19:13<3:22:58, 30.83s/it]
9%|███████████▉ | 38/432 [19:44<3:22:43, 30.87s/it]
{'loss': 0.9141, 'grad_norm': 0.09911656379699707, 'learning_rate': 0.00017209302325581395, 'memory/max_active (GiB)': 26.26, 'memory/max_allocated (GiB)': 26.26, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1863.8, 'epoch': 0.26}
9%|███████████▉ | 38/432 [19:44<3:22:43, 30.87s/it]
9%|████████████▎ | 39/432 [20:06<3:03:55, 28.08s/it]
{'loss': 1.0342, 'grad_norm': 0.07778877764940262, 'learning_rate': 0.00017674418604651164, 'memory/max_active (GiB)': 24.17, 'memory/max_allocated (GiB)': 24.17, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1619.33, 'epoch': 0.27}
9%|████████████▎ | 39/432 [20:06<3:03:55, 28.08s/it]
9%|████████████▌ | 40/432 [20:35<3:05:18, 28.36s/it]
{'loss': 0.9365, 'grad_norm': 0.09776122868061066, 'learning_rate': 0.0001813953488372093, 'memory/max_active (GiB)': 26.06, 'memory/max_allocated (GiB)': 26.06, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1708.88, 'epoch': 0.28}
9%|████████████▌ | 40/432 [20:35<3:05:18, 28.36s/it]
9%|████████████▉ | 41/432 [21:07<3:12:06, 29.48s/it]
{'loss': 0.9455, 'grad_norm': 0.09212527424097061, 'learning_rate': 0.000186046511627907, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1854.36, 'epoch': 0.28}
9%|████████████▉ | 41/432 [21:07<3:12:06, 29.48s/it]
10%|█████████████▏ | 42/432 [21:32<3:03:44, 28.27s/it]
{'loss': 0.9964, 'grad_norm': 0.1160384938120842, 'learning_rate': 0.00019069767441860466, 'memory/max_active (GiB)': 23.7, 'memory/max_allocated (GiB)': 23.7, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1775.41, 'epoch': 0.29}
10%|█████████████▏ | 42/432 [21:32<3:03:44, 28.27s/it]
10%|█████████████▌ | 43/432 [22:02<3:06:18, 28.74s/it]
{'loss': 0.8627, 'grad_norm': 0.06805545091629028, 'learning_rate': 0.00019534883720930232, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1830.25, 'epoch': 0.3}
10%|█████████████▌ | 43/432 [22:02<3:06:18, 28.74s/it]
10%|█████████████▊ | 44/432 [22:35<3:13:22, 29.90s/it]
{'loss': 0.9417, 'grad_norm': 0.06951376795768738, 'learning_rate': 0.0002, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1843.05, 'epoch': 0.31}
10%|█████████████▊ | 44/432 [22:35<3:13:22, 29.90s/it]
10%|██████████████▏ | 45/432 [23:07<3:18:05, 30.71s/it]
{'loss': 0.9017, 'grad_norm': 0.06728649139404297, 'learning_rate': 0.00019999673886943734, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1889.73, 'epoch': 0.31}
10%|██████████████▏ | 45/432 [23:07<3:18:05, 30.71s/it]
11%|██████████████▍ | 46/432 [23:39<3:19:09, 30.96s/it]
{'loss': 1.0087, 'grad_norm': 0.0888209342956543, 'learning_rate': 0.0001999869556904488, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1805.58, 'epoch': 0.32}
11%|██████████████▍ | 46/432 [23:39<3:19:09, 30.96s/it]
11%|██████████████▊ | 47/432 [24:08<3:16:18, 30.59s/it]
{'loss': 0.9246, 'grad_norm': 0.07477093487977982, 'learning_rate': 0.00019997065110111885, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1807.73, 'epoch': 0.33}
11%|██████████████▊ | 47/432 [24:08<3:16:18, 30.59s/it]
11%|███████████████ | 48/432 [24:39<3:15:53, 30.61s/it]
{'loss': 0.936, 'grad_norm': 0.08000776916742325, 'learning_rate': 0.00019994782616487538, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1732.16, 'epoch': 0.33}
11%|███████████████ | 48/432 [24:39<3:15:53, 30.61s/it]
11%|███████████████▍ | 49/432 [25:06<3:08:03, 29.46s/it]
{'loss': 0.9732, 'grad_norm': 0.2703610956668854, 'learning_rate': 0.00019991848237042035, 'memory/max_active (GiB)': 24.64, 'memory/max_allocated (GiB)': 24.64, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1814.28, 'epoch': 0.34}
11%|███████████████▍ | 49/432 [25:06<3:08:03, 29.46s/it]
12%|███████████████▋ | 50/432 [25:34<3:04:18, 28.95s/it]
{'loss': 0.9867, 'grad_norm': 0.08173573762178421, 'learning_rate': 0.00019988262163163264, 'memory/max_active (GiB)': 26.25, 'memory/max_allocated (GiB)': 26.25, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1746.16, 'epoch': 0.35}
12%|███████████████▋ | 50/432 [25:34<3:04:18, 28.95s/it]
12%|████████████████ | 51/432 [26:06<3:10:51, 30.06s/it]
{'loss': 0.9353, 'grad_norm': 0.06703449040651321, 'learning_rate': 0.00019984024628744328, 'memory/max_active (GiB)': 25.31, 'memory/max_allocated (GiB)': 25.31, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1803.49, 'epoch': 0.35}
12%|████████████████ | 51/432 [26:06<3:10:51, 30.06s/it]
12%|████████████████▎ | 52/432 [26:31<3:00:46, 28.54s/it]
{'loss': 0.9705, 'grad_norm': 0.0770621970295906, 'learning_rate': 0.0001997913591016829, 'memory/max_active (GiB)': 27.44, 'memory/max_allocated (GiB)': 27.44, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1805.23, 'epoch': 0.36}
12%|████████████████▎ | 52/432 [26:31<3:00:46, 28.54s/it]
12%|████████████████▋ | 53/432 [27:04<3:08:21, 29.82s/it]
{'loss': 0.9082, 'grad_norm': 0.08800782263278961, 'learning_rate': 0.00019973596326290137, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1787.79, 'epoch': 0.37}
12%|████████████████▋ | 53/432 [27:04<3:08:21, 29.82s/it]
12%|█████████████████ | 54/432 [27:38<3:16:01, 31.11s/it]
{'loss': 0.964, 'grad_norm': 0.0656328946352005, 'learning_rate': 0.00019967406238415998, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1888.13, 'epoch': 0.38}
12%|█████████████████ | 54/432 [27:38<3:16:01, 31.11s/it]
13%|█████████████████▎ | 55/432 [28:08<3:13:44, 30.84s/it]
{'loss': 0.918, 'grad_norm': 0.09178014099597931, 'learning_rate': 0.00019960566050279566, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1822.06, 'epoch': 0.38}
13%|█████████████████▎ | 55/432 [28:08<3:13:44, 30.84s/it]
13%|█████████████████▋ | 56/432 [28:34<3:03:56, 29.35s/it]
{'loss': 1.0078, 'grad_norm': 0.07544898241758347, 'learning_rate': 0.00019953076208015772, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1779.76, 'epoch': 0.39}
13%|█████████████████▋ | 56/432 [28:34<3:03:56, 29.35s/it]
13%|█████████████████▉ | 57/432 [29:07<3:09:09, 30.27s/it]
{'loss': 0.952, 'grad_norm': 0.07013165950775146, 'learning_rate': 0.0001994493720013169, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1875.9, 'epoch': 0.4}
13%|█████████████████▉ | 57/432 [29:07<3:09:09, 30.27s/it]
13%|██████████████████▎ | 58/432 [29:36<3:06:17, 29.89s/it]
{'loss': 0.9751, 'grad_norm': 0.2212851643562317, 'learning_rate': 0.00019936149557474666, 'memory/max_active (GiB)': 27.01, 'memory/max_allocated (GiB)': 27.01, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1808.54, 'epoch': 0.4}
13%|██████████████████▎ | 58/432 [29:36<3:06:17, 29.89s/it]
14%|██████████████████▌ | 59/432 [30:09<3:12:19, 30.94s/it]
{'loss': 0.9055, 'grad_norm': 0.0661102756857872, 'learning_rate': 0.00019926713853197695, 'memory/max_active (GiB)': 24.74, 'memory/max_allocated (GiB)': 24.74, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1802.3, 'epoch': 0.41}
14%|██████████████████▌ | 59/432 [30:09<3:12:19, 30.94s/it]
14%|██████████████████▉ | 60/432 [30:40<3:12:31, 31.05s/it]
{'loss': 0.9826, 'grad_norm': 0.08663811534643173, 'learning_rate': 0.0001991663070272206, 'memory/max_active (GiB)': 24.17, 'memory/max_allocated (GiB)': 24.17, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1707.88, 'epoch': 0.42}
14%|██████████████████▉ | 60/432 [30:40<3:12:31, 31.05s/it]
14%|███████████████████▏ | 61/432 [31:14<3:15:52, 31.68s/it]
{'loss': 0.9759, 'grad_norm': 0.07683200389146805, 'learning_rate': 0.0001990590076369715, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1895.98, 'epoch': 0.42}
14%|███████████████████▏ | 61/432 [31:14<3:15:52, 31.68s/it]
14%|███████████████████▌ | 62/432 [31:45<3:14:33, 31.55s/it]
{'loss': 0.9168, 'grad_norm': 0.07595925778150558, 'learning_rate': 0.00019894524735957622, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1797.18, 'epoch': 0.43}
14%|███████████████████▌ | 62/432 [31:45<3:14:33, 31.55s/it]
15%|███████████████████▊ | 63/432 [32:13<3:08:40, 30.68s/it]
{'loss': 0.9679, 'grad_norm': 0.07661418616771698, 'learning_rate': 0.00019882503361477705, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1830.69, 'epoch': 0.44}
15%|███████████████████▊ | 63/432 [32:13<3:08:40, 30.68s/it]
15%|████████████████████▏ | 64/432 [32:44<3:08:27, 30.73s/it]
{'loss': 0.9592, 'grad_norm': 0.08054457604885101, 'learning_rate': 0.00019869837424322829, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1798.08, 'epoch': 0.44}
15%|████████████████████▏ | 64/432 [32:44<3:08:27, 30.73s/it]
15%|████████████████████▍ | 65/432 [33:15<3:08:07, 30.76s/it]
{'loss': 0.9257, 'grad_norm': 0.08320043236017227, 'learning_rate': 0.00019856527750598493, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1795.23, 'epoch': 0.45}
15%|████████████████████▍ | 65/432 [33:15<3:08:07, 30.76s/it]
15%|████████████████████▊ | 66/432 [33:49<3:13:58, 31.80s/it]
{'loss': 0.8969, 'grad_norm': 0.0733579471707344, 'learning_rate': 0.00019842575208396372, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1861.37, 'epoch': 0.46}
15%|████████████████████▊ | 66/432 [33:49<3:13:58, 31.80s/it]
16%|█████████████████████ | 67/432 [34:22<3:14:20, 31.95s/it]
{'loss': 0.8604, 'grad_norm': 0.29595091938972473, 'learning_rate': 0.00019827980707737703, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1844.94, 'epoch': 0.47}
16%|█████████████████████ | 67/432 [34:22<3:14:20, 31.95s/it]
16%|█████████████████████▍ | 68/432 [34:51<3:09:19, 31.21s/it]
{'loss': 0.9479, 'grad_norm': 0.10486430674791336, 'learning_rate': 0.00019812745200513927, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1823.2, 'epoch': 0.47}
16%|█████████████████████▍ | 68/432 [34:51<3:09:19, 31.21s/it]
16%|█████████████████████▋ | 69/432 [35:28<3:18:46, 32.85s/it]
{'loss': 0.9287, 'grad_norm': 0.13543325662612915, 'learning_rate': 0.0001979686968042461, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1821.33, 'epoch': 0.48}
16%|█████████████████████▋ | 69/432 [35:28<3:18:46, 32.85s/it]
16%|██████████████████████ | 70/432 [36:00<3:16:21, 32.55s/it]
{'loss': 0.9248, 'grad_norm': 0.07873474061489105, 'learning_rate': 0.00019780355182912626, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1771.5, 'epoch': 0.49}
16%|██████████████████████ | 70/432 [36:00<3:16:21, 32.55s/it]
16%|██████████████████████▎ | 71/432 [36:30<3:11:56, 31.90s/it]
{'loss': 0.9172, 'grad_norm': 0.06926668435335159, 'learning_rate': 0.0001976320278509663, 'memory/max_active (GiB)': 26.06, 'memory/max_allocated (GiB)': 26.06, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1875.05, 'epoch': 0.49}
16%|██████████████████████▎ | 71/432 [36:30<3:11:56, 31.90s/it]
17%|██████████████████████▋ | 72/432 [37:06<3:19:18, 33.22s/it]
{'loss': 0.8823, 'grad_norm': 0.08082268387079239, 'learning_rate': 0.0001974541360570079, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1865.89, 'epoch': 0.5}
17%|██████████████████████▋ | 72/432 [37:06<3:19:18, 33.22s/it]
17%|██████████████████████▉ | 73/432 [37:37<3:14:13, 32.46s/it]
{'loss': 0.9185, 'grad_norm': 0.07178379595279694, 'learning_rate': 0.00019726988804981844, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1854.53, 'epoch': 0.51}
17%|██████████████████████▉ | 73/432 [37:37<3:14:13, 32.46s/it]
17%|███████████████████████▎ | 74/432 [38:09<3:12:00, 32.18s/it]
{'loss': 0.9461, 'grad_norm': 0.07196955382823944, 'learning_rate': 0.00019707929584653408, 'memory/max_active (GiB)': 28.15, 'memory/max_allocated (GiB)': 28.15, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1820.12, 'epoch': 0.51}
17%|███████████████████████▎ | 74/432 [38:09<3:12:00, 32.18s/it]
17%|███████████████████████▌ | 75/432 [38:35<3:01:38, 30.53s/it]
{'loss': 1.0447, 'grad_norm': 0.07153692096471786, 'learning_rate': 0.00019688237187807594, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1809.64, 'epoch': 0.52}
17%|███████████████████████▌ | 75/432 [38:35<3:01:38, 30.53s/it]
18%|███████████████████████▉ | 76/432 [39:04<2:57:30, 29.92s/it]
{'loss': 0.8106, 'grad_norm': 0.06721945106983185, 'learning_rate': 0.00019667912898833955, 'memory/max_active (GiB)': 23.22, 'memory/max_allocated (GiB)': 23.22, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1785.2, 'epoch': 0.53}
18%|███████████████████████▉ | 76/432 [39:04<2:57:30, 29.92s/it]
18%|████████████████████████▏ | 77/432 [39:35<2:58:54, 30.24s/it]
{'loss': 0.9299, 'grad_norm': 0.08322236686944962, 'learning_rate': 0.00019646958043335677, 'memory/max_active (GiB)': 24.64, 'memory/max_allocated (GiB)': 24.64, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1836.34, 'epoch': 0.53}
18%|████████████████████████▏ | 77/432 [39:35<2:58:54, 30.24s/it]
18%|████████████████████████▌ | 78/432 [40:07<3:02:40, 30.96s/it]
{'loss': 0.9262, 'grad_norm': 0.06773433834314346, 'learning_rate': 0.00019625373988043165, 'memory/max_active (GiB)': 27.95, 'memory/max_allocated (GiB)': 27.95, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1850.19, 'epoch': 0.54}
18%|████████████████████████▌ | 78/432 [40:07<3:02:40, 30.96s/it]
18%|████████████████████████▊ | 79/432 [40:40<3:05:38, 31.55s/it]
{'loss': 0.9067, 'grad_norm': 0.06558340042829514, 'learning_rate': 0.00019603162140724862, 'memory/max_active (GiB)': 27.2, 'memory/max_allocated (GiB)': 27.2, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1876.53, 'epoch': 0.55}
18%|████████████████████████▊ | 79/432 [40:40<3:05:38, 31.55s/it]
19%|█████████████████████████▏ | 80/432 [41:07<2:56:33, 30.10s/it]
{'loss': 0.8971, 'grad_norm': 0.06962298601865768, 'learning_rate': 0.0001958032395009545, 'memory/max_active (GiB)': 25.68, 'memory/max_allocated (GiB)': 25.68, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1767.39, 'epoch': 0.56}
19%|█████████████████████████▏ | 80/432 [41:07<2:56:33, 30.10s/it]
19%|█████████████████████████▌ | 81/432 [41:32<2:47:00, 28.55s/it]
{'loss': 0.9593, 'grad_norm': 0.08821487426757812, 'learning_rate': 0.00019556860905721362, 'memory/max_active (GiB)': 25.11, 'memory/max_allocated (GiB)': 25.11, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1789.22, 'epoch': 0.56}
19%|█████████████████████████▌ | 81/432 [41:32<2:47:00, 28.55s/it]
19%|█████████████████████████▊ | 82/432 [42:00<2:45:10, 28.32s/it]
{'loss': 0.9409, 'grad_norm': 0.07249249517917633, 'learning_rate': 0.00019532774537923617, 'memory/max_active (GiB)': 27.44, 'memory/max_allocated (GiB)': 27.44, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1764.78, 'epoch': 0.57}
19%|█████████████████████████▊ | 82/432 [42:00<2:45:10, 28.32s/it]
19%|██████████████████████████▏ | 83/432 [42:29<2:46:56, 28.70s/it]
{'loss': 0.8989, 'grad_norm': 0.08999690413475037, 'learning_rate': 0.00019508066417678018, 'memory/max_active (GiB)': 28.9, 'memory/max_allocated (GiB)': 28.9, 'memory/device_reserved (GiB)': 32.39, 'tokens_per_second_per_gpu': 1821.55, 'epoch': 0.58}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f57df1b1377e70b47a5de74cc6e3b67e7a535f2a9db8474749cee539ca959de
size 29547716064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:438f02a1b51403bf963ccd1a3f126c43bdd106faed70c4c15481ef343afd21bb
size 8988110304

3
img-assets/header.jpg Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:34a352d49d78cee433ab323a5ada82cf7fba1e87980eee32405ff99df1c78c58
size 337446

View File

@@ -0,0 +1,24 @@
{
"</tool_call>": 151658,
"<tool_call>": 151657,
"<|box_end|>": 151649,
"<|box_start|>": 151648,
"<|endoftext|>": 151643,
"<|file_sep|>": 151664,
"<|fim_middle|>": 151660,
"<|fim_pad|>": 151662,
"<|fim_prefix|>": 151659,
"<|fim_suffix|>": 151661,
"<|im_end|>": 151645,
"<|im_start|>": 151644,
"<|image_pad|>": 151655,
"<|object_ref_end|>": 151647,
"<|object_ref_start|>": 151646,
"<|quad_end|>": 151651,
"<|quad_start|>": 151650,
"<|repo_name|>": 151663,
"<|video_pad|>": 151656,
"<|vision_end|>": 151653,
"<|vision_pad|>": 151654,
"<|vision_start|>": 151652
}

View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

78
merged-hf/config.json Normal file
View File

@@ -0,0 +1,78 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "bfloat16",
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 48,
"model_type": "qwen2",
"num_attention_heads": 40,
"num_hidden_layers": 48,
"num_key_value_heads": 8,
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000.0,
"sliding_window": null,
"tie_word_embeddings": false,
"transformers_version": "4.57.3",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 152064
}

View File

@@ -0,0 +1,14 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"repetition_penalty": 1.05,
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8,
"transformers_version": "4.57.3"
}

151388
merged-hf/merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6942efc3db5b2aafff502b17fb261902269a4728c3e60ee003cbe65156a72e78
size 4986211280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b58851de2ada64e5ea46c942279cfe8703f81aa10c50a9a9766f408c92db24bc
size 4954847344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e37819a31f6a7291564b90f4d3893a5caa0fcf0c33e4efb76cf7ee087ec2f26
size 4954847392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:572f40e8bea1d50a0ab58ae9184c69dfb1989e715ca1accdc6aed478f4b7c4f7
size 4954847392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eaff6905eec560801bfca2311ff491a38899a477fe4be851fcc00f43c50aca4c
size 4954847392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1f83d0d9d38a1f8320c6d11c1bba3f4600130741e03467fc068e6ea08efea959
size 4734533160

View File

@@ -0,0 +1,587 @@
{
"metadata": {
"total_parameters": 14770033664,
"total_size": 29540067328
},
"weight_map": {
"lm_head.weight": "model-00006-of-00006.safetensors",
"model.embed_tokens.weight": "model-00001-of-00006.safetensors",
"model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.10.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.10.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.11.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.11.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.12.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.12.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.13.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.13.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.14.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.14.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.15.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.15.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.16.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.17.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.18.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.18.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.19.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.19.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.2.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.20.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.20.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.20.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.21.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.21.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.22.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.22.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.input_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.23.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.23.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.24.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.24.self_attn.k_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.q_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.v_proj.bias": "model-00003-of-00006.safetensors",
"model.layers.24.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
"model.layers.25.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.25.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.26.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.26.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.27.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.27.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.28.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.28.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.29.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.29.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.3.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.3.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.3.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.30.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.30.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.30.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.31.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.31.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.input_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.32.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.32.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.33.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.33.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.33.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.33.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.33.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.33.self_attn.k_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.q_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.v_proj.bias": "model-00004-of-00006.safetensors",
"model.layers.33.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
"model.layers.34.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.34.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.34.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.35.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.35.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.36.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.36.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.37.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.37.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.38.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.38.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.39.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.39.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.4.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.4.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.4.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.40.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.40.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.40.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.input_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.41.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.41.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.42.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.42.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.42.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.42.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.42.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.42.self_attn.k_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.q_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.v_proj.bias": "model-00005-of-00006.safetensors",
"model.layers.42.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
"model.layers.43.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.43.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.k_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.q_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.v_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.43.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.44.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.k_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.q_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.v_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.44.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.45.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.k_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.q_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.v_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.45.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.46.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.k_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.q_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.v_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.46.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.input_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.47.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.k_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.q_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.v_proj.bias": "model-00006-of-00006.safetensors",
"model.layers.47.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
"model.layers.5.input_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.5.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.5.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.6.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.6.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.6.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.6.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.6.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.6.self_attn.k_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.q_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.v_proj.bias": "model-00001-of-00006.safetensors",
"model.layers.6.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
"model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.k_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.q_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.v_proj.bias": "model-00002-of-00006.safetensors",
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
"model.norm.weight": "model-00006-of-00006.safetensors"
}
}

View File

@@ -0,0 +1,31 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"eos_token": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

3
merged-hf/tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c5ae00e602b8860cbd784ba82a8aa14e8feecec692e7076590d014d7b7fdafa
size 11421896

View File

@@ -0,0 +1,207 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 32768,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

1
merged-hf/vocab.json Normal file

File diff suppressed because one or more lines are too long

151388
merges.txt Normal file

File diff suppressed because it is too large Load Diff

7
mlx-q4/README.md Normal file
View File

@@ -0,0 +1,7 @@
---
language: en
tags:
- mlx
library_name: mlx
pipeline_tag: text-generation
---

24
mlx-q4/added_tokens.json Normal file
View File

@@ -0,0 +1,24 @@
{
"</tool_call>": 151658,
"<tool_call>": 151657,
"<|box_end|>": 151649,
"<|box_start|>": 151648,
"<|endoftext|>": 151643,
"<|file_sep|>": 151664,
"<|fim_middle|>": 151660,
"<|fim_pad|>": 151662,
"<|fim_prefix|>": 151659,
"<|fim_suffix|>": 151661,
"<|im_end|>": 151645,
"<|im_start|>": 151644,
"<|image_pad|>": 151655,
"<|object_ref_end|>": 151647,
"<|object_ref_start|>": 151646,
"<|quad_end|>": 151651,
"<|quad_start|>": 151650,
"<|repo_name|>": 151663,
"<|video_pad|>": 151656,
"<|vision_end|>": 151653,
"<|vision_pad|>": 151654,
"<|vision_start|>": 151652
}

View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

88
mlx-q4/config.json Normal file
View File

@@ -0,0 +1,88 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "bfloat16",
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 48,
"model_type": "qwen2",
"num_attention_heads": 40,
"num_hidden_layers": 48,
"num_key_value_heads": 8,
"quantization": {
"group_size": 64,
"bits": 4,
"mode": "affine"
},
"quantization_config": {
"group_size": 64,
"bits": 4,
"mode": "affine"
},
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000.0,
"sliding_window": null,
"tie_word_embeddings": false,
"transformers_version": "4.57.3",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 152064
}

View File

@@ -0,0 +1,14 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"repetition_penalty": 1.05,
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8,
"transformers_version": "4.57.3"
}

151388
mlx-q4/merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15ed4a295e8bff61c368ed0f1364d1b592c343e2034fa9d33f5e88d111ba7182
size 5353840945

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4985576339d0e33bfeefdcc993a85240f42fb1716b75810924e9701795950a52
size 2955654333

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,31 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"eos_token": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

3
mlx-q4/tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c5ae00e602b8860cbd784ba82a8aa14e8feecec692e7076590d014d7b7fdafa
size 11421896

View File

@@ -0,0 +1,207 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 32768,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

1
mlx-q4/vocab.json Normal file

File diff suppressed because one or more lines are too long

31
special_tokens_map.json Normal file
View File

@@ -0,0 +1,31 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"eos_token": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

32
test_model.py Normal file
View File

@@ -0,0 +1,32 @@
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model_name = "Qwen/Qwen2.5-Coder-14B-Instruct"
adapter_path = "./outputs/qwen25-coder-n8n"
print("Loading base model...")
base_model = AutoModelForCausalLM.from_pretrained(
base_model_name,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
print("Loading adapter...")
model = PeftModel.from_pretrained(base_model, adapter_path)
tokenizer = AutoTokenizer.from_pretrained(base_model_name)
system_prompt = "You are an expert n8n workflow generation assistant. Your goal is to create valid, efficient, and error-free n8n workflow JSONs based on the user's requirements. Always output ONLY the valid JSON workflow."
user_input = "Create a workflow that gets data from a webhook and sends it to Slack. Also have a sticky note as documentation."
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_input}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
print("Generating workflow...")
outputs = model.generate(**inputs, max_new_tokens=2048, do_sample=True, temperature=0.1)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c5ae00e602b8860cbd784ba82a8aa14e8feecec692e7076590d014d7b7fdafa
size 11421896

207
tokenizer_config.json Normal file
View File

@@ -0,0 +1,207 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 32768,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

1
vocab.json Normal file

File diff suppressed because one or more lines are too long