初始化项目，由ModelHub XC社区提供模型

Model: OctoMed/OctoMed-7B Source: Original Platform
2026-05-27 10:20:20 +08:00
commit aacbdde26c
17 changed files with 458698 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,35 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,301 @@
 ---
 license: apache-2.0
 language:
 - en
 pipeline_tag: image-text-to-text
 tags:
 - multimodal
 library_name: transformers
 base_model:
 - Qwen/Qwen2.5-VL-7B-Instruct
 ---
 # <img src="assets/OctoMed.svg" alt="OctoMed Logo" width="100" style="vertical-align:bottom; margin-right:0px;" /> OctoMed-7B
 ## Introduction
 OctoMed-7B is a high-performance multimodal medical reasoning model created through large-scale data curation and supervised fine-tuning (SFT). To support reliable clinical reasoning, we developed a scalable data pipeline that distills structured reasoning traces from DeepSeek-R1 and GPT-4o and produced the largest multimodal medical reasoning dataset to date with more than 8 million traces and 6.8 billion response tokens.
 Using Qwen2.5-VL-7B-Instruct as the base model, OctoMed-7B is trained on this curated corpus and achieves strong, robust performance on a wide range of out-of-distribution medical benchmarks.
 OctoMed-7B produces internal reasoning traces in \<think>...\</think> tokens before writing out its final answer. In general, the model has a tendency to think longer for harder or ill-defined questions, while sticking to shorter reasoning traces for easier queries.
 ## Evaluation
 ### Medical Benchmark Performances
 <p align="center">
    <img src="assets/performances.svg" alt="Medical Benchmark Performances" width="100%" />
 </p>
 **Notes:**  
 - Green = OSS smaller models (<10B), Cyan = large proprietary models.  
 - † = 10-sample majority vote ensemble result.
 ### Legacy Medical Benchmark Performance
 | Dataset  | Setting | Performance |
 |----------|---------|--------------|
 | VQA-RAD  | Open (Token F1)    | 64.23        |
 | VQA-RAD  | Closed (Accuracy)  | 85.66        |
 | SLAKE    | Open (Token F1)   | 84.96        |
 | SLAKE    | Closed (Accuracy) | 89.66        |
 We also train on the train splits of the VQA-RAD and SLAKE datasets and report the performances here. For these results, we apply a **direct** prompt by including the phrase **Answer in a short word or phrase.** at the end of each sample. GPT2 is used as the tokenizer to compute Token F1 for open-ended questions following prior work.
 ## Requirements
 We recommend installing the transformers version used in our experiments and other dependencies with this command:
 ```
 pip install transformers==4.57.1 accelerate==1.12.0 torchvision==0.24.1 qwen-vl-utils==0.0.14
 ```
 ## Quickstart
 Below, we provide a some examples to show how to use OctoMed-7B with 🤗 Transformers or vLLM.
 <details>
 <summary>Inference with HF Transformers 🤗</summary>
 Here we show a code snippet to show you how chat with OctoMed-7B using `transformers` and `qwen_vl_utils`:
 ```python
 import torch
 from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
 from qwen_vl_utils import process_vision_info
 # default: Load the model on the available device(s)
 model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "OctoMed/OctoMed-7B", dtype=torch.bfloat16, device_map="auto"
 )
 # We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image and video scenarios.
 # model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
 #     "OctoMed/OctoMed-7B",
 #     dtype=torch.bfloat16,
 #     attn_implementation="flash_attention_2",
 #     device_map="auto",
 # )
 # The default range for the number of visual tokens per image in the model is 4-16384.
 # You can set min_pixels and max_pixels according to your needs, such as a token range of 256-1280, to balance performance and cost.
 min_pixels = 262144
 max_pixels = 262144
 processor = AutoProcessor.from_pretrained("OctoMed/OctoMed-7B", min_pixels=min_pixels, max_pixels=max_pixels)
 # Text-Only Query
 # messages = [
 #     {
 #         "role": "user",
 #         "content": [
 #             {"type": "text", "text": "I've had a persistent dry cough for two weeks but no fever. Could this be allergies, and when should I see a doctor?"},
 #         ],
 #     }
 # ]
 # General Query
 # messages = [
 #     {
 #         "role": "user",
 #         "content": [
 #             {
 #                 "type": "image",
 #                 "image": "https://cdn.ncbi.nlm.nih.gov/pmc/blobs/51b2/10835941/13323b55fbb5/13256_2024_4349_Fig1_HTML.jpg",
 #             },
 #             {"type": "text", "text": "Describe this image."},
 #         ],
 #     }
 # ]
 # Multiple Choice Query
 messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://cdn.ncbi.nlm.nih.gov/pmc/blobs/51b2/10835941/13323b55fbb5/13256_2024_4349_Fig1_HTML.jpg",
            },
            {"type": "text", "text": "What orientation was the MRI in image B taken in?\nA. Axial\nB. Coronal\nC. Sagittal\nD. Oblique\n\nPlease reason step-by-step, and put your final answer within \\boxed{}."},
        ],
    }
 ]
 # Preparation for inference
 text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
 )
 image_inputs, video_inputs = process_vision_info(messages)
 inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
 )
 inputs = inputs.to(device="cuda")
 # Inference: Generation of the output
 generated_ids = model.generate(**inputs, max_new_tokens=8192)
 generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
 ]
 output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
 )
 print(output_text)
 ```
 </details>
 <details>
 <summary>Inference with vLLM</summary>
 Here we show an example of how to use OctoMed with vLLM (tested with vLLM==0.11.2 and transformers==4.57.1):
 ```python
 from vllm import LLM, SamplingParams
 from transformers import AutoProcessor
 min_pixels = 262144
 max_pixels = 262144
 processor = AutoProcessor.from_pretrained("OctoMed/OctoMed-7B", min_pixels=min_pixels, max_pixels=max_pixels)
 llm = LLM(
    model="OctoMed/OctoMed-7B",
    trust_remote_code=True,
    dtype="bfloat16",
    max_model_len=8192,
    tensor_parallel_size=4,
    gpu_memory_utilization=0.8,
    limit_mm_per_prompt={"image": 1}
 )
 # Set up sampling parameters
 sampling_params = SamplingParams(
    temperature=0.6,
    top_p=0.95,
    max_tokens=8192,
 )
 image_data = []
 # Text-Only Query
 messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Explain the difference between type 1 and type 2 diabetes."},
        ],
    }
 ]
 # General Query
 # image_data = ['https://cdn.ncbi.nlm.nih.gov/pmc/blobs/51b2/10835941/13323b55fbb5/13256_2024_4349_Fig1_HTML.jpg']
 # messages = [
 #     {
 #         "role": "user",
 #         "content": [
 #             {
 #                 "type": "image",
 #                 "image": image_data[0],
 #             },
 #             {"type": "text", "text": "Describe this image."},
 #         ],
 #     }
 # ]
 # Multiple Choice Query
 # image_data = ['https://cdn.ncbi.nlm.nih.gov/pmc/blobs/51b2/10835941/13323b55fbb5/13256_2024_4349_Fig1_HTML.jpg']
 # messages = [
 #     {
 #         "role": "user",
 #         "content": [
 #             {
 #                 "type": "image",
 #                 "image": image_data[0],
 #             },
 #             {"type": "text", "text": "What orientation was the MRI in image B taken in?\nA. Axial\nB. Coronal\nC. Sagittal\nD. Oblique\n\nPlease reason step-by-step, and put your final answer within \\boxed{}."},
 #         ],
 #     }
 # ]
 prompt = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True)
 if image_data:
    mm_prompt = {
        "prompt": prompt,
        "multi_modal_data": {"image": image_data}
    }
 else:
    mm_prompt = {"prompt": prompt}
 # Generate response
 outputs = llm.generate([mm_prompt], sampling_params)
 # Print the generated response
 for output in outputs:
    prompt = output.prompt
    generated_text = output.outputs[0].text
    print(f"Prompt: {prompt}")
    print(f"Generated text: {generated_text}")
    print("-" * 50)
 ```
 </details>
 ### Suggested Hyperparameters
 We suggest using the same settings used in evaluation to reproduce results:
 Format multiple choice questions with the following template:
 ```
 {optional image(s)}
 {question}
 {options, 1 on each line}
 Please reason step-by-step, and put your final answer within \\boxed{}.
 ```
 Example Prompt:
 ```
 {image(s)}
 What orientation was the MRI in image B taken in?
 A: Axial
 B: Coronal
 C: Sagittal
 D: Oblique
 Please reason step-by-step, and put your final answer within \\boxed{}.
 ```
 - Use the default system prompt ("You are a helpful assistant.")
 - Extract the answer by looking at the content within the last \\boxed{}.
 - Temperature of 0.6
 - Top-p of 0.95
 - min_pixels = 262144
 - max_pixels = 262144
 ### Known Issues
 * Model is sensitive to system prompt. We recommend using the default one.
 * The model is finetuned for multiple-choice VQA. The model may follow instructions for other tasks but is not extensively tested or post-trained to do so.
 We hope to address these concerns moving forward in future iterations!
 ## Citation
 If you find our work helpful, feel free to give us a cite.
 ```
@article{ossowski2025octomed,
  title={OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning},
  author={Ossowski, Timothy and Zhang, Sheng and Liu, Qianchu and Qin, Guanghui and Tan, Reuben and Naumann, Tristan and Hu, Junjie and Poon, Hoifung},
  journal={arXiv preprint arXiv:2511.23269},
  year={2025}
 }
 ```
--- a/assets/OctoMed.svg
+++ b/assets/OctoMed.svg
--- a/assets/performances.svg
+++ b/assets/performances.svg
--- a/chat_template.jinja
+++ b/chat_template.jinja
@@ -0,0 +1,7 @@
 {% set image_count = namespace(value=0) %}{% set video_count = namespace(value=0) %}{% for message in messages %}{% if loop.first and message['role'] != 'system' %}<|im_start|>system
 You are a helpful assistant.<|im_end|>
 {% endif %}<|im_start|>{{ message['role'] }}
 {% if message['content'] is string %}{{ message['content'] }}<|im_end|>
 {% else %}{% for content in message['content'] %}{% if content['type'] == 'image' or 'image' in content or 'image_url' in content %}{% set image_count.value = image_count.value + 1 %}{% if add_vision_id %}Picture {{ image_count.value }}: {% endif %}<|vision_start|><|image_pad|><|vision_end|>{% elif content['type'] == 'video' or 'video' in content %}{% set video_count.value = video_count.value + 1 %}{% if add_vision_id %}Video {{ video_count.value }}: {% endif %}<|vision_start|><|video_pad|><|vision_end|>{% elif 'text' in content %}{{ content['text'] }}{% endif %}{% endfor %}<|im_end|>
 {% endif %}{% endfor %}{% if add_generation_prompt %}<|im_start|>assistant
 {% endif %}
--- a/config.json
+++ b/config.json
@@ -0,0 +1,50 @@
 {
  "architectures": ["Qwen2_5_VLForConditionalGeneration"],
  "attention_dropout": 0.0,
  "bos_token_id": 151643,
  "eos_token_id": 151645,
  "vision_start_token_id": 151652,
  "vision_end_token_id": 151653,
  "vision_token_id": 151654,
  "image_token_id": 151655,
  "video_token_id": 151656,
  "hidden_act": "silu",
  "hidden_size": 3584,
  "initializer_range": 0.02,
  "intermediate_size": 18944,
  "max_position_embeddings": 128000,
  "max_window_layers": 28,
  "model_type": "qwen2_5_vl",
  "num_attention_heads": 28,
  "num_hidden_layers": 28,
  "num_key_value_heads": 4,
  "rms_norm_eps": 1e-06,
  "rope_theta": 1000000.0,
  "sliding_window": 32768,
  "tie_word_embeddings": false,
  "torch_dtype": "bfloat16",
  "transformers_version": "4.41.2",
  "use_cache": true,
  "use_sliding_window": false,
  "rope_scaling": {
    "type": "mrope",
    "mrope_section": [16, 24, 24]
  },
  "vision_config": {
    "depth": 32,
    "hidden_act": "silu",
    "hidden_size": 1280,
    "intermediate_size": 3420,
    "num_heads": 16,
    "in_chans": 3,
    "out_hidden_size": 3584,
    "patch_size": 14,
    "spatial_merge_size": 2,
    "spatial_patch_size": 14,
    "window_size": 112,
    "fullatt_block_indexes": [7, 15, 23, 31],
    "tokens_per_second": 2,
    "temporal_patch_size": 2
  },
  "vocab_size": 152064
 }
--- a/generation_config.json
+++ b/generation_config.json
@@ -0,0 +1,11 @@
 {
  "do_sample": true,
  "eos_token_id": [
    151645,
    151643
  ],
  "pad_token_id": 151643,
  "repetition_penalty": 1.05,
  "temperature": 1e-06,
  "transformers_version": "4.57.1"
 }
--- a/merges.txt
+++ b/merges.txt
--- a/model-00001-of-00004.safetensors
+++ b/model-00001-of-00004.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:d771df9eee7c5fd3f6b31b44fbecc669f8e7460916e91c5f1088037d769fa2df
 size 4968243304
--- a/model-00002-of-00004.safetensors
+++ b/model-00002-of-00004.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:90d9502e13084ef71c7af2f44cefb1acf63e4024066618e870b273f0f5eb4976
 size 4991495816
--- a/model-00003-of-00004.safetensors
+++ b/model-00003-of-00004.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:2e2982d12494e60eae87caf2fcb32aee485b87648e5d8a2a6a77b00050de70f8
 size 4932751040
--- a/model-00004-of-00004.safetensors
+++ b/model-00004-of-00004.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:ee28029c447d3af566050c0ce70daec97252933ad5056665284261403386151a
 size 1691924384
--- a/model.safetensors.index.json
+++ b/model.safetensors.index.json
@@ -0,0 +1,736 @@
 {
  "metadata": {
    "total_size": 16584333312
  },
  "weight_map": {
    "lm_head.weight": "model-00004-of-00004.safetensors",
    "model.embed_tokens.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.input_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.input_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.10.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.10.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.15.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.16.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.16.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.16.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.16.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.17.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.17.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.18.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.19.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.2.input_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.20.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.20.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.21.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.22.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.23.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.24.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.input_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.25.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.input_layernorm.weight": "model-00004-of-00004.safetensors",
    "model.layers.26.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.26.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
    "model.layers.26.self_attn.k_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.q_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.v_proj.bias": "model-00003-of-00004.safetensors",
    "model.layers.26.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
    "model.layers.27.input_layernorm.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.k_proj.bias": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.q_proj.bias": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.v_proj.bias": "model-00004-of-00004.safetensors",
    "model.layers.27.self_attn.v_proj.weight": "model-00004-of-00004.safetensors",
    "model.layers.3.input_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.input_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.5.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.5.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.5.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.5.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.5.self_attn.k_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.q_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.v_proj.bias": "model-00001-of-00004.safetensors",
    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
    "model.layers.6.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.6.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.7.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.input_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.k_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.q_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.v_proj.bias": "model-00002-of-00004.safetensors",
    "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
    "model.norm.weight": "model-00004-of-00004.safetensors",
    "visual.blocks.0.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.0.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.0.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.0.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.0.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.1.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.1.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.1.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.1.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.10.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.10.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.10.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.10.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.11.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.11.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.11.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.11.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.12.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.12.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.12.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.12.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.13.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.13.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.13.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.13.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.14.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.14.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.14.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.14.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.15.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.15.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.15.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.15.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.16.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.16.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.16.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.16.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.17.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.17.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.17.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.17.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.18.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.18.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.18.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.18.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.19.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.19.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.19.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.19.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.2.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.2.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.2.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.2.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.20.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.20.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.20.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.20.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.21.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.21.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.21.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.21.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.22.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.22.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.22.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.22.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.23.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.23.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.23.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.23.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.24.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.24.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.24.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.24.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.25.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.25.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.25.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.25.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.26.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.26.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.26.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.26.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.27.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.27.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.27.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.27.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.28.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.28.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.28.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.28.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.29.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.29.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.29.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.29.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.3.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.3.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.3.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.3.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.30.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.30.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.30.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.30.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.31.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.31.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.31.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.31.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.4.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.4.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.4.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.4.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.5.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.5.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.5.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.5.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.6.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.6.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.6.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.6.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.7.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.7.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.7.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.7.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.8.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.8.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.8.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.8.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.attn.proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.9.attn.proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.attn.qkv.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.9.attn.qkv.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.down_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.gate_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.up_proj.bias": "model-00001-of-00004.safetensors",
    "visual.blocks.9.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.norm1.weight": "model-00001-of-00004.safetensors",
    "visual.blocks.9.norm2.weight": "model-00001-of-00004.safetensors",
    "visual.merger.ln_q.weight": "model-00001-of-00004.safetensors",
    "visual.merger.mlp.0.bias": "model-00001-of-00004.safetensors",
    "visual.merger.mlp.0.weight": "model-00001-of-00004.safetensors",
    "visual.merger.mlp.2.bias": "model-00001-of-00004.safetensors",
    "visual.merger.mlp.2.weight": "model-00001-of-00004.safetensors",
    "visual.patch_embed.proj.weight": "model-00001-of-00004.safetensors"
  }
 }
--- a/processor_config.json
+++ b/processor_config.json
@@ -0,0 +1,39 @@
 {
  "crop_size": null,
  "data_format": "channels_first",
  "default_to_square": true,
  "device": null,
  "disable_grouping": null,
  "do_center_crop": null,
  "do_convert_rgb": true,
  "do_normalize": true,
  "do_pad": null,
  "do_rescale": true,
  "do_resize": true,
  "image_mean": [
    0.48145466,
    0.4578275,
    0.40821073
  ],
  "image_processor_type": "Qwen2VLImageProcessorFast",
  "image_std": [
    0.26862954,
    0.26130258,
    0.27577711
  ],
  "input_data_format": null,
  "max_pixels": 12845056,
  "merge_size": 2,
  "min_pixels": 3136,
  "pad_size": null,
  "patch_size": 14,
  "processor_class": "Qwen2_5_VLProcessor",
  "resample": 3,
  "rescale_factor": 0.00392156862745098,
  "return_tensors": null,
  "size": {
    "longest_edge": 12845056,
    "shortest_edge": 3136
  },
  "temporal_patch_size": 2
 }
--- a/tokenizer.json
+++ b/tokenizer.json
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,209 @@
 {
  "add_bos_token": false,
  "add_prefix_space": false,
  "added_tokens_decoder": {
    "151643": {
      "content": "<|endoftext|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151644": {
      "content": "<|im_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151645": {
      "content": "<|im_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151646": {
      "content": "<|object_ref_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151647": {
      "content": "<|object_ref_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151648": {
      "content": "<|box_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151649": {
      "content": "<|box_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151650": {
      "content": "<|quad_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151651": {
      "content": "<|quad_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151652": {
      "content": "<|vision_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151653": {
      "content": "<|vision_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151654": {
      "content": "<|vision_pad|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151655": {
      "content": "<|image_pad|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151656": {
      "content": "<|video_pad|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "151657": {
      "content": "<tool_call>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151658": {
      "content": "</tool_call>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151659": {
      "content": "<|fim_prefix|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151660": {
      "content": "<|fim_middle|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151661": {
      "content": "<|fim_suffix|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151662": {
      "content": "<|fim_pad|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151663": {
      "content": "<|repo_name|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "151664": {
      "content": "<|file_sep|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    }
  },
  "additional_special_tokens": [
    "<|im_start|>",
    "<|im_end|>",
    "<|object_ref_start|>",
    "<|object_ref_end|>",
    "<|box_start|>",
    "<|box_end|>",
    "<|quad_start|>",
    "<|quad_end|>",
    "<|vision_start|>",
    "<|vision_end|>",
    "<|vision_pad|>",
    "<|image_pad|>",
    "<|video_pad|>"
  ],
  "bos_token": null,
  "clean_up_tokenization_spaces": false,
  "eos_token": "<|im_end|>",
  "errors": "replace",
  "extra_special_tokens": {},
  "model_max_length": 131072,
  "pad_token": "<|endoftext|>",
  "padding_side": "right",
  "processor_class": "Qwen2_5_VLProcessor",
  "split_special_tokens": false,
  "tokenizer_class": "Qwen2Tokenizer",
  "unk_token": null
 }
--- a/vocab.json
+++ b/vocab.json