初始化项目,由ModelHub XC社区提供模型
Model: Andycurrent/Qwen2.5-VL-7B-Abliterated-Caption-it_GGUF Source: Original Platform
This commit is contained in:
43
.gitattributes
vendored
Normal file
43
.gitattributes
vendored
Normal file
@@ -0,0 +1,43 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it.mmproj-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Qwen2.5-VL-3B-Abliterated-Caption-it_Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it.f16.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:97ac20a6196b233a1aa3ba945096bd6a006e01f151c5c6a6f7d218ce2130509f
|
||||||
|
size 6178316288
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it.mmproj-f16.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it.mmproj-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:e29a914431ac3a9491968511453eeca6ace0c97f14d81d2252effe5e140d75cc
|
||||||
|
size 1338428672
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q2_K.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q2_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:b0c5bec46e47efda03c827c6ca534f5917e785e411a814ff55ece735601b2741
|
||||||
|
size 1274755072
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q3_K_M.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q3_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:cee9fd975ff0cc02631ab3d858fa50c0e3248741326390b941693727df1012f6
|
||||||
|
size 1590474752
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q4_K_M.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q4_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:9005f43a8b0d5188de2f5de3baf3ed3e4f21a7df17ce634666928710211454f4
|
||||||
|
size 1929902080
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q5_K_M.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q5_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:2cc3ee8e17ad6bc1062009621d6848d3d2bcdf6b231bd25aec300cbe9cc7589e
|
||||||
|
size 2224814080
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q6_K.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:052c7afeb221a76a018f5f9751463356fb6cac80780fd777816750f202c8ad98
|
||||||
|
size 2538158080
|
||||||
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q8_0.gguf
Normal file
3
Qwen2.5-VL-3B-Abliterated-Caption-it_Q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:118f173879f7cdf24d5cffd943eb703f1b3141affc93a4703269a35cee155779
|
||||||
|
size 3285475328
|
||||||
72
README.md
Normal file
72
README.md
Normal file
@@ -0,0 +1,72 @@
|
|||||||
|
---
|
||||||
|
license: apache-2.0
|
||||||
|
language:
|
||||||
|
- en
|
||||||
|
- zh
|
||||||
|
base_model:
|
||||||
|
- Qwen/Qwen2.5-VL-7B-Instruct
|
||||||
|
tags:
|
||||||
|
- Image-to-text
|
||||||
|
- text-generation
|
||||||
|
- conversational
|
||||||
|
- uncensored
|
||||||
|
---
|
||||||
|
|
||||||
|
### Qwen2.5-VL-7B-Abliterated-Caption-it_GGUF(Vision Language)
|
||||||
|
|
||||||
|
This repository hosts Qwen2.5-VL-Abliterated-Caption-GGUF, a quantized Vision-Language (Uncensored) model optimized for image understanding and caption generation with relaxed alignment constraints. The model is designed for local inference, experimentation, and research-oriented multimodal workflows.
|
||||||
|
|
||||||
|
It targets users who want direct, descriptive visual reasoning without heavy content moderation layers, packaged in a GGUF format for efficient CPU and edge-device deployment.
|
||||||
|
|
||||||
|
### Model Summary
|
||||||
|
|
||||||
|
- **Model Identifier**: Qwen2.5-VL-Abliterated-Caption-GGUF
|
||||||
|
- **Base Model**: Qwen2.5-VL (Vision-Language)
|
||||||
|
- **Architecture**: Transformer-based multimodal model (text + vision)
|
||||||
|
- **Original model**: prithivMLmods/Qwen2.5-VL-Abliterated-Caption-GGUF
|
||||||
|
- **Primary Function**: Image captioning and visual-text understanding
|
||||||
|
|
||||||
|
###Purpose & Design Goals
|
||||||
|
|
||||||
|
This variant prioritizes expressive visual descriptions and caption accuracy while minimizing restrictive alignment behaviors. The “abliterated” aspect indicates reduced policy-driven refusals, making the model more suitable for:
|
||||||
|
- Dataset generation
|
||||||
|
- Visual analysis research
|
||||||
|
- Creative or descriptive captioning tasks
|
||||||
|
- Offline or private multimodal pipelines
|
||||||
|
|
||||||
|
### Multimodal Interaction Format
|
||||||
|
|
||||||
|
The model follows a standard multimodal prompt structure compatible with Qwen-VL style templates. A typical interaction may include system context, a user query, and an image reference:
|
||||||
|
```
|
||||||
|
<|system|>
|
||||||
|
You are a visual captioning assistant.
|
||||||
|
<|user|>
|
||||||
|
Describe the image in detail.
|
||||||
|
<|vision_input|>
|
||||||
|
<image>
|
||||||
|
<|assistant|>
|
||||||
|
```
|
||||||
|
|
||||||
|
### Core Capabilities
|
||||||
|
|
||||||
|
- Detailed and literal image captioning
|
||||||
|
- Multimodal reasoning over visual scenes
|
||||||
|
- Object, action, and context recognition
|
||||||
|
- Long-form descriptive outputs
|
||||||
|
- Reduced refusal behavior compared to safety-aligned VL models
|
||||||
|
- Optimized for local inference via GGUF
|
||||||
|
|
||||||
|
### Recommended Use Cases
|
||||||
|
|
||||||
|
- **Image caption generation** – datasets, tagging, annotation
|
||||||
|
- **Visual analysis** – scene breakdowns, object relationships
|
||||||
|
- **Creative workflows** – storytelling from images
|
||||||
|
- **Research & evaluation** – alignment and multimodal behavior testing
|
||||||
|
- **Offline deployments** – no cloud or API dependency
|
||||||
|
|
||||||
|
|
||||||
|
### Credits & Acknowledgements
|
||||||
|
|
||||||
|
- Qwen team for the base Qwen2.5-VL architecture
|
||||||
|
- GGUF tooling and local inference ecosystem contributors
|
||||||
|
- Open-source multimodal research community
|
||||||
Reference in New Issue
Block a user