初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Faro-Yi-9B-DPO-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 01:37:12 +08:00
commit 3f31046ff0
27 changed files with 239 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ1_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO-f16.gguf filter=lfs diff=lfs merge=lfs -text
Faro-Yi-9B-DPO.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cf6a462499ea34c28f3d2c0664187b2f8d22b0f1a5af462933b8bb154da150e0
size 2181640608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:059af7259ca9bce41f56fb2774637066f42afe6d9c2fc2de9b351094364415f7
size 2014572960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d9bab9f2da1d6b2702bab32bd2a2c72673b479983b2abb6509f3e0294f78a297
size 3098112416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:df8e6edf64719033531a0c5c10fa67cd324d5bc8ac9d3339d27093247d7c3453
size 2875355552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:385b9da4044a0308481dda406387d995e2ff9f90ba9fdc18bc5962ad00fe51e1
size 2708009376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:29f44d924edd1ae62d91effd27db877cbb280b196e028f748ea130798b497af2
size 2460086688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f6285e7f1d0943484f7c0f0e0e373be3565496ed1b41a2ac82d1a0ce4da538a1
size 4055462304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15c6f8550c520e40a8da39a6406de0addc81c31d9fef8cef598ae4d115b3a908
size 3912577440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aba38421e3310f8014c083ae8cc52449961c4b4d040cc9f6edecdb2048372fd3
size 3717935520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:412fab2f1a5d8aab541f9626efa43490a3ec6fc17b13133e4b0b8626ff6818bd
size 3474321824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b130351ba19a592532d48047c67e201e6042fc56ffc90998ce90458be37b2533
size 5049577888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bc09a22866dd899d428670b36786fa3eef351dd63cd18b542365c63bd26f2780
size 4785009056

3
Faro-Yi-9B-DPO-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3647877dfd7b0c4e1e96e1f085c139ae74f0710b6faf114a780809e497109584
size 3354325408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6582b401e0da2cb8cd619ef66d53c41d0a91394e7df4366027973ea9b3951ed8
size 4690751904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:555f38ae89e4a6a06b053def964115a55eb9acc609c8308fd42077bd5c76bbba
size 4324405664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2f241dbdf318c59e897b961101a22351d2a1ab6c98d1d67ebccd4b7d31744f7a
size 3899208096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4cf89a113327b626feb8aa3a9d31aea144da8a52498dad84dd427bec2e40f08
size 5328957856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a17fa195c7cdebc95413f34a2bd2f3ce2672ba620092c210ea449b9a57e76256
size 5071860128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c47c4580683982a0aa7af3d46589ff1da17583798cbfc53441635aa27551821
size 6258258336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1ac105954c158b6f003772e7a855f9ec89e9d5a9ea43254e548c0a64e7e8a62e
size 6107853216

3
Faro-Yi-9B-DPO-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:933744905be3ef971bad747ae0f6d9ddd973154a2f0abe49ff6d66daf22c9c9f
size 7245640096

3
Faro-Yi-9B-DPO-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:759ee9cdfc6c86e6d689137a83d745ea145560773e4f7252afaad51f6cd05aa1
size 9383915936

3
Faro-Yi-9B-DPO-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7f6de0819e1f25a9038d262a54c682489b930a7dc4d497a4937b76106ecf33d8
size 17661112480

3
Faro-Yi-9B-DPO.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5003ad8be6d8ff37e243081db5801679eb97adc0dffee238029427169e934c04
size 6843289

107
README.md Normal file
View File

@@ -0,0 +1,107 @@
---
language:
- en
- zh
license: mit
datasets:
- wenbopan/Chinese-dpo-pairs
- Intel/orca_dpo_pairs
- argilla/ultrafeedback-binarized-preferences-cleaned
- jondurbin/truthy-dpo-v0.1
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Faro-Yi-9B-DPO
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b2965">b2965</a> for quantization.
Original model: https://huggingface.co/wenbopan/Faro-Yi-9B-DPO
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/b6ac44691e994344625687afe3263b3a)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Faro-Yi-9B-DPO-Q8_0.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q8_0.gguf) | Q8_0 | 9.38GB | Extremely high quality, generally unneeded but max available quant. |
| [Faro-Yi-9B-DPO-Q6_K.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q6_K.gguf) | Q6_K | 7.24GB | Very high quality, near perfect, *recommended*. |
| [Faro-Yi-9B-DPO-Q5_K_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q5_K_M.gguf) | Q5_K_M | 6.25GB | High quality, *recommended*. |
| [Faro-Yi-9B-DPO-Q5_K_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q5_K_S.gguf) | Q5_K_S | 6.10GB | High quality, *recommended*. |
| [Faro-Yi-9B-DPO-Q4_K_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q4_K_M.gguf) | Q4_K_M | 5.32GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Faro-Yi-9B-DPO-Q4_K_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q4_K_S.gguf) | Q4_K_S | 5.07GB | Slightly lower quality with more space savings, *recommended*. |
| [Faro-Yi-9B-DPO-IQ4_NL.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ4_NL.gguf) | IQ4_NL | 5.04GB | Decent quality, slightly smaller than Q4_K_S with similar performance *recommended*. |
| [Faro-Yi-9B-DPO-IQ4_XS.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ4_XS.gguf) | IQ4_XS | 4.78GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Faro-Yi-9B-DPO-Q3_K_L.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q3_K_L.gguf) | Q3_K_L | 4.69GB | Lower quality but usable, good for low RAM availability. |
| [Faro-Yi-9B-DPO-Q3_K_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q3_K_M.gguf) | Q3_K_M | 4.32GB | Even lower quality. |
| [Faro-Yi-9B-DPO-IQ3_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ3_M.gguf) | IQ3_M | 4.05GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Faro-Yi-9B-DPO-IQ3_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ3_S.gguf) | IQ3_S | 3.91GB | Lower quality, new method with decent performance, recommended over Q3_K_S quant, same size with better performance. |
| [Faro-Yi-9B-DPO-Q3_K_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q3_K_S.gguf) | Q3_K_S | 3.89GB | Low quality, not recommended. |
| [Faro-Yi-9B-DPO-IQ3_XS.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ3_XS.gguf) | IQ3_XS | 3.71GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Faro-Yi-9B-DPO-IQ3_XXS.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ3_XXS.gguf) | IQ3_XXS | 3.47GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Faro-Yi-9B-DPO-Q2_K.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-Q2_K.gguf) | Q2_K | 3.35GB | Very low quality but surprisingly usable. |
| [Faro-Yi-9B-DPO-IQ2_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ2_M.gguf) | IQ2_M | 3.09GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Faro-Yi-9B-DPO-IQ2_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ2_S.gguf) | IQ2_S | 2.87GB | Very low quality, uses SOTA techniques to be usable. |
| [Faro-Yi-9B-DPO-IQ2_XS.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ2_XS.gguf) | IQ2_XS | 2.70GB | Very low quality, uses SOTA techniques to be usable. |
| [Faro-Yi-9B-DPO-IQ2_XXS.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ2_XXS.gguf) | IQ2_XXS | 2.46GB | Lower quality, uses SOTA techniques to be usable. |
| [Faro-Yi-9B-DPO-IQ1_M.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ1_M.gguf) | IQ1_M | 2.18GB | Extremely low quality, *not* recommended. |
| [Faro-Yi-9B-DPO-IQ1_S.gguf](https://huggingface.co/bartowski/Faro-Yi-9B-DPO-GGUF/blob/main/Faro-Yi-9B-DPO-IQ1_S.gguf) | IQ1_S | 2.01GB | Extremely low quality, *not* recommended. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Faro-Yi-9B-DPO-GGUF --include "Faro-Yi-9B-DPO-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Faro-Yi-9B-DPO-GGUF --include "Faro-Yi-9B-DPO-Q8_0.gguf/*" --local-dir Faro-Yi-9B-DPO-Q8_0
```
You can either specify a new local-dir (Faro-Yi-9B-DPO-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}