初始化项目,由ModelHub XC社区提供模型

Model: bartowski/FastApply-7B-v1.0-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-14 15:10:13 +08:00
commit 75762ee7ff
28 changed files with 272 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0-f16.gguf filter=lfs diff=lfs merge=lfs -text
FastApply-7B-v1.0.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:72febdb27022868f1ed8bd9fb878de1906b376c72f43c4c3732ea5101684a21a
size 2780341248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61b3c657a3b0070a1487b0b485740e4a1a8358a5742927100f8dbc6794c7d52d
size 3574010880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0d624616fe0b4e67a488410a483dadcf6969ef9c8ee10c8d1b0342cab1374c0a
size 3346254848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c758e7f7c1788e4e491f5d964dfd0bc59cbf37c1e02276483f047c48d3b79fd
size 4218471424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a93cef8f2aac14fd386292f9ceb0a671add232c7ea4d4ed8890cfce69c252e0c
size 3015939072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f097b3f2d3b00f73380c5ed81a50f6f6ec18cd56b30c9e12e4a542346af17d9
size 3548163072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d0cc63911e999c9126c44876f261d0872b963e4e9c6aa8a8f22586ebd0b17645
size 4088458240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5703d4179db73a722fe591619627e33ed6d44ed564fb0e02b14328254cc68c2b
size 3808390144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:02adda7cc87468a957cba923843585e3652a6ab8354e6b857ab324b6bd8b5fb3
size 3492367360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ad80c7a9c4c70fa2bab7a4fd311688c856ec80281268a2c83ea1adc78a67fa8
size 4565330944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:27bc7a1859525c3b15b7b893aff7580856805f54cd693e93c98e2adef3d74520
size 4444120064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cfb60eee062a9f1aff22a5e373f88ec7fe15ae586c56824eb0305bcb84c87de4
size 4431389696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d04cebc2cd8f396267a4f6511922be6d5fa20d4885f8fdf9d2498da0e056fdd
size 4431389696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d64bd4ac6b68ce1b713c92ad0fb63ac9275e63fa52a32c361a0306fe084f4246
size 4431389696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3cd2b8bea1317aeb1c2e1bf5fbc73252ed08a3b4afd0e4e5821ca53b17051a63
size 5087562752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:da7536b434c36349a962bc0ce37990e2b07c3e1f73535d644ef734159a952013
size 4683072512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b9c5e4ceba73ba0854862a3056eac3233819d18668fe8d776ff6826743a4cfbe
size 4457767936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:507fc34f93291ecf7a63d6ffc0fcf8625b4113ebc46271fbbd1bd9bd0a853c6e
size 5781195776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f826f91f020d652e8d00368fcc88018b2e7c395c6327a8a4b8c4212f6e75a6da
size 5444830208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a17307f9291f4a615cb7042e49ccb0dc7ece434ed10e5612aae7cb63c28dcf6
size 5315175424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2281de622b358740afb8e8b9334efbb50359006db54a5527c7d49a6747aafcda
size 6254197760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b32c92dfe3df83efa2a9bfa0ddde5f7bcbec216cbebeaf7f3954156517d65caf
size 6518180864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:22489b4a57815f97aef7b52eaa6ddf6b4f6fd91466951bd541ce5bdea719752f
size 8098524160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4437f6457b3cfd39269061f4e438f248c65c4f72cd69848a22f7f9c670cc2fc3
size 15237851872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0dcd40c5b648f5c23a0dc72874cc22dcb58d69cc59615d8e71b949293a101557
size 4536678

136
README.md Normal file
View File

@@ -0,0 +1,136 @@
---
base_model: Kortix/FastApply-7B-v1.0
language:
- en
license: apache-2.0
pipeline_tag: text-generation
tags:
- text-generation-inference
- transformers
- unsloth
- qwen2
- trl
- sft
- fast-apply
- instant-apply
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of FastApply-7B-v1.0
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3972">b3972</a> for quantization.
Original model: https://huggingface.co/Kortix/FastApply-7B-v1.0
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [FastApply-7B-v1.0-f16.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-f16.gguf) | f16 | 15.24GB | false | Full F16 weights. |
| [FastApply-7B-v1.0-Q8_0.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q8_0.gguf) | Q8_0 | 8.10GB | false | Extremely high quality, generally unneeded but max available quant. |
| [FastApply-7B-v1.0-Q6_K_L.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q6_K_L.gguf) | Q6_K_L | 6.52GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [FastApply-7B-v1.0-Q6_K.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q6_K.gguf) | Q6_K | 6.25GB | false | Very high quality, near perfect, *recommended*. |
| [FastApply-7B-v1.0-Q5_K_L.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q5_K_L.gguf) | Q5_K_L | 5.78GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [FastApply-7B-v1.0-Q5_K_M.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q5_K_M.gguf) | Q5_K_M | 5.44GB | false | High quality, *recommended*. |
| [FastApply-7B-v1.0-Q5_K_S.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q5_K_S.gguf) | Q5_K_S | 5.32GB | false | High quality, *recommended*. |
| [FastApply-7B-v1.0-Q4_K_L.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_K_L.gguf) | Q4_K_L | 5.09GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [FastApply-7B-v1.0-Q4_K_M.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_K_M.gguf) | Q4_K_M | 4.68GB | false | Good quality, default size for must use cases, *recommended*. |
| [FastApply-7B-v1.0-Q3_K_XL.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q3_K_XL.gguf) | Q3_K_XL | 4.57GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [FastApply-7B-v1.0-Q4_K_S.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_K_S.gguf) | Q4_K_S | 4.46GB | false | Slightly lower quality with more space savings, *recommended*. |
| [FastApply-7B-v1.0-Q4_0.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_0.gguf) | Q4_0 | 4.44GB | false | Legacy format, generally not worth using over similarly sized formats |
| [FastApply-7B-v1.0-Q4_0_8_8.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.43GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [FastApply-7B-v1.0-Q4_0_4_8.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.43GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [FastApply-7B-v1.0-Q4_0_4_4.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.43GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [FastApply-7B-v1.0-IQ4_XS.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-IQ4_XS.gguf) | IQ4_XS | 4.22GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [FastApply-7B-v1.0-Q3_K_L.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q3_K_L.gguf) | Q3_K_L | 4.09GB | false | Lower quality but usable, good for low RAM availability. |
| [FastApply-7B-v1.0-Q3_K_M.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q3_K_M.gguf) | Q3_K_M | 3.81GB | false | Low quality. |
| [FastApply-7B-v1.0-IQ3_M.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-IQ3_M.gguf) | IQ3_M | 3.57GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [FastApply-7B-v1.0-Q2_K_L.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q2_K_L.gguf) | Q2_K_L | 3.55GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [FastApply-7B-v1.0-Q3_K_S.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q3_K_S.gguf) | Q3_K_S | 3.49GB | false | Low quality, not recommended. |
| [FastApply-7B-v1.0-IQ3_XS.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-IQ3_XS.gguf) | IQ3_XS | 3.35GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [FastApply-7B-v1.0-Q2_K.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-Q2_K.gguf) | Q2_K | 3.02GB | false | Very low quality but surprisingly usable. |
| [FastApply-7B-v1.0-IQ2_M.gguf](https://huggingface.co/bartowski/FastApply-7B-v1.0-GGUF/blob/main/FastApply-7B-v1.0-IQ2_M.gguf) | IQ2_M | 2.78GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/FastApply-7B-v1.0-GGUF --include "FastApply-7B-v1.0-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/FastApply-7B-v1.0-GGUF --include "FastApply-7B-v1.0-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (FastApply-7B-v1.0-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}