初始化项目,由ModelHub XC社区提供模型

Model: bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-22 12:00:06 +08:00
commit b93df9b3e6
28 changed files with 328 additions and 0 deletions

73
.gitattributes vendored Normal file
View File

@@ -0,0 +1,73 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-bf16.gguf filter=lfs diff=lfs merge=lfs -text
p-e-w_Qwen3-4B-Instruct-2507-heretic-imatrix.gguf filter=lfs diff=lfs merge=lfs -text

179
README.md Normal file
View File

@@ -0,0 +1,179 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: p-e-w/Qwen3-4B-Instruct-2507-heretic
license: apache-2.0
tags:
- heretic
- uncensored
- decensored
- abliterated
license_link: https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/blob/main/LICENSE
---
## Llamacpp imatrix Quantizations of Qwen3-4B-Instruct-2507-heretic by p-e-w
Using <a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggml-org/llama.cpp/releases/tag/b7049">b7049</a> for quantization.
Original model: https://huggingface.co/p-e-w/Qwen3-4B-Instruct-2507-heretic
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8) combined with a subset of combined_all_small.parquet from Ed Addario [here](https://huggingface.co/datasets/eaddario/imatrix-calibration/blob/main/combined_all_small.parquet)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggml-org/llama.cpp), or any other llama.cpp based project
## Prompt format
No chat template specified so default is used. This may be incorrect, check original model card for details.
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen3-4B-Instruct-2507-heretic-bf16.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-bf16.gguf) | bf16 | 8.05GB | false | Full BF16 weights. |
| [Qwen3-4B-Instruct-2507-heretic-Q8_0.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q8_0.gguf) | Q8_0 | 4.28GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen3-4B-Instruct-2507-heretic-Q6_K_L.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q6_K_L.gguf) | Q6_K_L | 3.40GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q6_K.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q6_K.gguf) | Q6_K | 3.31GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q5_K_L.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_L.gguf) | Q5_K_L | 2.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q5_K_M.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_M.gguf) | Q5_K_M | 2.89GB | false | High quality, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q5_K_S.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q5_K_S.gguf) | Q5_K_S | 2.82GB | false | High quality, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q4_1.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_1.gguf) | Q4_1 | 2.60GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Qwen3-4B-Instruct-2507-heretic-Q4_K_L.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_L.gguf) | Q4_K_L | 2.59GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q4_K_M.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_M.gguf) | Q4_K_M | 2.50GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q4_K_S.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_S.gguf) | Q4_K_S | 2.38GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q4_0.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_0.gguf) | Q4_0 | 2.38GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Qwen3-4B-Instruct-2507-heretic-IQ4_NL.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ4_NL.gguf) | IQ4_NL | 2.38GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Qwen3-4B-Instruct-2507-heretic-Q3_K_XL.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_XL.gguf) | Q3_K_XL | 2.33GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen3-4B-Instruct-2507-heretic-IQ4_XS.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ4_XS.gguf) | IQ4_XS | 2.27GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen3-4B-Instruct-2507-heretic-Q3_K_L.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_L.gguf) | Q3_K_L | 2.24GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen3-4B-Instruct-2507-heretic-Q3_K_M.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_M.gguf) | Q3_K_M | 2.08GB | false | Low quality. |
| [Qwen3-4B-Instruct-2507-heretic-IQ3_M.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_M.gguf) | IQ3_M | 1.96GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen3-4B-Instruct-2507-heretic-Q3_K_S.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q3_K_S.gguf) | Q3_K_S | 1.89GB | false | Low quality, not recommended. |
| [Qwen3-4B-Instruct-2507-heretic-IQ3_XS.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_XS.gguf) | IQ3_XS | 1.81GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen3-4B-Instruct-2507-heretic-Q2_K_L.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q2_K_L.gguf) | Q2_K_L | 1.76GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen3-4B-Instruct-2507-heretic-IQ3_XXS.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ3_XXS.gguf) | IQ3_XXS | 1.67GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Qwen3-4B-Instruct-2507-heretic-Q2_K.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-Q2_K.gguf) | Q2_K | 1.67GB | false | Very low quality but surprisingly usable. |
| [Qwen3-4B-Instruct-2507-heretic-IQ2_M.gguf](https://huggingface.co/bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF/blob/main/p-e-w_Qwen3-4B-Instruct-2507-heretic-IQ2_M.gguf) | IQ2_M | 1.51GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF --include "p-e-w_Qwen3-4B-Instruct-2507-heretic-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/p-e-w_Qwen3-4B-Instruct-2507-heretic-GGUF --include "p-e-w_Qwen3-4B-Instruct-2507-heretic-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (p-e-w_Qwen3-4B-Instruct-2507-heretic-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggml-org/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggml-org/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggml-org/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggml-org/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4e36a40601b35fb543397b6448f219ecd74c286584a09f6367589f8fe3235cf
size 1512982464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c46c7eab6aaf423d072d14aef842db49c9c362d0452442293a75c812a4ef6f5
size 1962894784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ee6391d19e3ccf632953b058620b2129c2a0e8353b6122f5254aadc9cf7fb595
size 1814373824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:35728751e3a6732e2a3588db91fdc3d060733654180a63a563064e1106b2cc22
size 1670186944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:471c0e85daa66c748d2a07d742a2bb83e5841ddab2d02b961fd275c33e954a65
size 2381342144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12ba6afc1790631cc5b6e6b1412011ef0cdb78171922e2932d04e861689280a7
size 2270750144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6161832af29bc7138542d74a84a0e220e13dbeb36e27ce9deed594f04f75a98c
size 1669498304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ae5f3e2a68b78abb75f5520bf98eb9cf31877950d5e855b21d3850775f4457c
size 1763698624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:70e4b7cc5244f15f9083be8bd94a50c21481d75bab38b848a13f2f53a2fa5eaf
size 2239784384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:841e704f85f318f0cf361cb6a3cd12853f0ec5d2eecc34b4eb95547f916db304
size 2075616704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9559d48fad8195d75dd0d9e2474bab8bd9be2aecb8002300e5f910394207cea5
size 1886995904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c785c7803bce3f5275c8d45563ce02f23dcac404c00667456c096687aceca561
size 2333984704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:06e827b60180825f16f4921a7bd62ae0b5b9d04433a4928685fc82ff558cc839
size 2375771584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb18ba2e43641c06fa0ed13fe959e0704f010e4d76ce30d7fd745776f0a72b46
size 2596627904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ead716d0dbfeab21f55a740639f21584ee5a5f5c76706b9585c7f4454eb84876
size 2591479744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d4d7fcd11d36495fa340f91b514ae3a47f5da21cf27c158d620d736e3bb35a61
size 2497279424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ea7a24f56a56673fc5201a5c2aa4d4417f718e6ea64ecc74725ab9d33b28fca
size 2383308224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa5cdec387da647220b9b8aee13add64d2debb4567722569aedb5f89f124843e
size 2983712704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b31a65fbaafaf54830e94295cd339cb9b00a215bc584280607e794d32b6975a
size 2889512384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:50462f2886b51d04624329741544cca014cb684472d362ad7d1b04703c35e9a6
size 2823710144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:53f4fd5b191fcca42e5fc45ec1dcf2f6ab6e6646c557b2efec63dcfcc56e5ee3
size 3306259904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6cd4265204694f9adb76ccff7b07a541b74dfa53eff7d0f3aac65ad194473613
size 3400460224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b990e208f13616ec08d6e9eb1a5ac785582f0e8a48ee1e22238efb3fe4e986e
size 4280403904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:582bb56e63b967d695a74626fa5ad260953935ee2b80e3dbb75cbb3e37a2bb3c
size 8051283584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9808c7178a992e52b2024023025a8390b4496b5793d69ec25e85727ba41c819
size 3872640