初始化项目,由ModelHub XC社区提供模型

Model: bartowski/PinkPixel_Crystal-Think-V2-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-21 03:46:13 +08:00
commit 762412361d
28 changed files with 301 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0434fd5c41be6a5143af7c50170e2bf38d212f2441595f15cc6641d6f07f9e93
size 1512980128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2276db7847bf8150df767634993fbea4c5ea603e84c309284af14a2479f9a74a
size 1962892448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b76e3bbf5687918a99025086ed6754843138ab358db1db335536deadad17664a
size 1814371488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f6c54a3de8e2a01b287bbebfb621fff202c61b84c532760ae53f7c043f027e86
size 1670184608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f895cdc438c05746d5de42c70f1ead3c67cd50d3de0eed75a7e9b598b29e05c4
size 2381339808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a65ee904797c387b3b18059b9fb02b1e8ae6c55c597c9cbf3b607f20a000cd0
size 2270747808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eb6fe4ad40be5ca55eefc05ac930990b52a12fd50733d087d1fc323ffbdf47dc
size 1669495968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e18aa3b1d170d5fb7cce663f3c8c74b079a57b25f5b5d0fee49a9996a74209be
size 1763696288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7de5e20472bb944d50eb109f7f4ab1fefa70542c1d218de4149f6e0c30634cf4
size 2239782048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d112c549bc3762ddaa85956946faaddffb62718593ba653f73e07f50fc357f3b
size 2075614368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa81559e592ad5d38d0e3a416f6edabc17d91db06001cbf08c95971049c9c2c4
size 1886993568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c972e706321505ab251ad47fed3db9165367d736f3844c77f631fedb44ee7649
size 2333982368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:51ecd04a759c7cf47c49074661e383c1f1b2bc935609ef066ab22a4f2ac514ba
size 2375769248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d58accd6b1f99dfe6fcf50206361a9f4d9cd43e6d0ed53e3c175b44072e95220
size 2596625568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe83fc6d2b06659b655e638ca3708857c4b783514e900ecbada81c273e03d366
size 2591477408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:10f2558089c90bc9ef8036ac0b1142ad8991902ec83840a00710fd654df19aaa
size 2497277088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:46b4eda34adb8c4a6e3db6d5423bb6f328d84c58d5161fff48de6515eaeb8d8d
size 2383305888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:81dadce5abe0a754b60c28972c108e3c4b3e84458e1623e4a4dfa43eaae3fc8c
size 2983710368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:145c08279142afb9a27ee1f6761b4501de8a2b1f5cd4b3a4e23929254a0bc1a6
size 2889510048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:824bb79950145bdb20aeab7773c74bbf21f43a1677e12fd394db1305ae81ed21
size 2823707808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:494a00d0cd8f31e5cc18f7aa1aa2e4e62c71c528fc6d1a138facf6141cb4ebf8
size 3306257568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e7a9b22b54fbb146241e858db50c3c0adec8bbef2594bf94ac82107c5ec5724c
size 3400457888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9dfcd3f17609dd346654c994644a9cfda963e1793d146791a6fc9f3e3ef7dd00
size 4280401568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:98fabc7e38283a3ddf4e60b3defbf6be05d26eea87355ba4596e14769d319915
size 8051281280

Binary file not shown.

181
README.md Normal file
View File

@@ -0,0 +1,181 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: PinkPixel/Crystal-Think-V2
base_model_relation: quantized
license: apache-2.0
language:
- en
tags:
- mathematical-reasoning
- qwen3
- lora
- grpo
- math
- reasoning
- fine-tuned
datasets:
- nvidia/OpenMathReasoning
---
## Llamacpp imatrix Quantizations of Crystal-Think-V2 by PinkPixel
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b5760">b5760</a> for quantization.
Original model: https://huggingface.co/PinkPixel/Crystal-Think-V2
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
No chat template specified so default is used. This may be incorrect, check original model card for details.
```
{system_prompt}<|endoftext|>{prompt}<|endoftext|><think>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Crystal-Think-V2-bf16.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-bf16.gguf) | bf16 | 8.05GB | false | Full BF16 weights. |
| [Crystal-Think-V2-Q8_0.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q8_0.gguf) | Q8_0 | 4.28GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Crystal-Think-V2-Q6_K_L.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q6_K_L.gguf) | Q6_K_L | 3.40GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Crystal-Think-V2-Q6_K.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q6_K.gguf) | Q6_K | 3.31GB | false | Very high quality, near perfect, *recommended*. |
| [Crystal-Think-V2-Q5_K_L.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q5_K_L.gguf) | Q5_K_L | 2.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Crystal-Think-V2-Q5_K_M.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q5_K_M.gguf) | Q5_K_M | 2.89GB | false | High quality, *recommended*. |
| [Crystal-Think-V2-Q5_K_S.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q5_K_S.gguf) | Q5_K_S | 2.82GB | false | High quality, *recommended*. |
| [Crystal-Think-V2-Q4_1.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q4_1.gguf) | Q4_1 | 2.60GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Crystal-Think-V2-Q4_K_L.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q4_K_L.gguf) | Q4_K_L | 2.59GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Crystal-Think-V2-Q4_K_M.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q4_K_M.gguf) | Q4_K_M | 2.50GB | false | Good quality, default size for most use cases, *recommended*. |
| [Crystal-Think-V2-Q4_K_S.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q4_K_S.gguf) | Q4_K_S | 2.38GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Crystal-Think-V2-Q4_0.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q4_0.gguf) | Q4_0 | 2.38GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Crystal-Think-V2-IQ4_NL.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ4_NL.gguf) | IQ4_NL | 2.38GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Crystal-Think-V2-Q3_K_XL.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q3_K_XL.gguf) | Q3_K_XL | 2.33GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Crystal-Think-V2-IQ4_XS.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ4_XS.gguf) | IQ4_XS | 2.27GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Crystal-Think-V2-Q3_K_L.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q3_K_L.gguf) | Q3_K_L | 2.24GB | false | Lower quality but usable, good for low RAM availability. |
| [Crystal-Think-V2-Q3_K_M.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q3_K_M.gguf) | Q3_K_M | 2.08GB | false | Low quality. |
| [Crystal-Think-V2-IQ3_M.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ3_M.gguf) | IQ3_M | 1.96GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Crystal-Think-V2-Q3_K_S.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q3_K_S.gguf) | Q3_K_S | 1.89GB | false | Low quality, not recommended. |
| [Crystal-Think-V2-IQ3_XS.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ3_XS.gguf) | IQ3_XS | 1.81GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Crystal-Think-V2-Q2_K_L.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q2_K_L.gguf) | Q2_K_L | 1.76GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Crystal-Think-V2-IQ3_XXS.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ3_XXS.gguf) | IQ3_XXS | 1.67GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Crystal-Think-V2-Q2_K.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-Q2_K.gguf) | Q2_K | 1.67GB | false | Very low quality but surprisingly usable. |
| [Crystal-Think-V2-IQ2_M.gguf](https://huggingface.co/bartowski/PinkPixel_Crystal-Think-V2-GGUF/blob/main/PinkPixel_Crystal-Think-V2-IQ2_M.gguf) | IQ2_M | 1.51GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/PinkPixel_Crystal-Think-V2-GGUF --include "PinkPixel_Crystal-Think-V2-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/PinkPixel_Crystal-Think-V2-GGUF --include "PinkPixel_Crystal-Think-V2-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (PinkPixel_Crystal-Think-V2-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}