初始化项目,由ModelHub XC社区提供模型

Model: bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-22 22:27:05 +08:00
commit eda477aefe
28 changed files with 264 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA-f16.gguf filter=lfs diff=lfs merge=lfs -text
LLAMA-3_8B_Unaligned_BETA.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:04b89155bd834134774ebf8c4dba0a141c7bbf475089d8ac425efb281fdf237d
size 2948290368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1f7c0c7adb9a0bbcc5a73475ba6ba11e34b7e7e78bc76e9765f446ec34b8e119
size 3784833920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fff0f90171a1b9206a8b0121065dbab1ccbb170625b97b2c84c1e22ffa62657
size 3518757760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e6928620daf4f456357b58818f61bff3d84907592f769d5a394df3601f5fd02
size 4447674048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e7d39c873dc3e6d43a04be4e1d35b02ea434c99b0f2cc50bce9f805c49f740ee
size 3179141184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ee4406344f58d965ebc932c3ffc5c825b7952a33824bc37388f10f9be59d969
size 3692173184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9bdc923d4b77c40c55346cadfb58907e080b1ace488334fa4df727649986fc8f
size 4321966976

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d87fa4d7b27a626e626df3dff6f57d923624d377982b98ba3578026945f4ab5
size 4018928512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:63fee745083853ad01c99accb40f07cd968e80baf27a6b95643d4426d66dcb0a
size 3664509824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:066cafd52542d9c427b3586e69d3f94144bf848984b557cffaf95d275b6ce431
size 4781643648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a16dc4ef9807fa2d7b38b80b32e5f46c967c89929c4dda017815228cff4acd14
size 4675903424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9580700375f8c29f2397a400ff9e7a8d974bf161ff180e60e0160f05ad01580
size 4661223360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ccfe21b4964acdeb57f87eb5a27911ce0a8e5284d63d49a4dec3f7e8aa8e572
size 4661223360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8b8d4d1bb2f42a9daaca050b69b57b0d0ccc225de908e111aeebd06540cae517
size 4661223360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:69c37763bfc0812acab977dde412fd1e47d88048a41af982630aa9e291c26272
size 5310650240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b88fb4537339996c04e4a1b6ef6a2d555c4103b6378e273ae9c6c5e77af67eb
size 4920745920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25fe3cfc6841195c181787f0e6f675f5da526b344d90773849b42922a2510e56
size 4692680640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e18f44eba73e926e336056ccb0c336c74074599310e1e50c0b187358534e544
size 6057236352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a1f18a291eb23ec68bcd32338843612c77a8e3d128ad68ec28227fa69f6f6920
size 5733000128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f813fb9102d65e9831cf11ad2d5fb2b9c6e0edcb313b323b743fb518fc4371b8
size 5599306688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb5b0907e0d2ea336bc3b805517333acc428f1d70eb1b1adceac8a1ee7cdf8a3
size 6596020224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a0c2d30e4e4ce065c1f7e2cda1eae276199d9335359f8e7b9b1b5ff817121c4d
size 6850484096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3f597194230da283019852fcd9137ca42de00164a420bee4ae29cf484190bec
size 8540788608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1a7543a48e17aeae5ecbeab90000e8c2b7f580aa09278da957f5ad6fe46657c0
size 16068924000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:81fb4ab9538790e5c558254b07fe8a001b223741bfd866f4d1839a64949d7377
size 4988170

128
README.md Normal file
View File

@@ -0,0 +1,128 @@
---
base_model: SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of LLAMA-3_8B_Unaligned_BETA
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3901">b3901</a> for quantization.
Original model: https://huggingface.co/SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## What's new:
Update tokenizer
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [LLAMA-3_8B_Unaligned_BETA-f16.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [LLAMA-3_8B_Unaligned_BETA-Q8_0.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [LLAMA-3_8B_Unaligned_BETA-Q6_K_L.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q6_K.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q5_K_L.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q5_K_M.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q5_K_S.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_K_L.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_K_M.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for must use cases, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q3_K_XL.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_K_S.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_0.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [LLAMA-3_8B_Unaligned_BETA-Q4_0_8_8.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_0_4_8.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [LLAMA-3_8B_Unaligned_BETA-Q4_0_4_4.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [LLAMA-3_8B_Unaligned_BETA-IQ4_XS.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [LLAMA-3_8B_Unaligned_BETA-Q3_K_L.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [LLAMA-3_8B_Unaligned_BETA-Q3_K_M.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [LLAMA-3_8B_Unaligned_BETA-IQ3_M.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [LLAMA-3_8B_Unaligned_BETA-Q2_K_L.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [LLAMA-3_8B_Unaligned_BETA-Q3_K_S.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [LLAMA-3_8B_Unaligned_BETA-IQ3_XS.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [LLAMA-3_8B_Unaligned_BETA-Q2_K.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [LLAMA-3_8B_Unaligned_BETA-IQ2_M.gguf](https://huggingface.co/bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF/blob/main/LLAMA-3_8B_Unaligned_BETA-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF --include "LLAMA-3_8B_Unaligned_BETA-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/LLAMA-3_8B_Unaligned_BETA-GGUF --include "LLAMA-3_8B_Unaligned_BETA-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (LLAMA-3_8B_Unaligned_BETA-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}