初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Chronos-Gold-12B-1.0-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-22 03:52:06 +08:00
commit 7f3ecdd6cd
24 changed files with 241 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0-f16.gguf filter=lfs diff=lfs merge=lfs -text
Chronos-Gold-12B-1.0.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b759c79dda513b57d716257ef112379ce392db5a1d3254c0bc60510b14c45b8a
size 4435023328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c580656a93e81f8c53f0c4f63f45c26f6c579d78a77a9687d817712ebda3f28
size 5722232288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25f9a93a7e2d02bfdc6008e5e9ad41d56c74914a16ccb0f00febf2b242281d76
size 5306488288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ea6889e31fc3cc4c8f0e6e82e5b5459cfae2a05beec2920c4167d31b974f699
size 6742709728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa3dc6e734ec13ae24f059d2b57cafb56c82aa564a09a78290a9707404cd86b5
size 4791047648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a985788a8dfadf364680ea1cdc90f1462639a685bb4f1f0a3091a60254a47acd
size 5446407648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1f55cae48d6e6763cb3d51b95bc35f8bc33c9464cc6046ceff43136341e7202a
size 6561502688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de80d01912dc83800973f4ed3d29538f317fd22f992332f26e04b49ac78d8386
size 6083089888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b122e79b2bb18e6aafd145c95cbc197ea534440b66b9a73dc573eb8e187eda9a
size 5534225888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cdaa04885133ebb3b838232b945184103cc10a0fb5e4acb34254b336f132eaca
size 7148705248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:78ee51680484170f285012ca391aa2bc803e5f53a3a30fe56d23d940971a77ab
size 7975278048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c01c64b14afb5cd7ae7486915010a11da8bb019d49408dcc4b4908907ef991b5
size 7477204448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aebd6ee4ad1a8e7454ca44d33441b9fa74f7475433320b69f640f24700094842
size 7120197088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b1555030da04c68336e5b6738886da82ee4cc9430e4fca41212333f2297c342
size 9141818848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6bcc07dcb68629dfa9d87302b92807a709a6b7cee25c2941da21c77de0768553
size 8727631328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9a79e7bd7532c61bd6213e8ef810a57d95763848cac3c13bdf79e4a7f4fb878c
size 8518735328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ac7d9eb70d835c6cc4345f1e0796d365d385689167c67e38299f7f3c1439bf1
size 10056209888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c804af03b7b9bdb8d8809a710909382fd0028e542ba83eb4e89e76f77f3b7f91
size 10381268448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d7e73b5ce54d8546a90da5dcf4df5021185e90a1758f550cda8eebfd0a80d66
size 13022369248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:83dcd4f2051b382bb2072ff8f36a9768f1c0c3e5837abdadb812819f6c55d08b
size 24504276160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6256b03776adafdf873dccffa73a335e5029973ccbc89d48d0231074a1887818
size 7054418

121
README.md Normal file
View File

@@ -0,0 +1,121 @@
---
base_model: elinas/Chronos-Gold-12B-1.0
library_name: transformers
license: cc-by-nc-4.0
pipeline_tag: text-generation
tags:
- general-purpose
- roleplay
- storywriting
- merge
- finetune
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Chronos-Gold-12B-1.0
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3600">b3600</a> for quantization.
Original model: https://huggingface.co/elinas/Chronos-Gold-12B-1.0
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Chronos-Gold-12B-1.0-f16.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Chronos-Gold-12B-1.0-Q8_0.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Chronos-Gold-12B-1.0-Q6_K_L.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Chronos-Gold-12B-1.0-Q6_K.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Chronos-Gold-12B-1.0-Q5_K_L.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Chronos-Gold-12B-1.0-Q5_K_M.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Chronos-Gold-12B-1.0-Q5_K_S.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Chronos-Gold-12B-1.0-Q4_K_L.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Chronos-Gold-12B-1.0-Q4_K_M.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Chronos-Gold-12B-1.0-Q3_K_XL.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Chronos-Gold-12B-1.0-Q4_K_S.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Chronos-Gold-12B-1.0-IQ4_XS.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Chronos-Gold-12B-1.0-Q3_K_L.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Chronos-Gold-12B-1.0-Q3_K_M.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Chronos-Gold-12B-1.0-IQ3_M.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Chronos-Gold-12B-1.0-Q3_K_S.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Chronos-Gold-12B-1.0-Q2_K_L.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Chronos-Gold-12B-1.0-IQ3_XS.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Chronos-Gold-12B-1.0-Q2_K.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Chronos-Gold-12B-1.0-IQ2_M.gguf](https://huggingface.co/bartowski/Chronos-Gold-12B-1.0-GGUF/blob/main/Chronos-Gold-12B-1.0-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Chronos-Gold-12B-1.0-GGUF --include "Chronos-Gold-12B-1.0-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Chronos-Gold-12B-1.0-GGUF --include "Chronos-Gold-12B-1.0-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Chronos-Gold-12B-1.0-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}