初始化项目,由ModelHub XC社区提供模型

Model: bartowski/YuLan-Mini-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 15:22:07 +08:00
commit 5587f10953
28 changed files with 344 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-f16.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini-f32.gguf filter=lfs diff=lfs merge=lfs -text
YuLan-Mini.imatrix filter=lfs diff=lfs merge=lfs -text

208
README.md Normal file
View File

@@ -0,0 +1,208 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
- zh
datasets:
- HuggingFaceFW/fineweb-edu
- bigcode/the-stack-v2
- mlfoundations/dclm-baseline-1.0
- math-ai/AutoMathText
- gair-prox/open-web-math-pro
- RUC-AIBOX/long_form_thought_data_5k
- internlm/Lean-Workbook
- internlm/Lean-Github
- deepseek-ai/DeepSeek-Prover-V1
- ScalableMath/Lean-STaR-base
- ScalableMath/Lean-STaR-plus
- ScalableMath/Lean-CoT-base
- ScalableMath/Lean-CoT-plus
- opencsg/chinese-fineweb-edu
- liwu/MNBVC
- vikp/textbook_quality_programming
- HuggingFaceTB/smollm-corpus
- OpenCoder-LLM/opc-annealing-corpus
- OpenCoder-LLM/opc-sft-stage1
- OpenCoder-LLM/opc-sft-stage2
- XinyaoHu/AMPS_mathematica
- deepmind/math_dataset
- mrfakename/basic-math-10m
- microsoft/orca-math-word-problems-200k
- AI-MO/NuminaMath-CoT
- HuggingFaceTB/cosmopedia
- MU-NLPC/Calc-ape210k
- manu/project_gutenberg
- storytracer/LoC-PD-Books
- allenai/dolma
base_model: yulan-team/YuLan-Mini
license: mit
---
## Llamacpp imatrix Quantizations of YuLan-Mini
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4381">b4381</a> for quantization.
Original model: https://huggingface.co/yulan-team/YuLan-Mini
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>
<|start_header_id|> system<|end_header_id|>
{system_prompt}<|eot_id|>
<|start_header_id|> user<|end_header_id|>
{prompt}<|eot_id|>
<|start_header_id|> assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [YuLan-Mini-f32.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-f32.gguf) | f32 | 9.70GB | false | Full F32 weights. |
| [YuLan-Mini-f16.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-f16.gguf) | f16 | 4.85GB | false | Full F16 weights. |
| [YuLan-Mini-Q8_0.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q8_0.gguf) | Q8_0 | 2.58GB | false | Extremely high quality, generally unneeded but max available quant. |
| [YuLan-Mini-Q6_K_L.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q6_K_L.gguf) | Q6_K_L | 2.58GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [YuLan-Mini-Q6_K.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q6_K.gguf) | Q6_K | 2.58GB | false | Very high quality, near perfect, *recommended*. |
| [YuLan-Mini-Q5_K_L.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q5_K_L.gguf) | Q5_K_L | 2.03GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [YuLan-Mini-Q5_K_M.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q5_K_M.gguf) | Q5_K_M | 1.97GB | false | High quality, *recommended*. |
| [YuLan-Mini-Q4_K_L.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q4_K_L.gguf) | Q4_K_L | 1.92GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [YuLan-Mini-Q5_K_S.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q5_K_S.gguf) | Q5_K_S | 1.88GB | false | High quality, *recommended*. |
| [YuLan-Mini-Q4_K_M.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q4_K_M.gguf) | Q4_K_M | 1.85GB | false | Good quality, default size for most use cases, *recommended*. |
| [YuLan-Mini-Q4_K_S.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q4_K_S.gguf) | Q4_K_S | 1.75GB | false | Slightly lower quality with more space savings, *recommended*. |
| [YuLan-Mini-Q3_K_XL.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q3_K_XL.gguf) | Q3_K_XL | 1.70GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [YuLan-Mini-Q3_K_L.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q3_K_L.gguf) | Q3_K_L | 1.61GB | false | Lower quality but usable, good for low RAM availability. |
| [YuLan-Mini-Q4_1.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q4_1.gguf) | Q4_1 | 1.60GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [YuLan-Mini-Q3_K_M.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q3_K_M.gguf) | Q3_K_M | 1.56GB | false | Low quality. |
| [YuLan-Mini-Q2_K_L.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q2_K_L.gguf) | Q2_K_L | 1.56GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [YuLan-Mini-IQ3_M.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-IQ3_M.gguf) | IQ3_M | 1.50GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [YuLan-Mini-Q4_0.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q4_0.gguf) | Q4_0 | 1.47GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [YuLan-Mini-IQ4_NL.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-IQ4_NL.gguf) | IQ4_NL | 1.47GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [YuLan-Mini-IQ4_XS.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-IQ4_XS.gguf) | IQ4_XS | 1.47GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [YuLan-Mini-IQ3_XS.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-IQ3_XS.gguf) | IQ3_XS | 1.47GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [YuLan-Mini-Q2_K.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q2_K.gguf) | Q2_K | 1.47GB | false | Very low quality but surprisingly usable. |
| [YuLan-Mini-IQ2_M.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-IQ2_M.gguf) | IQ2_M | 1.47GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [YuLan-Mini-Q3_K_S.gguf](https://huggingface.co/bartowski/YuLan-Mini-GGUF/blob/main/YuLan-Mini-Q3_K_S.gguf) | Q3_K_S | 1.46GB | false | Low quality, not recommended. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/YuLan-Mini-GGUF --include "YuLan-Mini-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/YuLan-Mini-GGUF --include "YuLan-Mini-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (YuLan-Mini-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

3
YuLan-Mini-IQ2_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25f9d9a80187529b3d3e9a4a6be390439f0a409ecac9563579f90433b2250fbc
size 1467847456

3
YuLan-Mini-IQ3_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92b1b1217f70441ac1c2c1604fc8968688c30785a69f9634734c05fd85e1ac9d
size 1501716256

3
YuLan-Mini-IQ3_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:df111a3c5654e9bc9f4fc39228906ccbc7b4d2a6c209d02b04c2eba3ff0e81c9
size 1467847456

3
YuLan-Mini-IQ4_NL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d6c0f3f8bb4c257eb4f3717c60312d439cc1ac6dea3469f8d0040c2d7afacf9
size 1470427936

3
YuLan-Mini-IQ4_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:161b64b8962205dcd51352c8bc3bc1bf7014260ab0a64404e0de20c68f830a5b
size 1470427936

3
YuLan-Mini-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a4e9eecf149b3f259b2ef2fc9d32c64993f3a8dc11c77cba590b7ab16510988
size 1467847456

3
YuLan-Mini-Q2_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2280d37082a0c1cb453f5f301455c73ed1aa1b14c29e716023fed66fa9fa4351
size 1562887456

3
YuLan-Mini-Q3_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:be561130ff83804a87a777f4b4e7babea5a0c736a031365066974e3a020c0dad
size 1605903136

3
YuLan-Mini-Q3_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4eaa7b5f1821ab2ed76c66b162ca282850136fb58f4cd6e77bf070ec310ded5e
size 1559984416

3
YuLan-Mini-Q3_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:450345556b6cf83b9819e056185b5309fff8ada27bef99843911b9dc82690bf5
size 1462686496

3
YuLan-Mini-Q3_K_XL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e12cd9edb7d6ee1364dc0917b24aff36ee353f8517f3a9979314a2ec6da82f3
size 1700943136

3
YuLan-Mini-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eb1f8b194894e16264e4974dda3d36798a5d5ef0b0d85cad2590e055093c388b
size 1466718496

3
YuLan-Mini-Q4_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b59ab06c025f540748d91b86d3673a460036a051c6156a04c1f6053e9e3c51d7
size 1602300256

3
YuLan-Mini-Q4_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f5be5434fb52a958caf63189b88fe535a7fa1e9031797e5a4430b0b39b8fd41
size 1917703456

3
YuLan-Mini-Q4_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f849d3d0533548cf2c4a6b3dc68a0f21765d9e262955f3b7e0f224fc765292e0
size 1846423456

3
YuLan-Mini-Q4_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:164fe3613a0f542255f6169335f8343e63082ccbac9781f03e4733d2c155d46d
size 1746130336

3
YuLan-Mini-Q5_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:270835421cd2fd230b4d8f20ffa85e3139ffdc60d2041ce474b2f689ce7f4694
size 2028018976

3
YuLan-Mini-Q5_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4b8ebf328e6a2cd23d9b94fcc42fbf4d38276b58adeeb03774981e0cdf2840ad
size 1968618976

3
YuLan-Mini-Q5_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2b498660b2b4eb34262334ee8861caa6e55a9197d3f41ce7942056478540367f
size 1881527776

3
YuLan-Mini-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4dc84d11b6c348602d0f0b57db0f01a2e5641079e3fceb3c9b7ee34fa9a8b458
size 2579596576

3
YuLan-Mini-Q6_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4dc84d11b6c348602d0f0b57db0f01a2e5641079e3fceb3c9b7ee34fa9a8b458
size 2579596576

3
YuLan-Mini-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:94d59c172312467b63eee6bf00eb65deea89afe380aceedc10cef2972d2c8ba1
size 2579596576

3
YuLan-Mini-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6637894025069ba7d2a38a112cb3396cdb73cd5e47a7e74f5ce604b8fe9f5bf6
size 4852002976

3
YuLan-Mini-f32.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:44c8a78173a86c4017f66b271dead0935d67f29c12d0662559dbf83e0ef2ec68
size 9699803040

3
YuLan-Mini.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:453603cef9299717cd5de9928e5e400750573997d68d550212ad0303657ee4c2
size 3668706

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}