初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Sailor2-1B-Chat-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-10-05 02:21:14 +08:00
commit d3effa4d55
29 changed files with 343 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat-f16.gguf filter=lfs diff=lfs merge=lfs -text
Sailor2-1B-Chat.imatrix filter=lfs diff=lfs merge=lfs -text

203
README.md Normal file
View File

@@ -0,0 +1,203 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
widget:
- text: 如何制作烤鱼?
example_title: Chinese
- text: How to bake fish?
example_title: English
- text: Bagaimana cara memanggang ikan?
example_title: Malay
- text: วิธีย่างปลา?
example_title: Thai
- text: Bagaimana membuat bakaran ikan?
example_title: Indonesian
- text: Làm thế nào để nướng cá?
example_title: Vietnamese
base_model: sail/Sailor2-1B-Chat
language:
- en
- zh
- id
- th
- vi
- ms
- lo
- my
- jv
- km
- su
- tl
license: apache-2.0
tags:
- multilingual
- sea
- sailor
- sft
- chat
- instruction
---
## Llamacpp imatrix Quantizations of Sailor2-1B-Chat
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4273">b4273</a> for quantization.
Original model: https://huggingface.co/sail/Sailor2-1B-Chat
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Sailor2-1B-Chat-f16.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-f16.gguf) | f16 | 1.98GB | false | Full F16 weights. |
| [Sailor2-1B-Chat-Q8_0.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q8_0.gguf) | Q8_0 | 1.06GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Sailor2-1B-Chat-Q6_K_L.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q6_K_L.gguf) | Q6_K_L | 1.01GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Sailor2-1B-Chat-Q6_K.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q6_K.gguf) | Q6_K | 1.01GB | false | Very high quality, near perfect, *recommended*. |
| [Sailor2-1B-Chat-Q5_K_L.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q5_K_L.gguf) | Q5_K_L | 0.83GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Sailor2-1B-Chat-Q5_K_M.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q5_K_M.gguf) | Q5_K_M | 0.79GB | false | High quality, *recommended*. |
| [Sailor2-1B-Chat-Q4_K_L.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_K_L.gguf) | Q4_K_L | 0.79GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Sailor2-1B-Chat-Q5_K_S.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q5_K_S.gguf) | Q5_K_S | 0.78GB | false | High quality, *recommended*. |
| [Sailor2-1B-Chat-Q4_K_M.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_K_M.gguf) | Q4_K_M | 0.74GB | false | Good quality, default size for most use cases, *recommended*. |
| [Sailor2-1B-Chat-Q3_K_XL.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q3_K_XL.gguf) | Q3_K_XL | 0.73GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Sailor2-1B-Chat-Q4_K_S.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_K_S.gguf) | Q4_K_S | 0.71GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Sailor2-1B-Chat-Q2_K_L.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q2_K_L.gguf) | Q2_K_L | 0.67GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Sailor2-1B-Chat-Q3_K_L.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q3_K_L.gguf) | Q3_K_L | 0.66GB | false | Lower quality but usable, good for low RAM availability. |
| [Sailor2-1B-Chat-Q3_K_M.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q3_K_M.gguf) | Q3_K_M | 0.64GB | false | Low quality. |
| [Sailor2-1B-Chat-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_0_8_8.gguf) | Q4_0_8_8 | 0.63GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [Sailor2-1B-Chat-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_0_4_8.gguf) | Q4_0_4_8 | 0.63GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [Sailor2-1B-Chat-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_0_4_4.gguf) | Q4_0_4_4 | 0.63GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [Sailor2-1B-Chat-Q4_0.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q4_0.gguf) | Q4_0 | 0.63GB | false | Legacy format, offers online repacking for ARM CPU inference. |
| [Sailor2-1B-Chat-IQ4_NL.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-IQ4_NL.gguf) | IQ4_NL | 0.63GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Sailor2-1B-Chat-IQ4_XS.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-IQ4_XS.gguf) | IQ4_XS | 0.62GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Sailor2-1B-Chat-IQ3_M.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-IQ3_M.gguf) | IQ3_M | 0.61GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Sailor2-1B-Chat-Q3_K_S.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q3_K_S.gguf) | Q3_K_S | 0.60GB | false | Low quality, not recommended. |
| [Sailor2-1B-Chat-IQ3_XS.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-IQ3_XS.gguf) | IQ3_XS | 0.60GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Sailor2-1B-Chat-Q2_K.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-Q2_K.gguf) | Q2_K | 0.60GB | false | Very low quality but surprisingly usable. |
| [Sailor2-1B-Chat-IQ2_M.gguf](https://huggingface.co/bartowski/Sailor2-1B-Chat-GGUF/blob/main/Sailor2-1B-Chat-IQ2_M.gguf) | IQ2_M | 0.58GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Sailor2-1B-Chat-GGUF --include "Sailor2-1B-Chat-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Sailor2-1B-Chat-GGUF --include "Sailor2-1B-Chat-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Sailor2-1B-Chat-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
New: Thanks to efforts made to have online repacking of weights in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921), you can now just use Q4_0 if your llama.cpp has been compiled for your ARM device.
Similarly, if you want to get slightly better performance, you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:da842c45b9fd1d74e9648be657fb8c52d362fb162ee8d9b9663c16bd84ee08f0
size 583190496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:021e55ce73be4e7fd00a2f07af3adfc3999250b20ade2383b9898e09906dc234
size 611500512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:23d63a415850562c3c4068c0f8910c015372af5e02bcb1efb0f29a7b9733caf2
size 603210720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:54382e2f05da93943d66a8cffd343f9b5892db00e59c995ceff5f1fbbdac52e2
size 631337952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f274323f906f7bb3928db47ba437750a1da39169dc0ba22516bb4eaec2399d6c
size 624800736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0fb06bde3d3a6da4637838847e19eb96e1e7c70f4b93be4a88f5178e9ab7741c
size 603210720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:275619d785084b0961146a2f30c2c4a90226a3383a897c640b941660393fde21
size 671278048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac7195b8833c564e1d0ec3e5cc62e9d3e92112d7680d68bb5bceac4d17dff71d
size 664712160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e875a1995d655ac781f936d6cc34fe00afe5b02e9849f1a11b5305959c628ca
size 637459424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6b5cb280e7cbe7d73ace031b562e3ca3cfc9b5a520c53c933aed76f0def305e
size 602522592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6a5d8ff30ecb329bcee052d84373890a212fefbc96bb71f25a9c0006cb4d023
size 732779488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0d58dfe5d5950f7fb30e985ed1c34e66f0abc693827b410ba2e3c65a32794bc9
size 631940064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7add916041f09af23b0b3d58dcbc500bd2e3349c136204c8023a1d90dc628ea1
size 630305760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:619d7751def16d604e82533d988521b98918581396bc888a96755b38c4ee7979
size 630305760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:78a99942d20ed601182f3af367c4950af8c88d776c267cbbf371af7eb58b0a9f
size 630305760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c50d2a9639ccd644f8f26dbb642a7ed48139b5a17cd3ccce84fe1031861df1df
size 789679072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:782e8abed13d51a2083eadfb2f6d94c2cd77940532f612a99e6f6bec9b3501d4
size 738628576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de895a0ddf4a27b178e5cc1ee1a39ff4bd2d79018b8454d836cb8d4f5984504e
size 713927648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d58465dc33f561eadf4b9121bccd63af968ae97ab5ae1f8be0f8951a0416702
size 834235360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec222df078725186fceaa08620bcd7ada1984c32f9234c65af3578504d2426ba
size 791693280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b26002af4e2656695d7f661232205402ef546fa96dd25912ec0903cafc110d2
size 776941536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:466d9a1b559e3cd728a59ee7cfc056278f9d1245cb6fd065f52b91e038ede333
size 1005536224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:466d9a1b559e3cd728a59ee7cfc056278f9d1245cb6fd065f52b91e038ede333
size 1005536224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:756c51aeb7636daa672d8ce5e312351384d7c7c2475b7763f62a7179c0074905
size 1056199648

3
Sailor2-1B-Chat-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a41d2eb757c12ec51d29058a4a8b6ef996a24ab069048ba7b420408ecfbcd958
size 1982376672

3
Sailor2-1B-Chat.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e3307fd5c50dc4d3b06a13cc42e1815825bc6fa255912b8fcb0c40e740f3cca
size 1977242

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}