初始化项目,由ModelHub XC社区提供模型

Model: bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 09:24:06 +08:00
commit cd14844da2
28 changed files with 287 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

164
README.md Normal file
View File

@@ -0,0 +1,164 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: TheDrummer/Gemma-3-R1-4B-v1
---
## Llamacpp imatrix Quantizations of Gemma-3-R1-4B-v1 by TheDrummer
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b6139">b6139</a> for quantization.
Original model: https://huggingface.co/TheDrummer/Gemma-3-R1-4B-v1
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8) combined with a subset of combined_all_small.parquet from Ed Addario [here](https://huggingface.co/datasets/eaddario/imatrix-calibration/blob/main/combined_all_small.parquet)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Gemma-3-R1-4B-v1-bf16.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-bf16.gguf) | bf16 | 7.77GB | false | Full BF16 weights. |
| [Gemma-3-R1-4B-v1-Q8_0.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q8_0.gguf) | Q8_0 | 4.13GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Gemma-3-R1-4B-v1-Q6_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q6_K_L.gguf) | Q6_K_L | 3.35GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Gemma-3-R1-4B-v1-Q6_K.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q6_K.gguf) | Q6_K | 3.19GB | false | Very high quality, near perfect, *recommended*. |
| [Gemma-3-R1-4B-v1-Q5_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q5_K_L.gguf) | Q5_K_L | 2.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Gemma-3-R1-4B-v1-Q5_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q5_K_M.gguf) | Q5_K_M | 2.83GB | false | High quality, *recommended*. |
| [Gemma-3-R1-4B-v1-Q5_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q5_K_S.gguf) | Q5_K_S | 2.76GB | false | High quality, *recommended*. |
| [Gemma-3-R1-4B-v1-Q4_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q4_K_L.gguf) | Q4_K_L | 2.65GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Gemma-3-R1-4B-v1-Q4_1.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q4_1.gguf) | Q4_1 | 2.56GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Gemma-3-R1-4B-v1-Q4_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q4_K_M.gguf) | Q4_K_M | 2.49GB | false | Good quality, default size for most use cases, *recommended*. |
| [Gemma-3-R1-4B-v1-Q3_K_XL.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q3_K_XL.gguf) | Q3_K_XL | 2.40GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Gemma-3-R1-4B-v1-Q4_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q4_K_S.gguf) | Q4_K_S | 2.38GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Gemma-3-R1-4B-v1-Q4_0.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q4_0.gguf) | Q4_0 | 2.37GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Gemma-3-R1-4B-v1-IQ4_NL.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ4_NL.gguf) | IQ4_NL | 2.36GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Gemma-3-R1-4B-v1-IQ4_XS.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ4_XS.gguf) | IQ4_XS | 2.26GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Gemma-3-R1-4B-v1-Q3_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q3_K_L.gguf) | Q3_K_L | 2.24GB | false | Lower quality but usable, good for low RAM availability. |
| [Gemma-3-R1-4B-v1-Q3_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q3_K_M.gguf) | Q3_K_M | 2.10GB | false | Low quality. |
| [Gemma-3-R1-4B-v1-IQ3_M.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ3_M.gguf) | IQ3_M | 1.99GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Gemma-3-R1-4B-v1-Q3_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q3_K_S.gguf) | Q3_K_S | 1.94GB | false | Low quality, not recommended. |
| [Gemma-3-R1-4B-v1-Q2_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q2_K_L.gguf) | Q2_K_L | 1.89GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Gemma-3-R1-4B-v1-IQ3_XS.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ3_XS.gguf) | IQ3_XS | 1.86GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Gemma-3-R1-4B-v1-Q2_K.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-Q2_K.gguf) | Q2_K | 1.73GB | false | Very low quality but surprisingly usable. |
| [Gemma-3-R1-4B-v1-IQ3_XXS.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ3_XXS.gguf) | IQ3_XXS | 1.69GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Gemma-3-R1-4B-v1-IQ2_M.gguf](https://huggingface.co/bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF/blob/main/TheDrummer_Gemma-3-R1-4B-v1-IQ2_M.gguf) | IQ2_M | 1.54GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF --include "TheDrummer_Gemma-3-R1-4B-v1-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/TheDrummer_Gemma-3-R1-4B-v1-GGUF --include "TheDrummer_Gemma-3-R1-4B-v1-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (TheDrummer_Gemma-3-R1-4B-v1-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c73fecf48de68fe96dd8d5491444d5792ac09a0157625496c0ea2ad3bf5b3e7c
size 1537982528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e62a07d8e6259599827ece3039ee8f8818f7ad41a431874b69081938b74aa9b
size 1986803008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:780d58066fdc5936c04c605f5b13997b009e6c53ee89c0de0876770da91828ec
size 1863390528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f221eb7d9de9af59216b645a9d84cb421896c5dc952a44cb2314f3eadb0bb85b
size 1689452608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f633c34e019a24f610105a62002cad8bde4be38113dd0288bec38c64ea4937dc
size 2363512128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a81778b332048b807bd4f414969022eb0a1ae5ce0b46f07deb9607bb3894a431
size 2263242048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d3d4d5ffbab2e036b729f008232f13e8b4adfdb45f1edc5abfaec8ad102099d
size 1729164608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7ff091ad022f426d7da63bc6600f5c53a24ce1e074cb80d376a9f62f5a72d8e7
size 1891733568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb97a305ed7b5b2c44a98b13dbefbf017ceb3fe4b5ac5016ddf0a7b1e302fc21
size 2236085568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e23384b013fd7d7a4769f2fda63e81c0fc1f5d2c4e592c9034ec3f85b0ef4f43
size 2098459968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:407335bcc087e8ac9ebc2c02c320868b13f0c6475b2e8d1a8472017bc41bca11
size 1937364288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e464098b91afe316bbe04f8c3948a954d4e511f585a9bd50a63ad1ad8905ea8b
size 2398654528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe7bf82ecde40fe0fc1eb8d499589676b9a30585a352ad54a67e77edff4e6298
size 2370065728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:31054d37772600780e1a91b22d980ad995602ac7d133190208749bdb455622c2
size 2564052288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e291fd612d630390f4f5365fb4fccfb58886a431ad849f95df3ba05444a03d7e
size 2652463168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:72a7dc5bddbdf6bbea0d47aea8573d6baa191f4ddebd75547091c991678bcd08
size 2489894208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:36274277f077ea30bc50a3b7a5aa2f7be206c849484c37463e437d9a46b14189
size 2377930048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b8836459220db3301ec5574dc3f8a77f37264613c55a7ccf21eaef3c0c4fb2da
size 2992267328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ce2a08687bb6009fd99633e82f1eda1899ba43790390194e345a01ec8356399f
size 2829698368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b35e60995b6339b856f08654dce8d38a4705fcbc193d8c8766dbce435fba16e
size 2764592448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:03447a9a331a934592dfdea2414931110448c0850b3d88d2d2eff84e580c623a
size 3190740288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:59d3377e6589dcf7881ad8cde4eebf40bc34aafe31d5e63a259b8611227d429a
size 3353309248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:58bd8e78955f42012d7d39e3458bd10f792bd8bbcabf95bceda8c113b248483b
size 4130402368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fce7b7cb85c09b03e3bbf1e4e389bae495386a7ec90f0da1a9b6277a5ff0859e
size 7767803680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e250c8ae830c08880b7a7db72df4f13c75b58dce81b7574390aecc82ec3b537
size 3448608

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}