初始化项目,由ModelHub XC社区提供模型

Model: bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 23:42:14 +08:00
commit 29265a9180
28 changed files with 270 additions and 0 deletions

24
.gitattributes vendored Normal file
View File

@@ -0,0 +1,24 @@
TheDrummer_Fallen-Gemma3-4B-v1-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
TheDrummer_Fallen-Gemma3-4B-v1-bf16.gguf filter=lfs diff=lfs merge=lfs -text

173
README.md Normal file
View File

@@ -0,0 +1,173 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license: other
base_model_relation: quantized
base_model: TheDrummer/Fallen-Gemma3-4B-v1
---
## Llamacpp imatrix Quantizations of Fallen-Gemma3-4B-v1 by TheDrummer
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4925">b4925</a> for quantization.
Original model: https://huggingface.co/TheDrummer/Fallen-Gemma3-4B-v1
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<bos><start_of_turn>user
{system_prompt}
{prompt}<end_of_turn>
<start_of_turn>model
<end_of_turn>
<start_of_turn>model
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Fallen-Gemma3-4B-v1-bf16.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-bf16.gguf) | bf16 | 7.77GB | false | Full BF16 weights. |
| [Fallen-Gemma3-4B-v1-Q8_0.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q8_0.gguf) | Q8_0 | 4.13GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Fallen-Gemma3-4B-v1-Q6_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q6_K_L.gguf) | Q6_K_L | 3.35GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q6_K.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q6_K.gguf) | Q6_K | 3.19GB | false | Very high quality, near perfect, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q5_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_L.gguf) | Q5_K_L | 2.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q5_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_M.gguf) | Q5_K_M | 2.83GB | false | High quality, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q5_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q5_K_S.gguf) | Q5_K_S | 2.76GB | false | High quality, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q4_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_L.gguf) | Q4_K_L | 2.65GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q4_1.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q4_1.gguf) | Q4_1 | 2.56GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Fallen-Gemma3-4B-v1-Q4_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_M.gguf) | Q4_K_M | 2.49GB | false | Good quality, default size for most use cases, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q3_K_XL.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_XL.gguf) | Q3_K_XL | 2.40GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Fallen-Gemma3-4B-v1-Q4_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_S.gguf) | Q4_K_S | 2.38GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q4_0.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q4_0.gguf) | Q4_0 | 2.37GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Fallen-Gemma3-4B-v1-IQ4_NL.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ4_NL.gguf) | IQ4_NL | 2.36GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Fallen-Gemma3-4B-v1-IQ4_XS.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ4_XS.gguf) | IQ4_XS | 2.26GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Fallen-Gemma3-4B-v1-Q3_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_L.gguf) | Q3_K_L | 2.24GB | false | Lower quality but usable, good for low RAM availability. |
| [Fallen-Gemma3-4B-v1-Q3_K_M.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_M.gguf) | Q3_K_M | 2.10GB | false | Low quality. |
| [Fallen-Gemma3-4B-v1-IQ3_M.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ3_M.gguf) | IQ3_M | 1.99GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Fallen-Gemma3-4B-v1-Q3_K_S.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q3_K_S.gguf) | Q3_K_S | 1.94GB | false | Low quality, not recommended. |
| [Fallen-Gemma3-4B-v1-Q2_K_L.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q2_K_L.gguf) | Q2_K_L | 1.89GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Fallen-Gemma3-4B-v1-IQ3_XS.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ3_XS.gguf) | IQ3_XS | 1.86GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Fallen-Gemma3-4B-v1-Q2_K.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-Q2_K.gguf) | Q2_K | 1.73GB | false | Very low quality but surprisingly usable. |
| [Fallen-Gemma3-4B-v1-IQ3_XXS.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ3_XXS.gguf) | IQ3_XXS | 1.69GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Fallen-Gemma3-4B-v1-IQ2_M.gguf](https://huggingface.co/bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF/blob/main/TheDrummer_Fallen-Gemma3-4B-v1-IQ2_M.gguf) | IQ2_M | 1.54GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF --include "TheDrummer_Fallen-Gemma3-4B-v1-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/TheDrummer_Fallen-Gemma3-4B-v1-GGUF --include "TheDrummer_Fallen-Gemma3-4B-v1-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (TheDrummer_Fallen-Gemma3-4B-v1-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:035237e052982b8b8d013800e78c0180112769aa9b0db940be22178fc53e591e
size 1537982304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:da314823b150da2076d5dfe64bc476ce0e8f22e8f7a6658be933eacc66f99ddb
size 1986802784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:892ba0796ca782dd01e9e47b7bfdf41629c06c0230c87ef2260bb383a91a2b75
size 1863390304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0590724b27aea584651e5bea3a11895f6ba98762994417454e6abbb0ddc4e733
size 1689452384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65f6d6edb08d7a2c095138c9029dc76796d765bf451de759cdc624781e437da8
size 2363511904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:738d9854bd433e4a7911b2bbfb2467d53eacd49968248ec5308ccdbbe53120f7
size 2263241824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f961da0e3585def78499ac03cff10285026dc4bd0e5b7be2b1c2546377b387fe
size 1729164384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:abb739ad00568318077662a85a353f9810d3feeaccfa2a6c1be858fced25401f
size 1891733344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9535af4e019e8509a097999b7259c3f7bc39cf907b61f163959a967cece48058
size 2236085344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dbd1f9e99a3476e2d7bb0a567323676e51d9a34d8ccccd8403cbeff7d6e4e608
size 2098459744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd95564d8c5c2123e3a813565ee51da5432ca6adb55a4199d9d6d849b4d519c0
size 1937364064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:855fd39554a12e61193cd71d09da56f995a711e34ad4abb2f623b43047aa7460
size 2398654304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d663406a9beb7093e0a6128c13d504a21b7f50d218a26b37b66941b5ba583c37
size 2370065504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7abd61f41470689c59662d53168fb6e9a7c3dbfe3f835d8c545c96d5ba4e8ca6
size 2564052064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6f60604d85403dac88388267e55207dd36d053b2fcdb84e43affd4b91ed7507
size 2652462944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:85490a97bda2d40437c8dade4a68bb58e760c1263a2fbc59191daef57ee2d6c3
size 2489893984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e177fb72f0e34d0dd302c59286a38020370f2a046eee01fefb6bbcf6c3a7aec
size 2377929824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:672a665eec10f618d6c43080eb9de2c16ec4fc9c722a5fa908e3bd60f418c6ff
size 2992267104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:28f4237c2a7e4d04820b2b25b6b9233c76ce18ba828057bbc0ff23b40e4f3f52
size 2829698144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c551821cf987b4a2506c54933dc21ed6098754f190aadebf452cc4f1bccf5482
size 2764592224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a7a08c7856d233ad1a330e4c52044f3bbd0a93dc955092f7652f494442dabdeb
size 3190740064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e8f1f8449a8ef2fdd2c9ff1c884093b83de456098d5cc4a3b169e66e3c04dcb5
size 3353309024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e49b169d531b9cdba19a1d56535ad4fda51d2d6047dff6816ef35bad0a34ee88
size 4130402144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dadab320b5e62189ea79509ee5a31d33cef5b86c1da6bbab4db4a159086a5c81
size 7767803456

Binary file not shown.

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}