初始化项目,由ModelHub XC社区提供模型

Model: Mungert/SmallThinker-21BA3B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-26 02:32:13 +08:00
commit fc85411b7b
32 changed files with 395 additions and 0 deletions

77
.gitattributes vendored Normal file
View File

@@ -0,0 +1,77 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq2_xs.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-imatrix.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq3_xs.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q2_k_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq2_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q2_k_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq4_nl.gguf.tmp filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq1_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq3_xxs.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq4_xs.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq2_xxs.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-bf16_q8_0.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-f16_q8_0.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q5_0.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q5_1.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q6_k_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-bf16.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q4_0.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-q4_1.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq1_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq2_m.gguf filter=lfs diff=lfs merge=lfs -text
SmallThinker-21BA3B-Instruct-iq3_m.gguf filter=lfs diff=lfs merge=lfs -text

230
README.md Normal file
View File

@@ -0,0 +1,230 @@
---
language:
- en
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
- moe
---
# <span style="color: #7FFF7F;">SmallThinker-21BA3B-Instruct GGUF Models</span>
## <span style="color: #7F7FFF;">Model Generation Details</span>
This model was generated using [llama.cpp](https://github.com/ggerganov/llama.cpp) at commit [`4cb208c9`](https://github.com/ggerganov/llama.cpp/commit/4cb208c93c1c938591a5b40354e2a6f9b94489bc).
---
## <span style="color: #7FFF7F;">Quantization Beyond the IMatrix</span>
I've been experimenting with a new quantization approach that selectively elevates the precision of key layers beyond what the default IMatrix configuration provides.
In my testing, standard IMatrix quantization underperforms at lower bit depths, especially with Mixture of Experts (MoE) models. To address this, I'm using the `--tensor-type` option in `llama.cpp` to manually "bump" important layers to higher precision. You can see the implementation here:
👉 [Layer bumping with llama.cpp](https://github.com/Mungert69/GGUFModelBuilder/blob/main/model-converter/tensor_list_builder.py)
While this does increase model file size, it significantly improves precision for a given quantization level.
### **I'd love your feedback—have you tried this? How does it perform for you?**
---
<a href="https://readyforquantum.com/huggingface_gguf_selection_guide.html" style="color: #7FFF7F;">
Click here to get info on choosing the right GGUF model format
</a>
---
<!--Begin Original Model Card-->
## Introduction
<p align="center">
&nbsp&nbsp🤗 <a href="https://huggingface.co/PowerInfer">Hugging Face</a>&nbsp&nbsp | &nbsp&nbsp🤖 <a href="https://modelscope.cn/organization/PowerInfer">ModelScope</a>&nbsp&nbsp | &nbsp&nbsp 📑 <a href="https://github.com/SJTU-IPADS/SmallThinker/blob/main/smallthinker-technical-report.pdf">Technical Report</a> &nbsp&nbsp
&nbsp&nbsp 📚 <a href="https://huggingface.co/papers/2507.20984">Paper</a> &nbsp&nbsp | &nbsp&nbsp 💻 <a href="https://github.com/SJTU-IPADS/SmallThinker">GitHub Repo</a> &nbsp&nbsp
</p>
SmallThinker is a family of **on-device native** Mixture-of-Experts (MoE) language models specially designed for local deployment,
co-developed by the **IPADS** and **School of AI at Shanghai Jiao Tong University** and **Zenergize AI**.
Designed from the ground up for resource-constrained environments,
SmallThinker brings powerful, private, and low-latency AI directly to your personal devices,
without relying on the cloud.
## Paper
The model was presented in the paper [SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment](https://huggingface.co/papers/2507.20984).
### Abstract
While frontier large language models (LLMs) continue to push capability boundaries, their deployment remains confined to GPU-powered cloud infrastructure. We challenge this paradigm with SmallThinker, a family of LLMs natively designed - not adapted - for the unique constraints of local devices: weak computational power, limited memory, and slow storage. Unlike traditional approaches that mainly compress existing models built for clouds, we architect SmallThinker from the ground up to thrive within these limitations. Our innovation lies in a deployment-aware architecture that transforms constraints into design principles. First, We introduce a two-level sparse structure combining fine-grained Mixture-of-Experts (MoE) with sparse feed-forward networks, drastically reducing computational demands without sacrificing model capacity. Second, to conquer the I/O bottleneck of slow storage, we design a pre-attention router that enables our co-designed inference engine to prefetch expert parameters from storage while computing attention, effectively hiding storage latency that would otherwise cripple on-device inference. Third, for memory efficiency, we utilize NoPE-RoPE hybrid sparse attention mechanism to slash KV cache requirements. We release SmallThinker-4B-A0.6B and SmallThinker-21B-A3B, which achieve state-of-the-art performance scores and even outperform larger LLMs. Remarkably, our co-designed system mostly eliminates the need for expensive GPU hardware: with Q4_0 quantization, both models exceed 20 tokens/s on ordinary consumer CPUs, while consuming only 1GB and 8GB of memory respectively. SmallThinker is publicly available at this http URL and this http URL .
## Performance
Note: The model is trained mainly on English.
| Model | MMLU | GPQA-diamond | MATH-500 | IFEVAL | LIVEBENCH | HUMANEVAL | Average |
|---|---|---|---|---|---|---|---|
| **SmallThinker-21BA3B-Instruct** | 84.43 | <u>55.05</u> | 82.4 | **85.77** | **60.3** | <u>89.63</u> | **76.26** |
| Gemma3-12b-it | 78.52 | 34.85 | 82.4 | 74.68 | 44.5 | 82.93 | 66.31 |
| Qwen3-14B | <u>84.82</u> | 50 | **84.6** | <u>85.21</u>| <u>59.5</u> | 88.41 | <u>75.42</u> |
| Qwen3-30BA3B | **85.1** | 44.4 | <u>84.4</u> | 84.29 | 58.8 | **90.24** | 74.54 |
| Qwen3-8B | 81.79 | 38.89 | 81.6 | 83.92 | 49.5 | 85.9 | 70.26 |
| Phi-4-14B | 84.58 | **55.45** | 80.2 | 63.22 | 42.4 | 87.2 | 68.84 |
For the MMLU evaluation, we use a 0-shot CoT setting.
All models are evaluated in non-thinking mode.
## Speed
| Model | Memory(GiB) | i9 14900 | 1+13 8ge4 | rk3588 (16G) | Raspberry PI 5 |
|---|---|---|---|---|---|
| SmallThinker 21B+sparse | 11.47 | 30.19 | 23.03 | 10.84 | 6.61 |
| SmallThinker 21B+sparse+limited memory | limit 8G | 20.30 | 15.50 | 8.56 | - |
| Qwen3 30B A3B | 16.20 | 33.52 | 20.18 | 9.07 | - |
| Qwen3 30B A3B+limited memory | limit 8G | 10.11 | 0.18 | 6.32 | - |
| Gemma 3n E2B | 1G, theoretically | 36.88 | 27.06 | 12.50 | 6.66 |
| Gemma 3n E4B | 2G, theoretically | 21.93 | 16.58 | 7.37 | 4.01 |
Note: i9 14900, 1+13 8ge4 use 4 threads, others use the number of threads that can achieve the maximum speed. All models here have been quantized to q4_0.
You can deploy SmallThinker with offloading support using [PowerInfer](https://github.com/SJTU-IPADS/PowerInfer/tree/main/smallthinker)
## Model Card
<div align="center">
| **Architecture** | Mixture-of-Experts (MoE) |
|:---:|:---:|
| **Total Parameters** | 21B |
| **Activated Parameters** | 3B |
| **Number of Layers** | 52 |
| **Attention Hidden Dimension** | 2560 |
| **MoE Hidden Dimension** (per Expert) | 768 |
| **Number of Attention Heads** | 28 |
| **Number of KV Heads** | 4 |
| **Number of Experts** | 64 |
| **Selected Experts per Token** | 6 |
| **Vocabulary Size** | 151,936 |
| **Context Length** | 16K |
| **Attention Mechanism** | GQA |
| **Activation Function** | ReGLU |
</div>
## How to Run
### Transformers
`transformers==4.53.3` is required, we are actively working to support the latest version.
The following contains a code snippet illustrating how to use the model generate content based on given inputs.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
path = "PowerInfer/SmallThinker-21BA3B-Instruct"
device = "cuda"
tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(path, torch_dtype=torch.bfloat16, device_map=device, trust_remote_code=True)
messages = [
{"role": "user", "content": "Give me a short introduction to large language model."},
]
model_inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(device)
model_outputs = model.generate(
model_inputs,
do_sample=True,
max_new_tokens=1024
)
output_token_ids = [
model_outputs[i][len(model_inputs[i]):] for i in range(len(model_inputs))
]
responses = tokenizer.batch_decode(output_token_ids, skip_special_tokens=True)[0]
print(responses)
```
### ModelScope
`ModelScope` adopts Python API similar to (though not entirely identical to) `Transformers`. For basic usage, simply modify the first line of the above code as follows:
```python
from modelscope import AutoModelForCausalLM, AutoTokenizer
```
## Statement
- Due to the constraints of its model size and the limitations of its training data, its responses may contain factual inaccuracies, biases, or outdated information.
- Users bear full responsibility for independently evaluating and verifying the accuracy and appropriateness of all generated content.
- SmallThinker does not possess genuine comprehension or consciousness and cannot express personal opinions or value judgments.
<!--End Original Model Card-->
---
# <span id="testllm" style="color: #7F7FFF;">🚀 If you find these models useful</span>
Help me test my **AI-Powered Quantum Network Monitor Assistant** with **quantum-ready security checks**:
👉 [Quantum Network Monitor](https://readyforquantum.com/?assistant=open&utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme)
The full Open Source Code for the Quantum Network Monitor Service available at my github repos ( repos with NetworkMonitor in the name) : [Source Code Quantum Network Monitor](https://github.com/Mungert69). You will also find the code I use to quantize the models if you want to do it yourself [GGUFModelBuilder](https://github.com/Mungert69/GGUFModelBuilder)
💬 **How to test**:
Choose an **AI assistant type**:
- `TurboLLM` (GPT-4.1-mini)
- `HugLLM` (Hugginface Open-source models)
- `TestLLM` (Experimental CPU-only)
### **What Im Testing**
Im pushing the limits of **small open-source models for AI network monitoring**, specifically:
- **Function calling** against live network services
- **How small can a model go** while still handling:
- Automated **Nmap security scans**
- **Quantum-readiness checks**
- **Network Monitoring tasks**
🟡 **TestLLM** Current experimental model (llama.cpp on 2 CPU threads on huggingface docker space):
-**Zero-configuration setup**
- ⏳ 30s load time (slow inference but **no API costs**) . No token limited as the cost is low.
- 🔧 **Help wanted!** If youre into **edge-device AI**, lets collaborate!
### **Other Assistants**
🟢 **TurboLLM** Uses **gpt-4.1-mini** :
- **It performs very well but unfortunatly OpenAI charges per token. For this reason tokens usage is limited.
- **Create custom cmd processors to run .net code on Quantum Network Monitor Agents**
- **Real-time network diagnostics and monitoring**
- **Security Audits**
- **Penetration testing** (Nmap/Metasploit)
🔵 **HugLLM** Latest Open-source models:
- 🌐 Runs on Hugging Face Inference API. Performs pretty well using the lastest models hosted on Novita.
### 💡 **Example commands you could test**:
1. `"Give me info on my websites SSL certificate"`
2. `"Check if my server is using quantum safe encyption for communication"`
3. `"Run a comprehensive security audit on my server"`
4. '"Create a cmd processor to .. (what ever you want)" Note you need to install a [Quantum Network Monitor Agent](https://readyforquantum.com/Download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) to run the .net code on. This is a very flexible and powerful feature. Use with caution!
### Final Word
I fund the servers used to create these model files, run the Quantum Network Monitor service, and pay for inference from Novita and OpenAI—all out of my own pocket. All the code behind the model creation and the Quantum Network Monitor project is [open source](https://github.com/Mungert69). Feel free to use whatever you find helpful.
If you appreciate the work, please consider [buying me a coffee](https://www.buymeacoffee.com/mahadeva) ☕. Your support helps cover service costs and allows me to raise token limits for everyone.
I'm also open to job opportunities or sponsorship.
Thank you! 😊

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73ff93b54550646202e388325952d2469358598983d23ae3d483d79b7db9b36c
size 43036667296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6bed1ae457e80b43cc63dc8a0ad39b485740db6aa3fdfa8e19005c6cea24df50
size 24122977696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:43f6385cecb71be1cd4d6ddf0a0062c04670faf545ff7e68189d462a1ba6b5ce
size 24122977696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6c9e3b6b538385e47cb159925917c018e038d7822c3aa02512ad117a0231d1f1
size 81359936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8470af0a6e5c87821426f8820e6907d0cbcf5f650f9b11dd97b542068caa8d0e
size 6207545088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:46d96464a74e37bbdd32c33c4896a276fffcc69a1c6dae8e415443464fe9953e
size 5696098048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e81e1d90d86dfc2ba6b884af1f5c415378b9c05d9fc5ccc9349c01d2270d8547
size 7839555328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e5a16dfb16a4a472d5d991d70bd73681a50470de60dcbc85bc7e1afa2a8d15ce
size 7475830528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa6262191310193cf6af1679838ab63b8f480d4c1773a4cdd40b98763c400966
size 6954655488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0872a1676313500f21ac111e2f83e3b5f735f977e793b339610d11a5aa713973
size 6513987328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fff4d247b988502cedec2430ef21b7378f69e64477f8cdd965108b4b715f8123
size 10263138048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a92219827c433656ef9ea8937a46e1f9ab2e160d4f382f2d4839d1f8674ddc43
size 9155784448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:41e00c5e82fd8ea6707195dc622a140e5e2620746e4626838fe890a2f1c32967
size 9109909248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f5b1e2cb074c517fc42820070d202d80e9e1940af3937907d12b3b3da746544
size 555799808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7efd5dedceff4ca92edb73ffacda989b19c2ddd9efd01632408d6b6508a02832
size 11584894208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9b3eb543b3556a0c05589feecc63f4b3b082886de5f452f8d469679502e46180
size 8206920448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ea85992655b03632ad464a3f582cc6e127ec781cbddd3375e5dbc711537b77b1
size 8166600448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33f7ba93ac2136d94d9adc372abd0b8fac715914c2f09df5c8934df2d9e9ffb5
size 10877963008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c9083daaf01358bbab58b05761aa2a1e22ceabded0c16211fa9fb195ba957995
size 10715664128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:59545815ee6a5ff7fe324434a0dd2fb6b0c8f72e1d9aa6fb4335d0d8436b8074
size 12133617408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:80910158c2bec5e138bccafb4ecadbbb82b23258a92fee8f335a521b1c6607db
size 13477228288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c94e598f681a23f089f80a996f1541cbaa15cb185f19d3862ae1c7a9a9fb0467
size 13577657088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed83a9fcbac5d853485bfb4d5c8c7efcb5b47aaaa0b938772f95e5175f5085fc
size 12916132608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:828e9518a9d5ffe24bf4d81ea96bd68dcfc02ad1f48f597951cd2016655b55e6
size 14820839168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12ccadc107f9357f8b449c444226bd8b0946c3c25b49f1e04389b2eba2e783ff
size 16164450048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5efce61b0c230d27319c9da6e959d3315594763112ff9c6815a9261006ce22e5
size 15778888448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a9cb49288ca670b8a1c01c36d9712759c732fa886917e8a57b4164de0ee720f
size 15542487808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:df58436a58d2c0d5f73b89ed333b68528a2e74769af8a24a90746990215a67b4
size 17676012288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2be1a351ec0232701824fef718c1b6382206fa5d2425562c6c36b7903f661521
size 22882504096

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}