初始化项目,由ModelHub XC社区提供模型

Model: Mungert/EXAONE-Deep-2.4B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-28 04:00:17 +08:00
commit dbc1423f52
30 changed files with 554 additions and 0 deletions

72
.gitattributes vendored Normal file
View File

@@ -0,0 +1,72 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_1_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_1.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_1_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-tq1_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-tq2_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_0_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q6_k_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_k_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq3_xxs.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq3_s.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-f16.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-bf16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq3_m.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B.imatrix filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq4_nl.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_k_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq3_xs.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-iq4_xs.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q3_k_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-f16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-bf16.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q4_1.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-f16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-f16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q6_k_m.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-bf16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-q5_0_l.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-Deep-2.4B-bf16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c4748c2106c7257321c02e6caba99c3857136ea3e6a632c08fc6a14b6cf022b
size 1806711200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d42995add28a36e54736420f9ff802727079012a3f5d5b178e5d4a1728b812c
size 2287064480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bb1c7ede27d9a8b9438d64fb8d20c507ed226ea59983d61a6e6ae74b552378b1
size 2806078592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b824f6518b71e6c81733c2c3fe71b2c103f52be7488e90e1fe22c54559778ed6
size 4815166592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b1b36600234a7ea279f97c2aa323f8db0cf1cf7ea66891cdc57231e7434fb7d
size 1806711200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2109d8ef5c6ce9a8452ebcba851f9a5e7656b70b56265147d8334f0171af60af
size 2287064480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:980f29e3e6e61553dfa4493e601b0d97d5d2b6bf9b2953dcbc3e989be213ffec
size 2806078592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e20c5dae83f7f0a3258317405575a810eb0bba8421669584cb5042f475a3901e
size 1180647840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3492b103ca52cff560604267341e034e130528f51ca32531fae6b568328d17c
size 1147224480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5f28403eb57a49427439ee914f21d7909b01cafa4a31d4ed31888fb7c6aab949
size 1096137120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cc9bccbdbcaae2c084baf45e25c5ec1108fddeea2786f4f1df46e4b0b296f293
size 1008114080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:546d07542f79ee57086b57d27c65e4c64eec1bf249b8e51e59920c4084b77bcb
size 1431461280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:37535dfbcbb88930cefd5ff1ed0536e6e51e9e3fc73c7f43384edcdbdd178d99
size 1366027680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b8261cb4c05d65f5746eba3e3afddaf771131403cf2c4f9d2cc506b8082a5c80
size 1249153440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:59092e13fc85b0fd45789f5c5acec43de530c435b58183f0e45725b16af2a52e
size 1140696480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:94a5784b9a5f4d9cb13f839a0f33ab81f88fa2027132f02b37cf9c0028a89681
size 1357733280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d42355814d3b5159515c171fba3297e4e13f813dbf5ee24e219ac26c4a526e8f
size 1508056480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d821e5550097c7c48858bd454bbdad16c475405dc50a1829d04946a926294651
size 1497463200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b46eb7344d4563fcff7ee4b176959749a2d0bc4d7127f0e00263e6d2766b6b92
size 1433017760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:22925bb61077c15431bb73e045dffd7644ae5a181c0cba3439f4399efadbce61
size 1658379680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a23dade3236946549544b68d34e9dd674e21f5a68537f8ff65a76a93bd6fb9d5
size 1808702880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f82614dfb4ddbe798be2ad08f8f6f0b803db5db0478c2da6c71e64967d462ac9
size 1730361760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e3dffab6a794eece3834235c09306c754a0eedb2eb6cdc4ae2e4684478792f9
size 1693195680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64b4c0caa4f4a9286f1e646cdc7023b7a226376c23c43c6388cfc10a4b7ae798
size 1977816480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c83c95a39d2e7082e5141e24e33bf5a67e289e4df86fb41cd71fc95bb82dcd83
size 2560318592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6ad4cdbfd01dc98e14f716b63f4963d144ff225d2181ef98d91c480a9aa870b1
size 671909280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d852ad293b834df367e7d12433e985675406f07ef61d2afef5333aa3d0179b93
size 772363680

3
EXAONE-Deep-2.4B.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5a60b2d4a02cdcda37b98bc5a79d943ad1d07000d839dc2081b5f7197d69ff0b
size 2710328

398
README.md Normal file
View File

@@ -0,0 +1,398 @@
---
base_model: LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct
language:
- en
- ko
library_name: transformers
license: other
license_name: exaone
license_link: LICENSE
pipeline_tag: text-generation
tags:
- lg-ai
- exaone
- exaone-deep
base_model_relation: finetune
---
# <span style="color: #7FFF7F;">EXAONE-Deep-2.4B GGUF Models</span>
## **Choosing the Right Model Format**
Selecting the correct model format depends on your **hardware capabilities** and **memory constraints**.
### **BF16 (Brain Float 16) Use if BF16 acceleration is available**
- A 16-bit floating-point format designed for **faster computation** while retaining good precision.
- Provides **similar dynamic range** as FP32 but with **lower memory usage**.
- Recommended if your hardware supports **BF16 acceleration** (check your devices specs).
- Ideal for **high-performance inference** with **reduced memory footprint** compared to FP32.
📌 **Use BF16 if:**
✔ Your hardware has native **BF16 support** (e.g., newer GPUs, TPUs).
✔ You want **higher precision** while saving memory.
✔ You plan to **requantize** the model into another format.
📌 **Avoid BF16 if:**
❌ Your hardware does **not** support BF16 (it may fall back to FP32 and run slower).
❌ You need compatibility with older devices that lack BF16 optimization.
---
### **F16 (Float 16) More widely supported than BF16**
- A 16-bit floating-point **high precision** but with less of range of values than BF16.
- Works on most devices with **FP16 acceleration support** (including many GPUs and some CPUs).
- Slightly lower numerical precision than BF16 but generally sufficient for inference.
📌 **Use F16 if:**
✔ Your hardware supports **FP16** but **not BF16**.
✔ You need a **balance between speed, memory usage, and accuracy**.
✔ You are running on a **GPU** or another device optimized for FP16 computations.
📌 **Avoid F16 if:**
❌ Your device lacks **native FP16 support** (it may run slower than expected).
❌ You have memory limitations.
---
### **Quantized Models (Q4_K, Q6_K, Q8, etc.) For CPU & Low-VRAM Inference**
Quantization reduces model size and memory usage while maintaining as much accuracy as possible.
- **Lower-bit models (Q4_K)** → **Best for minimal memory usage**, may have lower precision.
- **Higher-bit models (Q6_K, Q8_0)** → **Better accuracy**, requires more memory.
📌 **Use Quantized Models if:**
✔ You are running inference on a **CPU** and need an optimized model.
✔ Your device has **low VRAM** and cannot load full-precision models.
✔ You want to reduce **memory footprint** while keeping reasonable accuracy.
📌 **Avoid Quantized Models if:**
❌ You need **maximum accuracy** (full-precision models are better for this).
❌ Your hardware has enough VRAM for higher-precision formats (BF16/F16).
---
### **Very Low-Bit Quantization (IQ3_XS, IQ3_S, IQ3_M, Q4_K, Q4_0)**
These models are optimized for **extreme memory efficiency**, making them ideal for **low-power devices** or **large-scale deployments** where memory is a critical constraint.
- **IQ3_XS**: Ultra-low-bit quantization (3-bit) with **extreme memory efficiency**.
- **Use case**: Best for **ultra-low-memory devices** where even Q4_K is too large.
- **Trade-off**: Lower accuracy compared to higher-bit quantizations.
- **IQ3_S**: Small block size for **maximum memory efficiency**.
- **Use case**: Best for **low-memory devices** where **IQ3_XS** is too aggressive.
- **IQ3_M**: Medium block size for better accuracy than **IQ3_S**.
- **Use case**: Suitable for **low-memory devices** where **IQ3_S** is too limiting.
- **Q4_K**: 4-bit quantization with **block-wise optimization** for better accuracy.
- **Use case**: Best for **low-memory devices** where **Q6_K** is too large.
- **Q4_0**: Pure 4-bit quantization, optimized for **ARM devices**.
- **Use case**: Best for **ARM-based devices** or **low-memory environments**.
---
### **Summary Table: Model Format Selection**
| Model Format | Precision | Memory Usage | Device Requirements | Best Use Case |
|--------------|------------|---------------|----------------------|---------------|
| **BF16** | Highest | High | BF16-supported GPU/CPUs | High-speed inference with reduced memory |
| **F16** | High | High | FP16-supported devices | GPU inference when BF16 isnt available |
| **Q4_K** | Medium Low | Low | CPU or Low-VRAM devices | Best for memory-constrained environments |
| **Q6_K** | Medium | Moderate | CPU with more memory | Better accuracy while still being quantized |
| **Q8_0** | High | Moderate | CPU or GPU with enough VRAM | Best accuracy among quantized models |
| **IQ3_XS** | Very Low | Very Low | Ultra-low-memory devices | Extreme memory efficiency and low accuracy |
| **Q4_0** | Low | Low | ARM or low-memory devices | llama.cpp can optimize for ARM devices |
---
## **Included Files & Details**
### `EXAONE-Deep-2.4B-bf16.gguf`
- Model weights preserved in **BF16**.
- Use this if you want to **requantize** the model into a different format.
- Best if your device supports **BF16 acceleration**.
### `EXAONE-Deep-2.4B-f16.gguf`
- Model weights stored in **F16**.
- Use if your device supports **FP16**, especially if BF16 is not available.
### `EXAONE-Deep-2.4B-bf16-q8_0.gguf`
- **Output & embeddings** remain in **BF16**.
- All other layers quantized to **Q8_0**.
- Use if your device supports **BF16** and you want a quantized version.
### `EXAONE-Deep-2.4B-f16-q8_0.gguf`
- **Output & embeddings** remain in **F16**.
- All other layers quantized to **Q8_0**.
### `EXAONE-Deep-2.4B-q4_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q4_K**.
- Good for **CPU inference** with limited memory.
### `EXAONE-Deep-2.4B-q4_k_s.gguf`
- Smallest **Q4_K** variant, using less memory at the cost of accuracy.
- Best for **very low-memory setups**.
### `EXAONE-Deep-2.4B-q6_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q6_K** .
### `EXAONE-Deep-2.4B-q8_0.gguf`
- Fully **Q8** quantized model for better accuracy.
- Requires **more memory** but offers higher precision.
### `EXAONE-Deep-2.4B-iq3_xs.gguf`
- **IQ3_XS** quantization, optimized for **extreme memory efficiency**.
- Best for **ultra-low-memory devices**.
### `EXAONE-Deep-2.4B-iq3_m.gguf`
- **IQ3_M** quantization, offering a **medium block size** for better accuracy.
- Suitable for **low-memory devices**.
### `EXAONE-Deep-2.4B-q4_0.gguf`
- Pure **Q4_0** quantization, optimized for **ARM devices**.
- Best for **low-memory environments**.
- Prefer IQ4_NL for better accuracy.
# <span id="testllm" style="color: #7F7FFF;">🚀 If you find these models useful</span>
Please click like ❤ . Also Id really appreciate it if you could test my Network Monitor Assistant at 👉 [Network Monitor Assitant](https://readyforquantum.com).
💬 Click the **chat icon** (bottom right of the main and dashboard pages) . Choose a LLM; toggle between the LLM Types TurboLLM -> FreeLLM -> TestLLM.
### What I'm Testing
I'm experimenting with **function calling** against my network monitoring service. Using small open source models. I am into the question "How small can it go and still function".
🟡 **TestLLM** Runs the current testing model using llama.cpp on 6 threads of a Cpu VM (Should take about 15s to load. Inference speed is quite slow and it only processes one user prompt at a time—still working on scaling!). If you're curious, I'd be happy to share how it works! .
### The other Available AI Assistants
🟢 **TurboLLM** Uses **gpt-4o-mini** Fast! . Note: tokens are limited since OpenAI models are pricey, but you can [Login](https://readyforquantum.com) or [Download](https://readyforquantum.com/download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) the Quantum Network Monitor agent to get more tokens, Alternatively use the TestLLM .
🔵 **HugLLM** Runs **open-source Hugging Face models** Fast, Runs small models (≈8B) hence lower quality, Get 2x more tokens (subject to Hugging Face API availability)
### Final Word
I fund the servers used to create these model files, run the Quantum Network Monitor service, and pay for inference from Novita and OpenAI—all out of my own pocket. All the code behind the model creation and the Quantum Network Monitor project is [open source](https://github.com/Mungert69). Feel free to use whatever you find helpful.
If you appreciate the work, please consider [buying me a coffee](https://www.buymeacoffee.com/mahadeva) ☕. Your support helps cover service costs and allows me to raise token limits for everyone.
I'm also open to job opportunities or sponsorship.
Thank you! 😊
# EXAONE-Deep-2.4B
## Introduction
We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. The model is described in the paper [EXAONE Deep: Reasoning Enhanced Language Models](https://huggingface.co/papers/2503.12524) and the code is available [here](https://github.com/LG-AI-EXAONE/EXAONE-Deep). Evaluation results show that 1) EXAONE Deep **2.4B** outperforms other models of comparable size, 2) EXAONE Deep **7.8B** outperforms not only open-weight models of comparable scale but also a proprietary reasoning model OpenAI o1-mini, and 3) EXAONE Deep **32B** demonstrates competitive performance against leading open-weight models.
For more details, please refer to our [documentation](https://arxiv.org/abs/2503.12524), [blog](https://www.lgresearch.ai/news/view?seq=543) and [GitHub](https://github.com/LG-AI-EXAONE/EXAONE-Deep).
This repository contains the reasoning 2.4B language model with the following features:
- Number of Parameters (without embeddings): 2.14B
- Number of Layers: 30
- Number of Attention Heads: GQA with 32 Q-heads and 8 KV-heads
- Vocab Size: 102,400
- Context Length: 32,768 tokens
- Tie Word Embeddings: True (unlike 7.8B and 32B models)
> ### Note
> The EXAONE Deep models are trained with an optimized configuration,
> so we recommend following the [Usage Guideline](#usage-guideline) section to achieve optimal performance.
## Evaluation
The following table shows the evaluation results of reasoning tasks such as math and coding. The full evaluation results can be found in the [documentation](https://arxiv.org/abs/2503.12524).
<table>
<tr>
<th>Models</th>
<th>MATH-500 (pass@1)</th>
<th>AIME 2024 (pass@1 / cons@64)</th>
<th>AIME 2025 (pass@1 / cons@64)</th>
<th>CSAT Math 2025 (pass@1)</th>
<th>GPQA Diamond (pass@1)</th>
<th>Live Code Bench (pass@1)</th>
</tr>
<tr>
<td>EXAONE Deep 32B</td>
<td>95.7</td>
<td>72.1 / <strong>90.0</strong></td>
<td>65.8 / <strong>80.0</strong></td>
<td><strong>94.5</strong></td>
<td>66.1</td>
<td>59.5</td>
</tr>
<tr>
<td>DeepSeek-R1-Distill-Qwen-32B</td>
<td>94.3</td>
<td>72.6 / 83.3</td>
<td>55.2 / 73.3</td>
<td>84.1</td>
<td>62.1</td>
<td>57.2</td>
</tr>
<tr>
<td>QwQ-32B</td>
<td>95.5</td>
<td>79.5 / 86.7</td>
<td><strong>67.1</strong> / 76.7</td>
<td>94.4</td>
<td>63.3</td>
<td>63.4</td>
</tr>
<tr>
<td>DeepSeek-R1-Distill-Llama-70B</td>
<td>94.5</td>
<td>70.0 / 86.7</td>
<td>53.9 / 66.7</td>
<td>88.8</td>
<td>65.2</td>
<td>57.5</td>
</tr>
<tr>
<td>DeepSeek-R1 (671B)</td>
<td><strong>97.3</strong></td>
<td><strong>79.8</strong> / 86.7</td>
<td>66.8 / <strong>80.0</strong></td>
<td>89.9</td>
<td><strong>71.5</strong></td>
<td><strong>65.9</strong></td>
</tr>
<tr>
<th colspan="7" height="30px"></th>
</tr>
<tr>
<td>EXAONE Deep 7.8B</td>
<td><strong>94.8</strong></td>
<td><strong>70.0</strong> / <strong>83.3</strong></td>
<td><strong>59.6</strong> / <strong>76.7</strong></td>
<td><strong>89.9</strong></td>
<td><strong>62.6</strong></td>
<td><strong>55.2</strong></td>
</tr>
<tr>
<td>DeepSeek-R1-Distill-Qwen-7B</td>
<td>92.8</td>
<td>55.5 / 83.3</td>
<td>38.5 / 56.7</td>
<td>79.7</td>
<td>49.1</td>
<td>37.6</td>
</tr>
<tr>
<td>DeepSeek-R1-Distill-Llama-8B</td>
<td>89.1</td>
<td>50.4 / 80.0</td>
<td>33.6 / 53.3</td>
<td>74.1</td>
<td>49.0</td>
<td>39.6</td>
</tr>
<tr>
<td>OpenAI o1-mini</td>
<td>90.0</td>
<td>63.6 / 80.0</td>
<td>54.8 / 66.7</td>
<td>84.4</td>
<td>60.0</td>
<td>53.8</td>
</tr>
<tr>
<th colspan="7" height="30px"></th>
</tr>
<tr>
<td>EXAONE Deep 2.4B</td>
<td><strong>92.3</strong></td>
<td><strong>52.5</strong> / <strong>76.7</strong></td>
<td><strong>47.9</strong> / <strong>73.3</strong></td>
<td><strong>79.2</strong></td>
<td><strong>54.3</strong></td>
<td><strong>46.6</strong></td>
</tr>
<tr>
<td>DeepSeek-R1-Distill-Qwen-1.5B</td>
<td>83.9</td>
<td>28.9 / 52.7</td>
<td>23.9 / 36.7</td>
<td>65.6</td>
<td>33.8</td>
<td>16.9</td>
</tr>
</table>
## Deployment
EXAONE Deep models can be inferred in the various frameworks, such as:
- `TensorRT-LLM`
- `vLLM`
- `SGLang`
- `llama.cpp`
- `Ollama`
- `LM-Studio`
Please refer to our [EXAONE Deep GitHub](https://github.com/LG-AI-EXAONE/EXAONE-Deep) for more details about the inference frameworks.
## Quantization
We provide the pre-quantized EXAONE Deep models with **AWQ** and several quantization types in **GGUF** format. Please refer to our [EXAONE Deep collection](https://huggingface.co/collections/LGAI-EXAONE/exaone-deep-67d119918816ec6efa79a4aa) to find corresponding quantized models.
## Usage Guideline
To achieve the expected performance, we recommend using the following configurations:
1. Ensure the model starts with `<thought>
` for reasoning steps. The model's output quality may be degraded when you omit it. You can easily apply this feature by using `tokenizer.apply_chat_template()` with `add_generation_prompt=True`. Please check the example code on [Quickstart](#quickstart) section.
2. The reasoning steps of EXAONE Deep models enclosed by `<thought>
...
</thought>` usually have lots of tokens, so previous reasoning steps may be necessary to be removed in multi-turn situation. The provided tokenizer handles this automatically.
3. Avoid using system prompt, and build the instruction on the user prompt.
4. Additional instructions help the models reason more deeply, so that the models generate better output.
- For math problems, the instructions **"Please reason step by step, and put your final answer within \boxed{}."** are helpful.
- For more information on our evaluation setting including prompts, please refer to our [Documentation](https://arxiv.org/abs/2503.12524).
5. In our evaluation, we use `temperature=0.6` and `top_p=0.95` for generation.
6. When evaluating the models, it is recommended to test multiple times to assess the expected performance accurately.
## Limitation
The EXAONE language model has certain limitations and may occasionally generate inappropriate responses. The language model generates responses based on the output probability of tokens, and it is determined during learning from training data. While we have made every effort to exclude personal, harmful, and biased information from the training data, some problematic content may still be included, potentially leading to undesirable responses. Please note that the text generated by EXAONE language model does not reflects the views of LG AI Research.
- Inappropriate answers may be generated, which contain personal, harmful or other inappropriate information.
- Biased responses may be generated, which are associated with age, gender, race, and so on.
- The generated responses rely heavily on statistics from the training data, which can result in the generation of
semantically or syntactically incorrect sentences.
- Since the model does not reflect the latest information, the responses may be false or contradictory.
LG AI Research strives to reduce potential risks that may arise from EXAONE language models. Users are not allowed
to engage in any malicious activities (e.g., keying in illegal information) that may induce the creation of inappropriate
outputs violating LG AIs ethical principles when using EXAONE language models.
## License
The model is licensed under [EXAONE AI Model License Agreement 1.1 - NC](./LICENSE)
## Citation
```
@article{exaone-deep,
title={EXAONE Deep: Reasoning Enhanced Language Models},
author={{LG AI Research}},
journal={arXiv preprint arXiv:2503.12524},
year={2025}
}
```
## Contact
LG AI Research Technical Support: contact_us@lgresearch.ai