初始化项目,由ModelHub XC社区提供模型

Model: Mungert/phi-4-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 19:10:09 +08:00
commit 6bb8255991
37 changed files with 530 additions and 0 deletions

80
.gitattributes vendored Normal file
View File

@@ -0,0 +1,80 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
phi-4-f16.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-f16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-bf16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-f16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-bf16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-f16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-bf16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q2_k_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q3_k_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_k_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_k_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q6_k_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q6_k_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq4_xs.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq3_xs.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq4_nl.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_1.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_0_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q4_1_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_1.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_0_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q5_1_l.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq2_xs.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq2_xxs.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq2_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq2_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq1_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq1_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-tq1_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-tq2_0.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-q2_k_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq3_xxs.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq3_s.gguf filter=lfs diff=lfs merge=lfs -text
phi-4-iq3_m.gguf filter=lfs diff=lfs merge=lfs -text
phi-4.imatrix filter=lfs diff=lfs merge=lfs -text
phi-4-bf16.gguf filter=lfs diff=lfs merge=lfs -text

345
README.md Normal file
View File

@@ -0,0 +1,345 @@
---
license: mit
license_link: https://huggingface.co/microsoft/phi-4/resolve/main/LICENSE
language:
- en
pipeline_tag: text-generation
tags:
- phi
- nlp
- math
- code
- chat
- conversational
inference:
parameters:
temperature: 0
widget:
- messages:
- role: user
content: How should I explain the Internet?
library_name: transformers
---
# <span style="color: #7FFF7F;">phi-4 GGUF Models</span>
## **Choosing the Right Model Format**
Selecting the correct model format depends on your **hardware capabilities** and **memory constraints**.
### **BF16 (Brain Float 16) Use if BF16 acceleration is available**
- A 16-bit floating-point format designed for **faster computation** while retaining good precision.
- Provides **similar dynamic range** as FP32 but with **lower memory usage**.
- Recommended if your hardware supports **BF16 acceleration** (check your devices specs).
- Ideal for **high-performance inference** with **reduced memory footprint** compared to FP32.
📌 **Use BF16 if:**
✔ Your hardware has native **BF16 support** (e.g., newer GPUs, TPUs).
✔ You want **higher precision** while saving memory.
✔ You plan to **requantize** the model into another format.
📌 **Avoid BF16 if:**
❌ Your hardware does **not** support BF16 (it may fall back to FP32 and run slower).
❌ You need compatibility with older devices that lack BF16 optimization.
---
### **F16 (Float 16) More widely supported than BF16**
- A 16-bit floating-point **high precision** but with less of range of values than BF16.
- Works on most devices with **FP16 acceleration support** (including many GPUs and some CPUs).
- Slightly lower numerical precision than BF16 but generally sufficient for inference.
📌 **Use F16 if:**
✔ Your hardware supports **FP16** but **not BF16**.
✔ You need a **balance between speed, memory usage, and accuracy**.
✔ You are running on a **GPU** or another device optimized for FP16 computations.
📌 **Avoid F16 if:**
❌ Your device lacks **native FP16 support** (it may run slower than expected).
❌ You have memory limitations.
---
### **Quantized Models (Q4_K, Q6_K, Q8, etc.) For CPU & Low-VRAM Inference**
Quantization reduces model size and memory usage while maintaining as much accuracy as possible.
- **Lower-bit models (Q4_K)** → **Best for minimal memory usage**, may have lower precision.
- **Higher-bit models (Q6_K, Q8_0)** → **Better accuracy**, requires more memory.
📌 **Use Quantized Models if:**
✔ You are running inference on a **CPU** and need an optimized model.
✔ Your device has **low VRAM** and cannot load full-precision models.
✔ You want to reduce **memory footprint** while keeping reasonable accuracy.
📌 **Avoid Quantized Models if:**
❌ You need **maximum accuracy** (full-precision models are better for this).
❌ Your hardware has enough VRAM for higher-precision formats (BF16/F16).
---
### **Very Low-Bit Quantization (IQ3_XS, IQ3_S, IQ3_M, Q4_K, Q4_0)**
These models are optimized for **extreme memory efficiency**, making them ideal for **low-power devices** or **large-scale deployments** where memory is a critical constraint.
- **IQ3_XS**: Ultra-low-bit quantization (3-bit) with **extreme memory efficiency**.
- **Use case**: Best for **ultra-low-memory devices** where even Q4_K is too large.
- **Trade-off**: Lower accuracy compared to higher-bit quantizations.
- **IQ3_S**: Small block size for **maximum memory efficiency**.
- **Use case**: Best for **low-memory devices** where **IQ3_XS** is too aggressive.
- **IQ3_M**: Medium block size for better accuracy than **IQ3_S**.
- **Use case**: Suitable for **low-memory devices** where **IQ3_S** is too limiting.
- **Q4_K**: 4-bit quantization with **block-wise optimization** for better accuracy.
- **Use case**: Best for **low-memory devices** where **Q6_K** is too large.
- **Q4_0**: Pure 4-bit quantization, optimized for **ARM devices**.
- **Use case**: Best for **ARM-based devices** or **low-memory environments**.
---
### **Summary Table: Model Format Selection**
| Model Format | Precision | Memory Usage | Device Requirements | Best Use Case |
|--------------|------------|---------------|----------------------|---------------|
| **BF16** | Highest | High | BF16-supported GPU/CPUs | High-speed inference with reduced memory |
| **F16** | High | High | FP16-supported devices | GPU inference when BF16 isnt available |
| **Q4_K** | Medium Low | Low | CPU or Low-VRAM devices | Best for memory-constrained environments |
| **Q6_K** | Medium | Moderate | CPU with more memory | Better accuracy while still being quantized |
| **Q8_0** | High | Moderate | CPU or GPU with enough VRAM | Best accuracy among quantized models |
| **IQ3_XS** | Very Low | Very Low | Ultra-low-memory devices | Extreme memory efficiency and low accuracy |
| **Q4_0** | Low | Low | ARM or low-memory devices | llama.cpp can optimize for ARM devices |
---
## **Included Files & Details**
### `phi-4-bf16.gguf`
- Model weights preserved in **BF16**.
- Use this if you want to **requantize** the model into a different format.
- Best if your device supports **BF16 acceleration**.
### `phi-4-f16.gguf`
- Model weights stored in **F16**.
- Use if your device supports **FP16**, especially if BF16 is not available.
### `phi-4-bf16-q8_0.gguf`
- **Output & embeddings** remain in **BF16**.
- All other layers quantized to **Q8_0**.
- Use if your device supports **BF16** and you want a quantized version.
### `phi-4-f16-q8_0.gguf`
- **Output & embeddings** remain in **F16**.
- All other layers quantized to **Q8_0**.
### `phi-4-q4_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q4_K**.
- Good for **CPU inference** with limited memory.
### `phi-4-q4_k_s.gguf`
- Smallest **Q4_K** variant, using less memory at the cost of accuracy.
- Best for **very low-memory setups**.
### `phi-4-q6_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q6_K** .
### `phi-4-q8_0.gguf`
- Fully **Q8** quantized model for better accuracy.
- Requires **more memory** but offers higher precision.
### `phi-4-iq3_xs.gguf`
- **IQ3_XS** quantization, optimized for **extreme memory efficiency**.
- Best for **ultra-low-memory devices**.
### `phi-4-iq3_m.gguf`
- **IQ3_M** quantization, offering a **medium block size** for better accuracy.
- Suitable for **low-memory devices**.
### `phi-4-q4_0.gguf`
- Pure **Q4_0** quantization, optimized for **ARM devices**.
- Best for **low-memory environments**.
- Prefer IQ4_NL for better accuracy.
# <span id="testllm" style="color: #7F7FFF;">🚀 If you find these models useful</span>
Please click like ❤ . Also Id really appreciate it if you could test my Network Monitor Assistant at 👉 [Network Monitor Assitant](https://readyforquantum.com).
💬 Click the **chat icon** (bottom right of the main and dashboard pages) . Choose a LLM; toggle between the LLM Types TurboLLM -> FreeLLM -> TestLLM.
### What I'm Testing
I'm experimenting with **function calling** against my network monitoring service. Using small open source models. I am into the question "How small can it go and still function".
🟡 **TestLLM** Runs the current testing model using llama.cpp on 6 threads of a Cpu VM (Should take about 15s to load. Inference speed is quite slow and it only processes one user prompt at a time—still working on scaling!). If you're curious, I'd be happy to share how it works! .
### The other Available AI Assistants
🟢 **TurboLLM** Uses **gpt-4o-mini** Fast! . Note: tokens are limited since OpenAI models are pricey, but you can [Login](https://readyforquantum.com) or [Download](https://readyforquantum.com/download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) the Quantum Network Monitor agent to get more tokens, Alternatively use the TestLLM .
🔵 **HugLLM** Runs **open-source Hugging Face models** Fast, Runs small models (≈8B) hence lower quality, Get 2x more tokens (subject to Hugging Face API availability)
### Final Word
I fund the servers used to create these model files, run the Quantum Network Monitor service, and pay for inference from Novita and OpenAI—all out of my own pocket. All the code behind the model creation and the Quantum Network Monitor project is [open source](https://github.com/Mungert69). Feel free to use whatever you find helpful.
If you appreciate the work, please consider [buying me a coffee](https://www.buymeacoffee.com/mahadeva) ☕. Your support helps cover service costs and allows me to raise token limits for everyone.
I'm also open to job opportunities or sponsorship.
Thank you! 😊
# Phi-4 Model Card
[Phi-4 Technical Report](https://arxiv.org/pdf/2412.08905)
## Model Summary
| | |
|-------------------------|-------------------------------------------------------------------------------|
| **Developers** | Microsoft Research |
| **Description** | `phi-4` is a state-of-the-art open model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.<br><br>`phi-4` underwent a rigorous enhancement and alignment process, incorporating both supervised fine-tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures |
| **Architecture** | 14B parameters, dense decoder-only Transformer model |
| **Inputs** | Text, best suited for prompts in the chat format |
| **Context length** | 16K tokens |
| **GPUs** | 1920 H100-80G |
| **Training time** | 21 days |
| **Training data** | 9.8T tokens |
| **Outputs** | Generated text in response to input |
| **Dates** | October 2024 November 2024 |
| **Status** | Static model trained on an offline dataset with cutoff dates of June 2024 and earlier for publicly available data |
| **Release date** | December 12, 2024 |
| **License** | MIT |
## Intended Use
| | |
|-------------------------------|-------------------------------------------------------------------------|
| **Primary Use Cases** | Our model is designed to accelerate research on language models, for use as a building block for generative AI powered features. It provides uses for general purpose AI systems and applications (primarily in English) which require:<br><br>1. Memory/compute constrained environments.<br>2. Latency bound scenarios.<br>3. Reasoning and logic. |
| **Out-of-Scope Use Cases** | Our models is not specifically designed or evaluated for all downstream purposes, thus:<br><br>1. Developers should consider common limitations of language models as they select use cases, and evaluate and mitigate for accuracy, safety, and fairness before using within a specific downstream use case, particularly for high-risk scenarios.<br>2. Developers should be aware of and adhere to applicable laws or regulations (including privacy, trade compliance laws, etc.) that are relevant to their use case, including the models focus on English.<br>3. Nothing contained in this Model Card should be interpreted as or deemed a restriction or modification to the license the model is released under. |
## Data Overview
### Training Datasets
Our training data is an extension of the data used for Phi-3 and includes a wide variety of sources from:
1. Publicly available documents filtered rigorously for quality, selected high-quality educational data, and code.
2. Newly created synthetic, “textbook-like” data for the purpose of teaching math, coding, common sense reasoning, general knowledge of the world (science, daily activities, theory of mind, etc.).
3. Acquired academic books and Q&A datasets.
4. High quality chat format supervised data covering various topics to reflect human preferences on different aspects such as instruct-following, truthfulness, honesty and helpfulness.
Multilingual data constitutes about 8% of our overall data. We are focusing on the quality of data that could potentially improve the reasoning ability for the model, and we filter the publicly available documents to contain the correct level of knowledge.
#### Benchmark datasets
We evaluated `phi-4` using [OpenAIs SimpleEval](https://github.com/openai/simple-evals) and our own internal benchmarks to understand the models capabilities, more specifically:
* **MMLU:** Popular aggregated dataset for multitask language understanding.
* **MATH:** Challenging competition math problems.
* **GPQA:** Complex, graduate-level science questions.
* **DROP:** Complex comprehension and reasoning.
* **MGSM:** Multi-lingual grade-school math.
* **HumanEval:** Functional code generation.
* **SimpleQA:** Factual responses.
## Safety
### Approach
`phi-4` has adopted a robust safety post-training approach. This approach leverages a variety of both open-source and in-house generated synthetic datasets. The overall technique employed to do the safety alignment is a combination of SFT (Supervised Fine-Tuning) and iterative DPO (Direct Preference Optimization), including publicly available datasets focusing on helpfulness and harmlessness as well as various questions and answers targeted to multiple safety categories.
### Safety Evaluation and Red-Teaming
Prior to release, `phi-4` followed a multi-faceted evaluation approach. Quantitative evaluation was conducted with multiple open-source safety benchmarks and in-house tools utilizing adversarial conversation simulation. For qualitative safety evaluation, we collaborated with the independent AI Red Team (AIRT) at Microsoft to assess safety risks posed by `phi-4` in both average and adversarial user scenarios. In the average user scenario, AIRT emulated typical single-turn and multi-turn interactions to identify potentially risky behaviors. The adversarial user scenario tested a wide range of techniques aimed at intentionally subverting the models safety training including jailbreaks, encoding-based attacks, multi-turn attacks, and adversarial suffix attacks.
Please refer to the technical report for more details on safety alignment.
## Model Quality
To understand the capabilities, we compare `phi-4` with a set of models over OpenAIs SimpleEval benchmark.
At the high-level overview of the model quality on representative benchmarks. For the table below, higher numbers indicate better performance:
| **Category** | **Benchmark** | **phi-4** (14B) | **phi-3** (14B) | **Qwen 2.5** (14B instruct) | **GPT-4o-mini** | **Llama-3.3** (70B instruct) | **Qwen 2.5** (72B instruct) | **GPT-4o** |
|------------------------------|---------------|-----------|-----------------|----------------------|----------------------|--------------------|-------------------|-----------------|
| Popular Aggregated Benchmark | MMLU | 84.8 | 77.9 | 79.9 | 81.8 | 86.3 | 85.3 | **88.1** |
| Science | GPQA | **56.1** | 31.2 | 42.9 | 40.9 | 49.1 | 49.0 | 50.6 |
| Math | MGSM<br>MATH | 80.6<br>**80.4** | 53.5<br>44.6 | 79.6<br>75.6 | 86.5<br>73.0 | 89.1<br>66.3* | 87.3<br>80.0 | **90.4**<br>74.6 |
| Code Generation | HumanEval | 82.6 | 67.8 | 72.1 | 86.2 | 78.9* | 80.4 | **90.6** |
| Factual Knowledge | SimpleQA | 3.0 | 7.6 | 5.4 | 9.9 | 20.9 | 10.2 | **39.4** |
| Reasoning | DROP | 75.5 | 68.3 | 85.5 | 79.3 | **90.2** | 76.7 | 80.9 |
\* These scores are lower than those reported by Meta, perhaps because simple-evals has a strict formatting requirement that Llama models have particular trouble following. We use the simple-evals framework because it is reproducible, but Meta reports 77 for MATH and 88 for HumanEval on Llama-3.3-70B.
## Usage
### Input Formats
Given the nature of the training data, `phi-4` is best suited for prompts using the chat format as follows:
```bash
<|im_start|>system<|im_sep|>
You are a medieval knight and must provide explanations to modern people.<|im_end|>
<|im_start|>user<|im_sep|>
How should I explain the Internet?<|im_end|>
<|im_start|>assistant<|im_sep|>
```
### With `transformers`
```python
import transformers
pipeline = transformers.pipeline(
"text-generation",
model="microsoft/phi-4",
model_kwargs={"torch_dtype": "auto"},
device_map="auto",
)
messages = [
{"role": "system", "content": "You are a medieval knight and must provide explanations to modern people."},
{"role": "user", "content": "How should I explain the Internet?"},
]
outputs = pipeline(messages, max_new_tokens=128)
print(outputs[0]["generated_text"][-1])
```
## Responsible AI Considerations
Like other language models, `phi-4` can potentially behave in ways that are unfair, unreliable, or offensive. Some of the limiting behaviors to be aware of include:
* **Quality of Service:** The model is trained primarily on English text. Languages other than English will experience worse performance. English language varieties with less representation in the training data might experience worse performance than standard American English. `phi-4` is not intended to support multilingual use.
* **Representation of Harms & Perpetuation of Stereotypes:** These models can over- or under-represent groups of people, erase representation of some groups, or reinforce demeaning or negative stereotypes. Despite safety post-training, these limitations may still be present due to differing levels of representation of different groups or prevalence of examples of negative stereotypes in training data that reflect real-world patterns and societal biases.
* **Inappropriate or Offensive Content:** These models may produce other types of inappropriate or offensive content, which may make it inappropriate to deploy for sensitive contexts without additional mitigations that are specific to the use case.
* **Information Reliability:** Language models can generate nonsensical content or fabricate content that might sound reasonable but is inaccurate or outdated.
* **Limited Scope for Code:** Majority of `phi-4` training data is based in Python and uses common packages such as `typing`, `math`, `random`, `collections`, `datetime`, `itertools`. If the model generates Python scripts that utilize other packages or scripts in other languages, we strongly recommend users manually verify all API uses.
Developers should apply responsible AI best practices and are responsible for ensuring that a specific use case complies with relevant laws and regulations (e.g. privacy, trade, etc.). Using safety services like [Azure AI Content Safety](https://azure.microsoft.com/en-us/products/ai-services/ai-content-safety) that have advanced guardrails is highly recommended. Important areas for consideration include:
* **Allocation:** Models may not be suitable for scenarios that could have consequential impact on legal status or the allocation of resources or life opportunities (ex: housing, employment, credit, etc.) without further assessments and additional debiasing techniques.
* **High-Risk Scenarios:** Developers should assess suitability of using models in high-risk scenarios where unfair, unreliable or offensive outputs might be extremely costly or lead to harm. This includes providing advice in sensitive or expert domains where accuracy and reliability are critical (ex: legal or health advice). Additional safeguards should be implemented at the application level according to the deployment context.
* **Misinformation:** Models may produce inaccurate information. Developers should follow transparency best practices and inform end-users they are interacting with an AI system. At the application level, developers can build feedback mechanisms and pipelines to ground responses in use-case specific, contextual information, a technique known as Retrieval Augmented Generation (RAG).
* **Generation of Harmful Content:** Developers should assess outputs for their context and use available safety classifiers or custom solutions appropriate for their use case.
* **Misuse:** Other forms of misuse such as fraud, spam, or malware production may be possible, and developers should ensure that their applications do not violate applicable laws and regulations.

3
phi-4-bf16-q4_k.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6bc60e57bfffce1c4cb396acd412e757d5e3a4481ce0c51a213021f84b8ad153
size 10397831712

3
phi-4-bf16-q6_k.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0dc82a30d8c80fc53129377791d920fe3ee49b471b738c8858ba27f3a7f960f4
size 13242503712

3
phi-4-bf16-q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:506e0c358c0dcc57da90bd252cbc9a25a91d7d6f33b7bb0d629a40a794343050
size 16543879424

3
phi-4-bf16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:992ce0e843c2ea299a0a10fd4381409be893cef48f04228676fa934eabfcb948
size 29323399424

3
phi-4-f16-q4_k.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ba322a592a97f61d09a622e3fc0cd063a42793363cd049c17d928bc860cbe17
size 10397831712

3
phi-4-f16-q6_k.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87ecfe9a05d87ea0d45459039ef746d2909bce3c8e4096536dba37517df8a1ec
size 13242503712

3
phi-4-f16-q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:26128f604ceae7cb77a8015abc48bb51a70cae2358bc0f629548a9460aa4d08f
size 16543879424

3
phi-4-iq1_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1a5bca0476835e4d99307b179d662ea21d49aae5c392fa7fa9e6a9626ecea374
size 3600069152

3
phi-4-iq1_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a75dc3fffdefe478a2bdda6d15121a0171e1546628423176264a3153d2b2ee9
size 3315909152

3
phi-4-iq2_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:27f139ffabda764084026b2e0d87ea6cbe7822e5e8a962f1fddefc6a11524cd1
size 5110428192

3
phi-4-iq2_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:297a08b56ec6e9eb6abc5923d6f9cd018b1be51b3b7f50114f7f5db6d55a56c6
size 4731548192

3
phi-4-iq2_xs.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:890d88d3d9f858f5eddcbf772a247c23bad1bcf47f621c59a7c7c18ca420af49
size 4485317152

3
phi-4-iq2_xxs.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0351b3b17703bfe7c57437d02b676957f24333b5e7f7194a537a688db31ef4d4
size 4073669152

3
phi-4-iq3_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c07ccb7f621088b8f9c87dbd0d308ba802cbb4837ad46d4aeef012b1c42f3d1f
size 6913835552

3
phi-4-iq3_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c824694fa26d9728fc578f1570779bba8a9cc65a6afba8c04024e2bb3c058cd
size 6504747552

3
phi-4-iq3_xs.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7f5eebd6d4bcd477ee827d33da450a3e8198575bb236bb3af6cdd5d28732e64
size 6246699552

3
phi-4-iq3_xxs.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f260464bcae06864652ba7d35bb44797acd180757dfe6a329e54356800278e2a
size 5846684192

3
phi-4-iq4_nl.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87ccaee804827432f052c8978da55f22d0ab55a723793f27715bb57e13ab0355
size 8383418912

3
phi-4-iq4_xs.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f275f0c6ad614f2c4529e61e1e6b1285654e4b3ffccf48586351f6efee66d1a
size 7941378592

3
phi-4-q2_k_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a1e82d3c8ddfa5652cb69b0afa188e8fedbf927b39ce82a5b2067ffd011d02b6
size 5175636512

3
phi-4-q3_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c978f8380cfdb0a8223f613cd3f7e8b0d81ba35d883dd2e2c04e0b1f4550e42a
size 7363269152

3
phi-4-q3_k_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8c5ffa10a46193b1aa73430b27bca827832866475378542d593d6c63688acaa4
size 6504747552

3
phi-4-q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e7fa09504627fead00d25df99d56d07dd462d7166d5431fe529158913c11c600
size 8250954272

3
phi-4-q4_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e656cb8c0b7c01805ac781e2561260756766818529ab00a7a75c8a559ad94133
size 9167147552

3
phi-4-q4_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fc8201b03376875035c0fd95e5c8bda959c349ba76720281ed04def622dcfcd4
size 9053114912

3
phi-4-q4_k_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a229279b34b6919156fb43b37c1762163b58baaa6e6e7cbc0a67d2f5a2a94bb0
size 8440762912

3
phi-4-q5_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6186793d45a3689a642c804c6cc09446a04db46831ecd3506b5d0d3e4e5f87f
size 10083340832

3
phi-4-q5_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a33a6bd3b8c2ba2658736acf1addd3a5957b1be0e13e7e09cd50d0536e2d94e
size 10999534112

3
phi-4-q5_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25bb049e4357069c3a7ea836f75a9dc33ff919c2e0e8924eb6fa4fe987bc5e74
size 10604188192

3
phi-4-q5_k_s.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3620e2065d7ebe21380636ad3fb41ad3695e1ca1d806c71113f2aa411f739b8d
size 10151580192

3
phi-4-q6_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6acb66d848ae5858e800c0ffe7c805709612a93529f7820f579e83b965052ed9
size 12030251552

3
phi-4-q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd7d20380713b76773f09a5e8739098c7d8d627c2ba86abf3bb30da078703e45
size 15580500224

3
phi-4-tq1_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4bd272f3f1a79f301030d6b3fc1a17af9ae4ed33f147e22377920c3096e5ad8
size 3591098912

3
phi-4-tq2_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7cd8a5b6db9285ec3ea318a8f69f684324c29c2913b8da692f9cddde751142d8
size 4230074912

3
phi-4.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57fcf7e1beb22782d41003f1d0b6dcdd9698068de5f6902ca42f8b48421d5087
size 5330288