初始化项目,由ModelHub XC社区提供模型

Model: positron-ai/microsoft_phi-4-ingest-best-gptq
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-10-01 13:54:04 +08:00
commit 93073e6d56
22 changed files with 603550 additions and 0 deletions

84
README.md Normal file
View File

@@ -0,0 +1,84 @@
---
base_model: microsoft/phi-4
library_name: transformers
license: mit
pipeline_tag: text-generation
tags:
- positron-ai
- quantized
- gptq
---
# Positron AI Quantized Build
This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure.
## Recommended Use
Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability.
For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.
## Artifact Summary
| Field | Value |
|---|---|
| Base model | microsoft/phi-4 |
| Published artifact | microsoft_phi-4-ingest-best-gptq |
| Quantization method | GPTQ |
| Quantization format | gptq |
| Source precision | n/a |
| Target runtime | tron |
| Hardware target | FPGA |
| Release date | 2026-06-30 |
| License | mit |
## Quantization Details
| Field | Value |
|---|---|
| Weight precision | 4-bit |
| Activation precision | not quantized |
| Bits | 4 |
| Group size | 64 |
| Symmetric quantization | true |
| Activation ordering / desc_act | false |
| Damp percent | 0.05 |
| Calibration dataset | Universal mixed-domain set |
| Calibration samples | 256 |
| Calibration sequence length | 4096 |
| MoE experts per token | n/a |
| Quantization toolchain | GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8 |
## Validation Results
| Metric | Result | Reference | Notes |
|---|---|---|---|
| Mean KL-divergence | 0.0183 | microsoft/phi-4 | Mean across the prompt suite |
| P95 KL-divergence | 0.0753 | microsoft/phi-4 | Mean of per-prompt 95th-percentile token KL-divergence |
| Top-1 agreement | 0.9523 | microsoft/phi-4 | Greedy top-1 token agreement |
| Perplexity / NLL delta | +1.7% | microsoft/phi-4 | Same prompt suite as KL-divergence |
| MMLU mean | pending | n/a | Evaluation pending |
KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark.
## Evaluation Methodology
| Field | Value |
|---|---|
| Evaluation date | 2026-06-30 |
| Evaluation suite | Positron mixed-domain prompt suite |
| Number of prompts | 12 |
| Runtime | tron |
| Device | FPGA |
| Pass criteria | Measurement only (no fixed KL-divergence threshold) |
## Known Limitations
- KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark.
- Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference.
- MMLU evaluation is pending; results will be added when available.
## Provenance
This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.