初始化项目,由ModelHub XC社区提供模型
Model: positron-ai/microsoft_phi-4-ingest-best-gptq Source: Original Platform
This commit is contained in:
84
README.md
Normal file
84
README.md
Normal file
@@ -0,0 +1,84 @@
|
||||
---
|
||||
base_model: microsoft/phi-4
|
||||
library_name: transformers
|
||||
license: mit
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- positron-ai
|
||||
- quantized
|
||||
- gptq
|
||||
---
|
||||
|
||||
# Positron AI Quantized Build
|
||||
|
||||
This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure.
|
||||
|
||||
## Recommended Use
|
||||
|
||||
Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability.
|
||||
|
||||
For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.
|
||||
|
||||
## Artifact Summary
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Base model | microsoft/phi-4 |
|
||||
| Published artifact | microsoft_phi-4-ingest-best-gptq |
|
||||
| Quantization method | GPTQ |
|
||||
| Quantization format | gptq |
|
||||
| Source precision | n/a |
|
||||
| Target runtime | tron |
|
||||
| Hardware target | FPGA |
|
||||
| Release date | 2026-06-30 |
|
||||
| License | mit |
|
||||
|
||||
## Quantization Details
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Weight precision | 4-bit |
|
||||
| Activation precision | not quantized |
|
||||
| Bits | 4 |
|
||||
| Group size | 64 |
|
||||
| Symmetric quantization | true |
|
||||
| Activation ordering / desc_act | false |
|
||||
| Damp percent | 0.05 |
|
||||
| Calibration dataset | Universal mixed-domain set |
|
||||
| Calibration samples | 256 |
|
||||
| Calibration sequence length | 4096 |
|
||||
| MoE experts per token | n/a |
|
||||
| Quantization toolchain | GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8 |
|
||||
|
||||
## Validation Results
|
||||
|
||||
| Metric | Result | Reference | Notes |
|
||||
|---|---|---|---|
|
||||
| Mean KL-divergence | 0.0183 | microsoft/phi-4 | Mean across the prompt suite |
|
||||
| P95 KL-divergence | 0.0753 | microsoft/phi-4 | Mean of per-prompt 95th-percentile token KL-divergence |
|
||||
| Top-1 agreement | 0.9523 | microsoft/phi-4 | Greedy top-1 token agreement |
|
||||
| Perplexity / NLL delta | +1.7% | microsoft/phi-4 | Same prompt suite as KL-divergence |
|
||||
| MMLU mean | pending | n/a | Evaluation pending |
|
||||
|
||||
KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark.
|
||||
|
||||
## Evaluation Methodology
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Evaluation date | 2026-06-30 |
|
||||
| Evaluation suite | Positron mixed-domain prompt suite |
|
||||
| Number of prompts | 12 |
|
||||
| Runtime | tron |
|
||||
| Device | FPGA |
|
||||
| Pass criteria | Measurement only (no fixed KL-divergence threshold) |
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark.
|
||||
- Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference.
|
||||
- MMLU evaluation is pending; results will be added when available.
|
||||
|
||||
## Provenance
|
||||
|
||||
This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.
|
||||
Reference in New Issue
Block a user