85 lines
3.0 KiB
Markdown
85 lines
3.0 KiB
Markdown
|
|
---
|
||
|
|
base_model: microsoft/phi-4
|
||
|
|
library_name: transformers
|
||
|
|
license: mit
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
tags:
|
||
|
|
- positron-ai
|
||
|
|
- quantized
|
||
|
|
- gptq
|
||
|
|
---
|
||
|
|
|
||
|
|
# Positron AI Quantized Build
|
||
|
|
|
||
|
|
This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure.
|
||
|
|
|
||
|
|
## Recommended Use
|
||
|
|
|
||
|
|
Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability.
|
||
|
|
|
||
|
|
For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.
|
||
|
|
|
||
|
|
## Artifact Summary
|
||
|
|
|
||
|
|
| Field | Value |
|
||
|
|
|---|---|
|
||
|
|
| Base model | microsoft/phi-4 |
|
||
|
|
| Published artifact | microsoft_phi-4-ingest-best-gptq |
|
||
|
|
| Quantization method | GPTQ |
|
||
|
|
| Quantization format | gptq |
|
||
|
|
| Source precision | n/a |
|
||
|
|
| Target runtime | tron |
|
||
|
|
| Hardware target | FPGA |
|
||
|
|
| Release date | 2026-06-30 |
|
||
|
|
| License | mit |
|
||
|
|
|
||
|
|
## Quantization Details
|
||
|
|
|
||
|
|
| Field | Value |
|
||
|
|
|---|---|
|
||
|
|
| Weight precision | 4-bit |
|
||
|
|
| Activation precision | not quantized |
|
||
|
|
| Bits | 4 |
|
||
|
|
| Group size | 64 |
|
||
|
|
| Symmetric quantization | true |
|
||
|
|
| Activation ordering / desc_act | false |
|
||
|
|
| Damp percent | 0.05 |
|
||
|
|
| Calibration dataset | Universal mixed-domain set |
|
||
|
|
| Calibration samples | 256 |
|
||
|
|
| Calibration sequence length | 4096 |
|
||
|
|
| MoE experts per token | n/a |
|
||
|
|
| Quantization toolchain | GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8 |
|
||
|
|
|
||
|
|
## Validation Results
|
||
|
|
|
||
|
|
| Metric | Result | Reference | Notes |
|
||
|
|
|---|---|---|---|
|
||
|
|
| Mean KL-divergence | 0.0183 | microsoft/phi-4 | Mean across the prompt suite |
|
||
|
|
| P95 KL-divergence | 0.0753 | microsoft/phi-4 | Mean of per-prompt 95th-percentile token KL-divergence |
|
||
|
|
| Top-1 agreement | 0.9523 | microsoft/phi-4 | Greedy top-1 token agreement |
|
||
|
|
| Perplexity / NLL delta | +1.7% | microsoft/phi-4 | Same prompt suite as KL-divergence |
|
||
|
|
| MMLU mean | pending | n/a | Evaluation pending |
|
||
|
|
|
||
|
|
KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark.
|
||
|
|
|
||
|
|
## Evaluation Methodology
|
||
|
|
|
||
|
|
| Field | Value |
|
||
|
|
|---|---|
|
||
|
|
| Evaluation date | 2026-06-30 |
|
||
|
|
| Evaluation suite | Positron mixed-domain prompt suite |
|
||
|
|
| Number of prompts | 12 |
|
||
|
|
| Runtime | tron |
|
||
|
|
| Device | FPGA |
|
||
|
|
| Pass criteria | Measurement only (no fixed KL-divergence threshold) |
|
||
|
|
|
||
|
|
## Known Limitations
|
||
|
|
|
||
|
|
- KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark.
|
||
|
|
- Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference.
|
||
|
|
- MMLU evaluation is pending; results will be added when available.
|
||
|
|
|
||
|
|
## Provenance
|
||
|
|
|
||
|
|
This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.
|