--- base_model: microsoft/phi-4 library_name: transformers license: mit pipeline_tag: text-generation tags: - positron-ai - quantized - gptq --- # Positron AI Quantized Build This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure. ## Recommended Use Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability. For general-purpose GPU inference, compare against the original model and other quantized formats before deployment. ## Artifact Summary | Field | Value | |---|---| | Base model | microsoft/phi-4 | | Published artifact | microsoft_phi-4-ingest-best-gptq | | Quantization method | GPTQ | | Quantization format | gptq | | Source precision | n/a | | Target runtime | tron | | Hardware target | FPGA | | Release date | 2026-06-30 | | License | mit | ## Quantization Details | Field | Value | |---|---| | Weight precision | 4-bit | | Activation precision | not quantized | | Bits | 4 | | Group size | 64 | | Symmetric quantization | true | | Activation ordering / desc_act | false | | Damp percent | 0.05 | | Calibration dataset | Universal mixed-domain set | | Calibration samples | 256 | | Calibration sequence length | 4096 | | MoE experts per token | n/a | | Quantization toolchain | GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8 | ## Validation Results | Metric | Result | Reference | Notes | |---|---|---|---| | Mean KL-divergence | 0.0183 | microsoft/phi-4 | Mean across the prompt suite | | P95 KL-divergence | 0.0753 | microsoft/phi-4 | Mean of per-prompt 95th-percentile token KL-divergence | | Top-1 agreement | 0.9523 | microsoft/phi-4 | Greedy top-1 token agreement | | Perplexity / NLL delta | +1.7% | microsoft/phi-4 | Same prompt suite as KL-divergence | | MMLU mean | pending | n/a | Evaluation pending | KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark. ## Evaluation Methodology | Field | Value | |---|---| | Evaluation date | 2026-06-30 | | Evaluation suite | Positron mixed-domain prompt suite | | Number of prompts | 12 | | Runtime | tron | | Device | FPGA | | Pass criteria | Measurement only (no fixed KL-divergence threshold) | ## Known Limitations - KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark. - Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference. - MMLU evaluation is pending; results will be added when available. ## Provenance This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.