Files
ModelHub XC 93073e6d56 初始化项目,由ModelHub XC社区提供模型
Model: positron-ai/microsoft_phi-4-ingest-best-gptq
Source: Original Platform
2026-10-01 13:54:04 +08:00

3.0 KiB

base_model, library_name, license, pipeline_tag, tags
base_model library_name license pipeline_tag tags
microsoft/phi-4 transformers mit text-generation
positron-ai
quantized
gptq

Positron AI Quantized Build

This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure.

Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability.

For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.

Artifact Summary

Field Value
Base model microsoft/phi-4
Published artifact microsoft_phi-4-ingest-best-gptq
Quantization method GPTQ
Quantization format gptq
Source precision n/a
Target runtime tron
Hardware target FPGA
Release date 2026-06-30
License mit

Quantization Details

Field Value
Weight precision 4-bit
Activation precision not quantized
Bits 4
Group size 64
Symmetric quantization true
Activation ordering / desc_act false
Damp percent 0.05
Calibration dataset Universal mixed-domain set
Calibration samples 256
Calibration sequence length 4096
MoE experts per token n/a
Quantization toolchain GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8

Validation Results

Metric Result Reference Notes
Mean KL-divergence 0.0183 microsoft/phi-4 Mean across the prompt suite
P95 KL-divergence 0.0753 microsoft/phi-4 Mean of per-prompt 95th-percentile token KL-divergence
Top-1 agreement 0.9523 microsoft/phi-4 Greedy top-1 token agreement
Perplexity / NLL delta +1.7% microsoft/phi-4 Same prompt suite as KL-divergence
MMLU mean pending n/a Evaluation pending

KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark.

Evaluation Methodology

Field Value
Evaluation date 2026-06-30
Evaluation suite Positron mixed-domain prompt suite
Number of prompts 12
Runtime tron
Device FPGA
Pass criteria Measurement only (no fixed KL-divergence threshold)

Known Limitations

  • KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark.
  • Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference.
  • MMLU evaluation is pending; results will be added when available.

Provenance

This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.