license, language, library_name
license language library_name
mit
en
transformers

This is a QCQA version of the original model facebook/opt-125m. In this version, the original MHA architecture is preserved but instead of having a single K/V head, different K/V heads corresponding to the same group have the same mean-pooled K or V values. It has upto 6 groups of KV heads per layer instead of original 12 KV heads in the MHA implementation. This implementation is supposed to more efficient than corresponding GQA one.

Description
Model synced from source: xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache
Readme 619 KiB
Languages
Text 100%