xc-llm-kunlun

Author	SHA1	Message	Date
Li Wei	71bd70ad6c	[Feature] support compressed-tensors w4a16 quantization (#154 ) - native int4 kimi model inference is supported Signed-off-by: Li Wei <liwei.109@outlook.com>	2026-01-27 19:56:22 +08:00
fromck	74d4f804e8	add 2 kernels and optimize the calculation of topk_indices (#134 ) Co-authored-by: chengxiaokang <chengxiaokang@baidu.com>	2026-01-22 10:29:28 +08:00
Li Wei	8f56cbf3ed	[refactor]update Kunlun classes with monkey patch (#122 ) Signed-off-by: Li Wei <liwei.109@outlook.com>	2026-01-19 20:24:19 +08:00
baoqian426	eb40e8a07a	[Bugfix] fix can not import compressed_tensors (#87 ) Co-authored-by: root <root@rdtest-node1150.bcc-zwlt.baidu.com>	2026-01-07 11:32:10 +08:00
Li Wei	1c1b84d78c	[fix]update compressed-tensors scheme Deepseek v3.2 is supported now Signed-off-by: Li Wei <liwei.109@outlook.com>	2026-01-06 22:30:27 +08:00
Li Wei	515a4eeda9	[dev] support compressed-tensors w8a8 quantization (#75 ) * [dev] support compressed-tensors w8a8 quantization Co-authored-by: Li Wei <liwei.109@outlook.com> * [refact]update KunlunScaleMMKernel impl * [rebase]resolve conflicts and remove redundant code --------- Co-authored-by: tangshiwen <tangshiwen@baidu.com>	2026-01-06 13:51:53 +08:00