xc-llm-ascend/worker at v0.18.0 - xc-llm-ascend - Gitea: Git with a cup of tea

EngineX/xc-llm-ascend

Files

History

starkwj e17006077a fix multiproc executor determine kv cache memory & update Dockerfile

2026-04-24 12:56:40 +00:00

..

[v0.18.0][CI] Fix releases/v0.18.0 ci test only support vllm v0.18.0 (#7686 )

2026-03-26 18:36:04 +08:00

__init__.py

[Misc][V0 Deprecation] Remove Cache Engine Used for V0 Worker (#1878 )

2025-07-19 09:42:32 +08:00

block_table.py

[Hybrid] support prefix cache for Qwen3.5/Next with --mamba-cache-mode align (#7103 )

2026-03-15 09:44:09 +08:00

model_runner_v1.py

[Bugfix]v0.18.0 support FlashComm1 & DCP for Qwen (#7726 )

2026-03-29 15:59:19 +08:00

npu_input_batch.py

[Hybrid] support prefix cache for Qwen3.5/Next with --mamba-cache-mode align (#7103 )

2026-03-15 09:44:09 +08:00

pcp_utils.py

feat(attention_cp): support chunked prefill for Qwen3Next with PCP&DCP (#6900 )

2026-03-09 17:55:09 +08:00

worker.py

fix multiproc executor determine kv cache memory & update Dockerfile

2026-04-24 12:56:40 +00:00