|
|
c2de1c83b0
|
Utilize chunked prefill + K-tiling techniques to ensure 100K context
|
2026-06-05 17:00:41 +08:00 |
|
|
|
2d1ef50992
|
chunked prefill support and memory opts
|
2026-06-05 16:03:34 +08:00 |
|
|
|
8c047a70ea
|
some modifications to ensure 50K context input
|
2026-06-04 17:56:29 +08:00 |
|
|
|
1c33ef1355
|
add paged_attn
|
2026-05-29 16:53:39 +08:00 |
|