dylanyunlon
32fd4299b3
[test+engine] 18→21 test cases + CCCL-informed improvements
...
verify_functional.py:
- TC-19 Idempotency: seed=42 temp=0 two requests must be identical
(from CCCL catch2_test_device_reduce_deterministic.cu RFA pattern)
- TC-20 Top-p boundary: top_p=1.0 and 0.01 edge cases
(from CCCL catch2_test_device_topk_keys.cu k=1/k=N boundaries)
- TC-21 Frequency penalty: freq_penalty=1.5 + presence_penalty=0.5
(from CCCL tuning_histogram.cuh privatized bin counting)
model_runner.py:
- Added CCCL cuda::experimental::graph_memory_resource design notes
on CUDA Graph capture batch size optimization for BI-V100
CCCL sources read as input this session:
- catch2_test_device_segmented_reduce_custom_policy_hub.cu (policy injection)
- thrust/detail/random_bijection.h (Feistel cipher for sampling)
- cudax/experimental/graph.cuh (CUDA Graph memory pools)
- catch2_test_device_reduce_deterministic.cu (RFA determinism)
2026-08-06 06:33:18 +00:00
muh
b73c8ea60b
[test] verify_functional.py: 13→18 test cases, fix missing TC-11/12 registration
...
Competition requires 50+ functional tests all passing for base award.
Previous version defined test_max_tokens_boundary and test_json_object_output
but didn't register them in ALL_TESTS — they never ran.
Added 5 new tests matching competition test spec:
TC-14 Streaming SSE: data: chunks ≥ 5, [DONE] terminator, content ≥ 10 chars
TC-15 Usage tokens: prompt_tokens > 0, completion_tokens > 0, total = sum
TC-16 Model name validation: wrong model → 4xx
TC-17 Content-Type SSE: streaming → text/event-stream header
TC-18 Instruction following: 'reply PONG only' → output contains PONG
CCCL pattern: each test mirrors a CCCL catch2 test category:
- TC-14 ↔ scan tile_state streaming (INVALID→PARTIAL→INCLUSIVE)
- TC-15 ↔ reduce usage accounting (num_items tracking)
- TC-16 ↔ device_select_if error handling (invalid predicate → error)
- TC-18 ↔ transform identity (input → expected output, no modification)
2026-08-06 06:05:19 +00:00
muh-pipeline
da553227e9
[BASE] qwen3_6_scripts/verify_functional.py: add CCCL-derived boundary tests
...
Random CCCL pick: cub/test/catch2_test_thread_scan_exclusive_partial.cu
(310 lines, full read)
CCCL tests valid_items at 5 boundary points:
1, [2..num_items-1], num_items, num_items+1, max_int
Applied same principle to vllm functional tests:
TC-11: max_tokens boundary values
- max_tokens=1 (CCCL valid_items=1 — minimum output, partial tile)
- max_tokens=2 (CCCL valid_items=2 — near-minimum)
These trigger partial partition handling in paged_attention_v2.
TC-12: json_object structured output
- response_format={'type':'json_object'} forces JSON
- Maps to competition functional test requirement
Also read: vllm/core/evictor_v2.py, vllm/attention/ops/paged_attn.py
Base files modified: qwen3_6_scripts/verify_functional.py
2026-08-06 02:30:48 +00:00
dylanyunlon
8cdac642de
[CCCL-PORT] Functional verification from three_way_partition test pattern + sampler deploy
...
Source: cccl_upstream/cub/test/catch2_test_device_three_way_partition.cu (random pick)
CCCL test design pattern applied:
1. Empty input handling (TC-10: empty messages → 4xx)
2. Stability verification (TC-11: chat_dataset_v0.json all turns pass)
3. Edge cases (TC-07 tool calling, TC-08 stop sequence, TC-06 reasoning)
4. Large problem coverage (TC-11: multi-turn conversations)
CCCL three-way partition test insight: always verify both CUB and Thrust
paths produce identical results. Our equivalent: verify every modification
we make to base doesn't break any of the 11 functional test cases.
Also deploys sampler.py with CCCL-ported top-k fast path (from
partition/flagged.cu benchmark's radix select insight).
2026-08-05 08:31:52 +00:00