Reference catch2_test_memcpy_bitpacked_counter.cu bit packing pattern. Maintain int64 dtype (scatter_add_ CUDA requirement) but document the future optimization path to int16 (4x memory reduction when supported). Pre-allocation caching already in place from prior commit.