Files
project_6/cccl_upstream/.agent/skills/sass-diff/SKILL.md
muh-bot 2a7ca101d7 feat(cccl): integrate missing CCCL directories — python/, ci/, .agent/, docs/, test/
Sparse-checkout from NVIDIA/cccl main branch to complete cccl_upstream:

Added:
- python/cuda_cccl/ (226 files) — Python bindings for device-level algorithms
  Critical for muh toolchain: cuda.compute.reduce_into, scan, radix_sort, etc.
  Includes 204 .py files with full test coverage for all 27 algorithms
- ci/ (163 files) — Build/test infrastructure
  build_cub.sh, test_cub.sh, build_and_test_targets.sh, matrix.yaml
  Directly maps to our [INFRA-CI] and [INFRA-BUILD] items
- .agent/skills/ (7 files) — NVIDIA's own agent skills for CCCL
  cccl-style/SKILL.md, cccl-test/SKILL.md, sass-diff/SKILL.md
- docs/ (491 files) — Official CCCL documentation
  CI references, CMake guides, Python compute docs, libcudacxx PTX docs
- test/ (12 files) — Top-level integration tests (cuda_smoke, stdpar)
- Root configs: .clang-format, .clang-tidy, CONTRIBUTING.md, pyproject.toml
- CLAUDE.md symlink → AGENTS.md (NVIDIA's standard)

cccl_upstream now mirrors full NVIDIA/cccl structure:
  Before: 42M (cub + thrust + libcudacxx + cudax + c + examples + benchmarks)
  After:  53M (+python +ci +docs +.agent +test +configs)

This completes the CCCL base needed for:
- [muh-bench] items: ci/util/build_and_test_targets.sh for targeted builds
- [CCCL-verify] items: python/cuda_cccl/tests/ as reference implementations
- [CCCL-test] items: ci/test_cub.sh, ci/test_thrust.sh
- Agent workflow: .agent/skills/ for consistent style and test patterns
2026-08-07 02:34:33 +00:00

2.4 KiB

name, description
name description
sass-diff Use when asked to check for SASS (or PTX) changes between commits, branches, or a local changeset; guides normalization, comparison, and reporting of CUDA disassembly diffs.

SASS Diffs

Use this when asked to check for SASS changes between commits, branches or a local changeset.

Goal

Detect relevant changes in generated CUDA machine code (i.e. SASS) while filtering noise from addresses, symbols, metadata, etc. Any non-trivial change must be detected.

Inputs to establish

  • Compilation target under test
  • The CUDA SM architectures to compile for. Try to detect this from the code and offer the user a list of suggestions. The user must confirm or provide this list.
  • Baseline source (e.g. the previous commit/branch or the current commit without the changes in the working copy).
  • Comparison source (e.g. the current commit/branch or the current commit with the changes in the working copy).
  • Whether a SASS (default) or PTX diff is requested.

Disassembly listing generation

  • Compile both, the baseline and comparison source, with the same compiler flags and options. When not specified otherwise, lookup the options from compile_commands.json or the current build system (i.e. CMake files). Make sure the CUDA SM architectures (CMAKE_CUDA_ARCHITECTURES) are set to the user-provided or approved list.
  • Dump the disassembly from the binaries produced in the previous set using cuobjdump -sass or cuobjdump -ptx.

Comparison rules (what matters)

Ignore as trivial:

  • Register renaming with identical instruction sequence and operands.
  • Pure label renumbering or reordering of identical basic blocks.
  • Formatting-only differences or reordered symbol tables.
  • Changes to symbol names (global function names)

Reporting

  • If any non-trivial change was detected, report the top 5 regions where a non-trivial change was detected, including the name of the kernel they appeared in.
  • Provide a short summary of the diff type, including opcode changes, memory access size/cache policy changes, control-flow changes, register-count changes, spills/local memory, shared memory, and occupancy-relevant resource deltas.
  • Explicitly state if only noise was detected.
  • If you are not sure if the differences are impactful, show it and ask the user for guidance.
  • Keep the disassembly dumps available and tell the user where they can find them.