Files
project_6/cccl_upstream/docs/thrust/function_objects.rst

62 lines
2.2 KiB
ReStructuredText
Raw Normal View History

feat(cccl): integrate missing CCCL directories — python/, ci/, .agent/, docs/, test/ Sparse-checkout from NVIDIA/cccl main branch to complete cccl_upstream: Added: - python/cuda_cccl/ (226 files) — Python bindings for device-level algorithms Critical for muh toolchain: cuda.compute.reduce_into, scan, radix_sort, etc. Includes 204 .py files with full test coverage for all 27 algorithms - ci/ (163 files) — Build/test infrastructure build_cub.sh, test_cub.sh, build_and_test_targets.sh, matrix.yaml Directly maps to our [INFRA-CI] and [INFRA-BUILD] items - .agent/skills/ (7 files) — NVIDIA's own agent skills for CCCL cccl-style/SKILL.md, cccl-test/SKILL.md, sass-diff/SKILL.md - docs/ (491 files) — Official CCCL documentation CI references, CMake guides, Python compute docs, libcudacxx PTX docs - test/ (12 files) — Top-level integration tests (cuda_smoke, stdpar) - Root configs: .clang-format, .clang-tidy, CONTRIBUTING.md, pyproject.toml - CLAUDE.md symlink → AGENTS.md (NVIDIA's standard) cccl_upstream now mirrors full NVIDIA/cccl structure: Before: 42M (cub + thrust + libcudacxx + cudax + c + examples + benchmarks) After: 53M (+python +ci +docs +.agent +test +configs) This completes the CCCL base needed for: - [muh-bench] items: ci/util/build_and_test_targets.sh for targeted builds - [CCCL-verify] items: python/cuda_cccl/tests/ as reference implementations - [CCCL-test] items: ci/test_cub.sh, ci/test_thrust.sh - Agent workflow: .agent/skills/ for consistent style and test patterns
2026-08-07 02:34:33 +00:00
.. _thrust-module-api-function-objects:
Function Objects
=================
.. toctree::
:glob:
:maxdepth: 2
function_objects/adaptors
function_objects/placeholder
function_objects/predefined
.. _address-stability:
Copyable arguments
------------------
The C++ language allows to take the address of a parameter and depend on this value for the correctness of a code.
Consider this example:
.. code-block:: cpp
const int n = 10;
thrust::device_vector<int> a(n, 1);
thrust::device_vector<int> b(n);
int* a_ptr = thrust::raw_pointer_cast(a.data());
int* b_ptr = thrust::raw_pointer_cast(b.data());
thrust::transform(thrust::device, a.begin(), a.end(), a.begin(),
[a_ptr, b_ptr](const int& e) {
const auto i = &e - a_ptr; // &e expected to point into global memory
return e + b_ptr[i];
});
Here, :code:`thrust::transform` is invoked on the range of elements in :code:`a`.
The lambda function computes the index :code:`i` based on the start of the buffer held by :code:`a`
and the address of the parameter :code:`e`,
thus assuming that the reference :code:`e` points into the same memory block that :code:`a` holds,
e.g., global memory for the CUDA system.
While this example is contrived, such uses of Thrust exist and are currently valid.
We strongly urge users though to not rely on parameter addresses,
and we reserve the right to disallow this guarantee in the future.
Relying on the address of a parameter constrains the internal implementation
to serve the arguments to the callable directly from the input buffer,
which inhibits optimizations, like bulk copies or vectorized loads.
To permit the implementation to take advantage of such features,
a function object can be marked using :code:`proclaim_copyable_arguments`:
.. code-block:: cpp
thrust::transform(thrust::device, a.begin(), a.end(), a.begin(),
cuda::std::proclaim_copyable_arguments([](const int& a, const int& b) {
return a + b;
}));
Wrapping a function object in :code:`proclaim_copyable_arguments` will attach a marker that the implementation can detect,
and use for optimization.
Many function objects in libcu++, CUB and Thrust are marked by default,
but it does not hurt to mark them explicitly.