Sparse-checkout from NVIDIA/cccl main branch to complete cccl_upstream: Added: - python/cuda_cccl/ (226 files) — Python bindings for device-level algorithms Critical for muh toolchain: cuda.compute.reduce_into, scan, radix_sort, etc. Includes 204 .py files with full test coverage for all 27 algorithms - ci/ (163 files) — Build/test infrastructure build_cub.sh, test_cub.sh, build_and_test_targets.sh, matrix.yaml Directly maps to our [INFRA-CI] and [INFRA-BUILD] items - .agent/skills/ (7 files) — NVIDIA's own agent skills for CCCL cccl-style/SKILL.md, cccl-test/SKILL.md, sass-diff/SKILL.md - docs/ (491 files) — Official CCCL documentation CI references, CMake guides, Python compute docs, libcudacxx PTX docs - test/ (12 files) — Top-level integration tests (cuda_smoke, stdpar) - Root configs: .clang-format, .clang-tidy, CONTRIBUTING.md, pyproject.toml - CLAUDE.md symlink → AGENTS.md (NVIDIA's standard) cccl_upstream now mirrors full NVIDIA/cccl structure: Before: 42M (cub + thrust + libcudacxx + cudax + c + examples + benchmarks) After: 53M (+python +ci +docs +.agent +test +configs) This completes the CCCL base needed for: - [muh-bench] items: ci/util/build_and_test_targets.sh for targeted builds - [CCCL-verify] items: python/cuda_cccl/tests/ as reference implementations - [CCCL-test] items: ci/test_cub.sh, ci/test_thrust.sh - Agent workflow: .agent/skills/ for consistent style and test patterns
51 lines
2.0 KiB
ReStructuredText
51 lines
2.0 KiB
ReStructuredText
.. _cccl-runtime-legacy-resources:
|
|
.. _libcudacxx-extended-api-memory-resources-legacy-resources:
|
|
|
|
Legacy resources
|
|
================
|
|
|
|
Legacy memory resources provide synchronous allocation interfaces backed by the CUDA Runtime's legacy allocation APIs.
|
|
They are primarily intended for compatibility with older toolkits or platforms that do not support the newer memory
|
|
pool-based resources. Prefer the modern memory resources where available.
|
|
|
|
For the full memory resource model and property system, see
|
|
:ref:`Memory Resources (Extended API) <libcudacxx-extended-api-memory-resources>`.
|
|
|
|
:cpp:class:`cuda::mr::legacy_pinned_memory_resource`
|
|
------------------------------------------------------
|
|
.. _libcudacxx-memory-resource-legacy-pinned-memory-resource:
|
|
|
|
Provides pinned (page-locked) host allocations using ``cudaMallocHost`` and ``cudaFreeHost``. This resource is
|
|
*synchronous-only* and is intended as a compatibility fallback. For CUDA 12.9 and later, prefer
|
|
:cpp:any:`cuda::pinned_memory_resource`.
|
|
|
|
.. code:: cpp
|
|
|
|
#include <cuda/memory_resource>
|
|
|
|
void use_legacy_pinned() {
|
|
cuda::mr::legacy_pinned_memory_resource resource{};
|
|
void* ptr = resource.allocate_sync(1024, 64);
|
|
// Use memory...
|
|
resource.deallocate_sync(ptr, 1024, 64);
|
|
}
|
|
|
|
:cpp:class:`cuda::mr::legacy_managed_memory_resource`
|
|
-------------------------------------------------------
|
|
.. _libcudacxx-memory-resource-legacy-managed-memory-resource:
|
|
|
|
Provides managed (unified) allocations using ``cudaMallocManaged`` and ``cudaFree``. This resource is
|
|
*synchronous-only* and accepts the CUDA attachment flags (``cudaMemAttachGlobal`` / ``cudaMemAttachHost``). Prefer
|
|
:cpp:any:`cuda::managed_memory_resource` when available.
|
|
|
|
.. code:: cpp
|
|
|
|
#include <cuda/memory_resource>
|
|
|
|
void use_legacy_managed() {
|
|
cuda::mr::legacy_managed_memory_resource resource{cudaMemAttachGlobal};
|
|
void* ptr = resource.allocate_sync(1024, 64);
|
|
// Use memory...
|
|
resource.deallocate_sync(ptr, 1024, 64);
|
|
}
|