.. _cudax-places: Places ====== .. contents:: :depth: 2 Places are abstractions that represent where code executes and where data resides across the non-uniform memory of a CUDA system. They provide a unified interface for managing execution affinity, stream pools, memory allocation, and device context switching -- independently of any task-based programming model. Places come in two flavors: - **Execution places** (``exec_place``) determine where code is executed. - **Data places** (``data_place``) specify where data is located in memory. The places API is part of the ``cuda::experimental::places`` C++ namespace and can be used standalone via the ``cuda/experimental/places.cuh`` header, without pulling in the full CUDASTF task-graph framework. For backward compatibility, all places types are also available in the ``cuda::experimental::stf`` namespace. .. _places-execution-places: Execution places ---------------- An *execution place* describes a location where computation can occur. The following factory methods create the most common execution places: - ``exec_place::device(id)`` -- a specific CUDA device - ``exec_place::host()`` -- the host CPU - ``exec_place::current_device()`` -- the CUDA device that is currently active - ``exec_place::cuda_context(ctx, devid)`` -- an externally-owned CUDA driver context; the device ordinal is derived from the context when ``devid`` is omitted When an execution place is activated, it sets the appropriate CUDA context (e.g. calls ``cudaSetDevice``). Each execution place also has an *affine* data place: the memory location naturally associated with it. For a device execution place the affine data place is the device's global memory; for the host it is pinned host memory (RAM). A CUDA-context execution place is non-owning. The caller must keep the ``CUcontext`` alive while the place and any streams obtained from it are in use. .. _places-data-places: Data places ----------- A *data place* describes a memory location where data can reside. The following factory methods are available: - ``data_place::device(id)`` -- global memory of a specific CUDA device - ``data_place::host()`` -- pinned host memory - ``data_place::managed()`` -- CUDA managed (unified) memory - ``data_place::affine()`` -- the data place naturally associated with the current execution place The *affine* data place is the default: when no data place is specified, data is placed in the memory that is local to the execution place. For example, a task running on device 0 will access data in device 0's global memory by default. Non-affine placement is also supported: data can be placed on a different device or in host memory regardless of where the computation runs. This is useful for sparse accesses (leveraging CUDA Unified Memory page faulting) or for addressing memory capacity constraints. Non-affine placement assumes the hardware and OS support such accesses (NVLINK, UVM, etc.). .. _places-container-keys: Places as container keys ------------------------ Both ``exec_place`` and ``data_place`` can be used as keys in standard associative containers. The library provides the required comparison and hash support: - **``std::map``** and **``std::set``** use ``operator<`` (strict weak ordering) for keys. Both place types implement ``operator<``, so they can be used as ordered map or set keys. - **``std::unordered_map``** and **``std::unordered_set``** require a hash function and equality. The library specializes ``cuda::experimental::stf::hash`` for both place types (accessible from both the ``stf`` and ``places`` namespaces), and both implement ``operator==``. This allows, for example, maintaining per-place handles (e.g. CUBLAS or CUSOLVER handles keyed by ``exec_place``) or per-place caches keyed by ``data_place``, using either ordered or hash-based containers as needed. The following snippet shows lazy creation of a CUBLAS handle per execution place using an ``std::unordered_map`` keyed by ``exec_place``: .. code:: c++ #include #include using namespace cuda::experimental::places; cublasHandle_t& get_cublas_handle(const exec_place& ep = exec_place::current_device()) { static std::unordered_map> handles; auto& h = handles[ep]; if (h == cublasHandle_t{}) { exec_place_scope scope(ep); cuda_safe_call(cublasCreate(&h)); } return h; } .. _places-activate: Setting the current device or context -------------------------------------- The ``exec_place::activate()`` method provides a generic alternative to ``cudaSetDevice()`` that works uniformly across different execution place types. This is useful when you want to set the current CUDA device or context without using tasks. The method returns an ``exec_place`` representing the previous state, which can be used to restore the original device or context. **Behavior by execution place type:** - **Device places** (``exec_place::device(id)``): Calls ``cudaSetDevice(id)`` - **Green context places**: Sets the current CUDA driver context via ``cuCtxSetCurrent()`` - **Host places**: No-op **Basic usage with devices:** .. code:: cpp exec_place place = exec_place::device(1); exec_place prev = place.activate(); // Switch to device 1 // ... perform operations on device 1 ... place.deactivate(prev); // Restore previous device **Alternative restoration pattern:** You can also restore by calling ``activate()`` on the returned place: .. code:: cpp exec_place place = exec_place::device(1); exec_place prev = place.activate(); // ... work on device 1 ... prev.activate(); // Equivalent to place.deactivate(prev) **Usage with green contexts (CUDA 12.4+):** Green contexts provide SM-level partitioning of GPU resources. The ``activate()``/``deactivate()`` methods handle the underlying driver context management: .. code:: cpp // Create green contexts with 8 SMs each green_context_helper gc(8, device_id); auto view = gc.get_view(0); exec_place gc_place = exec_place::green_ctx(view); exec_place prev = gc_place.activate(); // Sets green context as current // ... GPU work runs with SM affinity ... gc_place.deactivate(prev); // Restore original context **RAII scope for scoped activation:** For exception-safe code or when you want automatic restoration, use the ``exec_place_scope`` RAII helper: .. code:: cpp { exec_place_scope scope(exec_place::device(1)); // Device 1 is now active // ... perform operations on device 1 ... } // Previous device is automatically restored when scope goes out of scope The guard automatically restores the previous execution place when it goes out of scope, making it useful for exception-safe code. .. _places-stream-management: Stream management with execution places ---------------------------------------- Execution places can be used independently of any task system to manage CUDA streams in a structured way. This is useful when you want to use place abstractions (devices, green contexts) for stream management without the full task-based programming model. Stream pools for pooled places (``device(N)``, ``host()``) live in an ``exec_place_resources`` registry that the caller owns. Pass the registry to ``exec_place::pick_stream`` to get a CUDA stream; the per-place pool inside the registry is created lazily on first request and is destroyed when the registry is destroyed. The method accepts an optional ``for_computation`` hint (defaults to ``true``) that may select between computation and data transfer stream pools to improve overlapping. This is purely a performance hint, and it does not affect correctness. Not all execution places enforce it. .. code:: cpp #include using namespace cuda::experimental::places; // Standalone use: own the registry yourself. exec_place_resources resources; // Get a stream from the current device exec_place place = exec_place::current_device(); cudaStream_t stream = place.pick_stream(resources); // Use the stream for CUDA operations myKernel<<>>(d_data); // Get streams from specific devices (sharing the same registry) cudaStream_t stream_dev0 = exec_place::device(0).pick_stream(resources); cudaStream_t stream_dev1 = exec_place::device(1).pick_stream(resources); Inside a CUDASTF context, the context's ``async_resources_handle`` already holds an ``exec_place_resources`` registry. Convenience overloads accept the handle directly so call sites do not have to dereference it: .. code:: cpp cudaStream_t stream = place.pick_stream(ctx.async_resources()); Stream pools are populated lazily -- CUDA streams are only created when first requested via ``pick_stream(resources)`` (or ``pick_stream(ctx.async_resources())`` inside CUDASTF). Self-contained places (``exec_place::cuda_stream(s)``, green-context places) ignore the registry and return their own embedded pool instead, so the user-provided ``cudaStream_t`` / ``CUgreenCtx`` must outlive any place that wraps it. .. _places-memory-allocation: Memory allocation with data places ------------------------------------ Data places provide a unified interface for memory allocation that works across different memory types (host, device, managed) and place extensions (green contexts, user-defined places). This allows you to allocate memory while benefiting from the place abstraction. The ``data_place::allocate()`` and ``data_place::deallocate()`` methods provide raw memory allocation. The stream parameter defaults to ``nullptr``, which is convenient for non-stream-ordered allocations (host, managed) where the stream is ignored: .. code:: cpp #include using namespace cuda::experimental::places; // Allocate on host (pinned memory) - stream defaults to nullptr void* host_ptr = data_place::host().allocate(1024); // ... use host_ptr ... data_place::host().deallocate(host_ptr, 1024); // Allocate on a specific device (stream-ordered) cudaStream_t stream; cudaStreamCreate(&stream); void* dev_ptr = data_place::device(0).allocate(1024, stream); // ... use dev_ptr with stream ... data_place::device(0).deallocate(dev_ptr, 1024, stream); cudaStreamDestroy(stream); // Allocate managed memory - stream defaults to nullptr void* managed_ptr = data_place::managed().allocate(1024); // ... use managed_ptr from host or device ... data_place::managed().deallocate(managed_ptr, 1024); **Stream-ordered vs immediate allocations:** Different data places have different allocation behaviors: - **Host** (``data_place::host()``): Uses ``cudaMallocHost()`` / ``cudaFreeHost()`` - immediate, stream parameter is ignored - **Managed** (``data_place::managed()``): Uses ``cudaMallocManaged()`` / ``cudaFree()`` - immediate, stream parameter is ignored (note: ``cudaFree`` may introduce implicit synchronization) - **Device** (``data_place::device(id)``): Uses ``cudaMallocAsync()`` / ``cudaFreeAsync()`` - stream-ordered - **Extensions** (green contexts, etc.): Behavior depends on the extension implementation You can query whether a place uses stream-ordered allocation with ``allocation_is_stream_ordered()``: .. code:: cpp data_place place = data_place::device(0); if (place.allocation_is_stream_ordered()) { // Allocation is stream-ordered - synchronize via the stream void* ptr = place.allocate(size, stream); myKernel<<>>(ptr); place.deallocate(ptr, size, stream); cudaStreamSynchronize(stream); } else { // Allocation is immediate - stream is ignored, safe to use right away void* ptr = place.allocate(size); // ... use ptr ... place.deallocate(ptr, size); } This abstraction is particularly useful when writing generic code that needs to work with different types of places, including custom place extensions. Geometry-aware allocation with allocate_nd ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ Some places need to know the shape of the tensor being allocated, not just its size: a composite data place distributes the allocation according to a partitioner that maps *element coordinates* to places. ``allocate_nd()`` takes the tensor extents (dimension 0 varying fastest) and the element size: .. code:: cpp // 2-D tensor of nx x ny doubles, distributed by the place's partitioner void* ptr = place.allocate_nd(dim4(nx, ny), sizeof(double)); // ... place.deallocate(ptr, nx * ny * sizeof(double)); For most places this is equivalent to ``allocate(prod(dims) * elemsize)``. For composite places it is required: the byte-count ``allocate()`` throws there, since a byte count alone cannot carry the geometry the partitioner needs. A caller that genuinely has untyped bytes states that explicitly with ``allocate_nd(dim4(nbytes), 1)``, which distributes the buffer with byte granularity. This raw-byte form applies to composite places built from scale-free partitioners only; a composite place backed by a structured partition (see :ref:`places-structured-partitions`) accepts exactly the extents of the tensor the partition was built for and rejects anything else, including a flat byte count. .. _places-vmm: VMM-based allocation with mem_create ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ For advanced use cases involving CUDA's Virtual Memory Management (VMM) API, ``data_place`` also provides the ``mem_create()`` method. This is a lower-level interface used internally by localized arrays (``composite_slice``) to create physical memory segments that are then mapped into a contiguous virtual address space. Unlike ``allocate()``, which returns a usable pointer directly, ``mem_create()`` returns a ``CUmemGenericAllocationHandle`` that must be subsequently mapped with ``cuMemMap()`` before use: .. code:: cpp #include using namespace cuda::experimental::places; // Create a physical memory handle for device 0 CUmemGenericAllocationHandle handle; data_place::device(0).mem_create(&handle, size); // The handle must be mapped to a virtual address before use // (see CUDA VMM documentation for cuMemMap, cuMemSetAccess, etc.) **When to use each method:** - Use ``allocate()`` for most cases - it provides ready-to-use memory with stream-ordered semantics where applicable. - Use ``mem_create()`` only when you need explicit control over virtual memory mapping, such as creating localized arrays that span multiple devices with a unified virtual address space. **Limitations of mem_create:** - Only supports device memory and host memory (pinned) - Managed memory is **not supported** by the VMM API - The returned handle requires additional VMM API calls to be usable Custom place extensions can override ``mem_create()`` to provide specialized VMM allocation behavior (e.g., memory localization for hardware partitions). .. _places-grid: Grid of places -------------- It is possible to manipulate places which are a collection of multiple places. In particular, it is possible to define an execution place which corresponds to multiple device execution places. Creating grids of places ^^^^^^^^^^^^^^^^^^^^^^^^ A grid of execution places is an ``exec_place`` that contains multiple underlying places. Grids are created with the ``make_grid`` free function: .. code:: c++ // Create a 1D grid from a vector of places exec_place grid = make_grid(std::vector{ exec_place::device(0), exec_place::device(1) }); The ``exec_place::all_devices()`` helper creates a grid of all available CUDA devices: .. code:: c++ exec_place all = exec_place::all_devices(); Similarly, ``exec_place::n_devices(n)`` creates a grid from the first ``n`` devices: .. code:: c++ exec_place first_four = exec_place::n_devices(4); It is possible to retrieve the total number of elements in a grid using the ``size()`` method, and individual places with ``get_place(i)``: .. code:: c++ exec_place grid = exec_place::all_devices(); for (size_t i = 0; i < grid.size(); i++) { exec_place dev = grid.get_place(i); // ... } Shaped grids ^^^^^^^^^^^^ Grids of places need not be 1D arrays. They can be structured as a multi-dimensional grid described with a ``dim4`` class by passing it to ``make_grid`` or ``n_devices``: .. code:: c++ // Create a shaped grid: 8 devices arranged as a 2x2x2 cube exec_place cube = exec_place::n_devices(8, dim4(2, 2, 2)); // Or from an explicit vector exec_place shaped = make_grid(my_places, dim4(4, 2)); Note that the total size of the ``dim4`` must match the number of places. It is possible to query the *shape* of the grid using ``get_dims()``, which returns a ``dim4`` object. Individual places can be accessed by multi-dimensional position using ``get_place(pos4)``. Reshaping and collapsing grid axes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ An existing grid can be viewed with different dimensions using ``reshape()``. The new dimensions must contain exactly the same number of places: .. code:: c++ exec_place cube = make_grid(my_places, dim4(2, 3, 4)); exec_place flat = cube.reshape(dim4(24)); Reshaping changes only the grid coordinate system. It preserves dimension-0- fastest linear order, so ``flat.get_place(i) == cube.get_place(i)`` for every linear index ``i``. It does not reorder, replicate, or remove places. ``collapse_axes(first, last)`` is a convenience operation that combines a contiguous inclusive range of axes. The collapsed extent is the product of the selected extents; later axes shift left and trailing extents become one: .. code:: c++ exec_place grid = make_grid(my_places, dim4(2, 3, 4)); exec_place grid_6x4 = grid.collapse_axes(0, 1); // dim4(6, 4) exec_place grid_2x12 = grid.collapse_axes(1, 2); // dim4(2, 12) exec_place grid_24 = grid.collapse_axes(0, 3); // dim4(24) These operations are useful when a partition should consume several axes of a processor grid as one logical axis. They are coordinate transformations, not :ref:`places-partitioning`: the latter decomposes a place into constituent resources. .. _places-partitioning: Partitioning grids ^^^^^^^^^^^^^^^^^^ The ``place_partition`` class partitions an execution place at a given granularity. This is useful for splitting a multi-device grid into its constituent devices, or for partitioning a device into green contexts or CUDA streams. The partitioning granularity is specified by ``place_partition_scope``: - ``place_partition_scope::cuda_device`` -- partition into individual devices - ``place_partition_scope::green_context`` -- partition into green contexts (CUDA 12.4+) - ``place_partition_scope::cuda_stream`` -- partition into CUDA streams .. code:: c++ exec_place grid = exec_place::all_devices(); // Partition into individual devices place_partition devices(grid, place_partition_scope::cuda_device); for (auto& dev : devices) { // dev is an exec_place for a single device } // Convert back to an exec_place grid exec_place new_grid = devices.to_exec_place(); The ``exec_place::partition_by_scope()`` method provides a shorthand that returns a new ``exec_place`` grid directly: .. code:: c++ exec_place grid = exec_place::all_devices(); exec_place by_device = grid.partition_by_scope(place_partition_scope::cuda_device); .. _places-data-partitioning: Data partitioning policies ^^^^^^^^^^^^^^^^^^^^^^^^^^ When using a grid of places with CUDASTF constructs such as ``parallel_for``, *data partitioning policies* express how data and index spaces are dispatched over the different places of a grid. .. code:: c++ class MyPartition : public partitioner_base { public: template static const S_out apply(const S_in& in, pos4 position, dim4 grid_dims); void get_executor(pos4* result, pos4 data_coords, dim4 data_dims, dim4 grid_dims); }; A partitioning class must implement an ``apply`` method which takes: - a reference to a shape of type ``S_in`` - a position within a grid of execution places, described using an object of type ``pos4`` - the dimension of this grid expressed as a ``dim4`` object ``apply`` returns a shape which corresponds to the subset of the ``in`` shape associated to this entry of the grid. Note that the output shape type ``S_out`` may be different from the ``S_in`` type of the input shape. To support different types of shapes, appropriate overloads of the ``apply`` method should be implemented. This ``apply`` method is typically used by the ``parallel_for`` construct in order to dispatch indices over the different places. A partitioning class must also implement the ``get_executor`` virtual method which allows localized data allocators. This method indicates, for each entry of a shape, on which place this entry should *preferably* be allocated. ``get_executor`` writes a ``pos4`` coordinate in the execution place grid into ``*result``, and its input arguments are: - a coordinate within the shape described as a ``pos4`` object - the dimension of the shape expressed as a ``dim4`` object - the dimension of the execution place grid expressed as a ``dim4`` object Defining the ``get_executor`` makes it possible to map a piece of data over an execution place grid. The ``get_executor`` method of a partitioning policy in an execution place grid therefore defines the *affine data place* of a logical data accessed on that grid. Predefined partitioning policies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There are currently two policies readily available: - ``tiled_partition`` dispatches entries of a shape using a *tiled* layout. For multi-dimensional shapes, the outermost dimension is dispatched into contiguous tiles of size ``TILE_SIZE``. - ``blocked_partition`` dispatches entries of the shape using a *blocked* layout, where each entry of the grid of places receives approximately the same contiguous portion of the shape, dispatched along the outermost dimension. This illustrates how a 2D shape is dispatched over 3 places using the blocked layout: .. code:: text __________________________________ | | | | | | | | | | | | | P 0 | P 1 | P 2 | | | | | | | | | |___________|___________|_________| This illustrates how a 2D shape is dispatched over 3 places using a tiled layout, where the dimension of the tiles is indicated by the ``TILE_SIZE`` parameter: .. code:: text ________________________________ | | | | | | | | | | | | | | | | | | | | | | P 0 | P 1 | P 2 | P 0 | P 1 |P2| | | | | | | | | | | | | | | |_____|_____|_____|_____|_____|__| .. _places-structured-partitions: Structured partitions --------------------- The classic partitioning policies above are *scale-free*: ``blocked_partition`` splits whatever shape it is handed, knows nothing about the tensor it will be applied to, and always dispatches along the outermost dimension. A *structured partition* (``cute_partition``) is the complementary tool: it describes, dimension by dimension, how **one specific tensor** maps onto a grid of places. .. code:: c++ using namespace cuda::experimental::places; // A 3-D tensor: dimension 1 blocked over the places of the grid, // dimensions 0 and 2 not distributed auto part = make_partition( dim4(nx, ny, nz), partition_spec{whole, blocked<0>, whole}, grid.get_dims()); Each entry in ``partition_spec`` selects a policy for the corresponding tensor dimension: ``whole`` (not distributed), ``blocked``, ``cyclic``, or ``block_cyclic(block_size)``. Rank, policy, mesh-axis, and leaf counts are preserved in the C++ type; tensor extents, strides, and block sizes remain runtime values. This is strictly more expressive than the classic policies -- splitting dimension 1 of a 3-D tensor, or mixing policies across dimensions, cannot be stated with ``blocked_partition``. The reference shape, padding, and predication ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ The first argument of ``make_partition`` is the tensor's extents: unlike a classic policy, a structured partition is **bound to one reference shape**, and remains the authority on it. This is a deliberate trade, and the source of most of the type's properties: - Split dimensions are *padded up to divisibility* (a 10-element dimension blocked over 3 places is treated as 12, in chunks of 4). Padding makes the underlying layout exact and bijective, which is what keeps every query closed-form: validation is a linear pass over the layout, and the owner of a coordinate is a chain of divisions and modulos. - Coordinates beyond the true extents (the *padding phantoms*) own no bytes and do no work: consumers discard them by comparing coordinates against the true extents. This is the *predication* idiom of CUTLASS/CuTe ("partition the rounded-up shape, predicate the boundary") rather than per-place clamping, which would break the layout's uniformity. Ownership can be queried directly, and -- more importantly -- a candidate mapping can be **scored before any memory is committed**: .. code:: c++ pos4 owner = part.owner(pos4(x, y, z)); // grid position owning (x,y,z) // Dry run: same block-majority decision procedure as a real allocation localized_stats stats = evaluate_localized_placement(grid, part, sizeof(double)); // stats.bytes_per_place, stats.accuracy() (estimated fraction of local bytes), // stats.nallocs, ... -- tune the spec, then allocate Placement through a structured partition ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ A structured partition can back a composite data place. Because the partition is bound to one tensor, such a place is *per-tensor* -- allocate with the partition's exact extents (compare with the classic composite place, which is a reusable shape-free policy): .. code:: c++ data_place dp = make_composite_data_place(grid, part); void* ptr = dp.allocate_nd(dim4(nx, ny, nz), sizeof(double)); // physical pages land on the place owning them, per the partition dp.deallocate(ptr, nx * ny * nz * sizeof(double)); Two structured composite places built from equal partitions compare equal, so they denote the same data placement wherever data places are compared. Conventions and limits ^^^^^^^^^^^^^^^^^^^^^^ - Extents follow the **dimension-0-fastest** linearization of ``dim4::get_index()`` (the convention of STF slices). A row-major front-end must present its *whole* description in this order -- the extents, the per-dimension ``partition_spec``, and any coordinates passed to ``owner()`` reverse together, since reversing only the extents would silently re-target each policy at the wrong axis. - At most 4 tensor dimensions (the ``pos4``/``dim4`` domain). - Typed partitions and their kernel-facing sub-shapes store exactly their layout leaves. Runtime interfaces (including C/Python opaque handles) erase them to a canonical descriptor only at the data-place boundary. - The partition object is trivially copyable and its queries are host/device callable. The ``partitioned_axpy`` example shows the intended workflow end to end: express the partition once, evaluate it, run tasks over data placed by it, and perform a raw geometry-aware allocation. Computing over structured partitions ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ The same ``parallel_for`` entry point that accepts the classic policies accepts a structured partition instance, which then decides **both** the per-place kernel decomposition and (through the task's affine data place) the placement of the data those kernels touch -- one object, both sides: .. code:: c++ // Every place computes exactly the coordinates it owns ctx.parallel_for(part, grid, lX.shape(), lX.write()) ->*[] __device__(size_t x, size_t y, size_t z, auto X) { ... }; The shape argument may also be a ``box`` describing a *region within the tensor the partition was built for* (validated by containment) -- e.g. the interior of a stencil domain. Each place still enumerates its own coordinates; those outside the region (like the padding phantoms of uneven extents) are skipped by a per-coordinate predicate, so iteration stays aligned with data ownership rather than re-splitting the region: .. code:: c++ box interior({1ul, nx - 1}, {1ul, ny - 1}, {1ul, nz - 1}); ctx.parallel_for(part, grid, interior, lX.rw())->*...; Predication has a cost proportional to the *rejected* fraction of the enumerated coordinates, which makes it the right tool for regions that are dense in their bounds (interiors: the rejected boundary shell is a surface-to-volume fraction) and the wrong tool for thin regions. For boundary-style updates -- a face of the domain, say -- prefer one of: - **fuse** the boundary handling into the volumetric kernel's body when the condition is cheap (application-dependent); - iterate the face with a **classic scale-free policy** (tight, no rejected coordinates) while an explicit dependency keeps placement on the partition's composite place: .. code:: c++ auto dist = make_composite_data_place(grid, part); box face({0ul, nx}, {0ul, ny}, {0ul, 1ul}); ctx.parallel_for(blocked_partition(), grid, face, lX.rw(dist))->*...; The face's few remote writes (places computing parts of a face another place owns) are typically negligible against the volumetric traffic. The ``fdtd_mgpu`` example demonstrates the full pattern: a single ``make_partition`` call decides which dimension splits for every task -- initialization over the full shape, updates over interior boxes, a point source -- and places the fields' data, so changing the distribution of the whole simulation is editing one ``partition_spec`` entry.