diff --git a/experimental/cuda-lang/docs/source/execution.rst b/experimental/cuda-lang/docs/source/execution.rst new file mode 100644 index 0000000..ee59359 --- /dev/null +++ b/experimental/cuda-lang/docs/source/execution.rst @@ -0,0 +1,110 @@ +.. SPDX-FileCopyrightText: Copyright (c) <2026> NVIDIA CORPORATION & AFFILIATES. All rights reserved. +.. +.. SPDX-License-Identifier: Apache-2.0 + +.. currentmodule:: cuda.lang + +.. _execution-execution-model: + +Execution Model +=============== + +Abstract Machine +---------------- + +.. _thread: +.. _block: +.. _grid: + +A *SIMT kernel* is executed using the following thread hierarchy: + +- A *grid* is a 1D, 2D, or 3D collection of thread blocks. +- A *thread-block cluster* is an optional 1D, 2D, or 3D group of blocks within + the grid. +- A *thread block* is a 1D, 2D, or 3D collection of threads. +- A *warp* groups threads within a block for SIMT execution. +- A *thread* executes the body of the |kernel|. + +A thread can query its position in the thread, block, warp, and cluster +hierarchy using the functions listed under +:ref:`SIMT Model `. + +Threads in a block can communicate through shared memory and synchronize +explicitly. Blocks in the same cluster can map each other's shared memory and +synchronize at cluster scope. Global memory is accessible to all blocks, +subject to CUDA synchronization and memory-ordering rules. See the +:ref:`Data Model ` for the corresponding memory spaces. + + +.. _execution-execution-spaces: + +Execution Spaces +---------------- + +|cuda lang| code runs on one or more *targets*. A target is an execution +environment defined by its hardware resources and programming model. + +.. _host code: +.. _SIMT code: + +The set of targets where a construct can be used is called its *execution +space*. |cuda lang| defines two execution spaces: + +- *Host code* --- Python code that prepares data and launches kernels on a CPU. +- *SIMT code* --- code compiled for CUDA SIMT targets and executed by GPU + threads. + +Some constructs span multiple execution spaces. For example, :func:`cdiv` is +usable in both host code and SIMT code. + +A function whose decorator explicitly specifies its execution space is called +an *annotated function*. + + +.. _execution-simt-functions: + +SIMT Functions +-------------- + +A *SIMT function* is a helper function usable from SIMT code. The +``@cl.function`` decorator explicitly marks such a function. An undecorated +Python function called from a kernel or SIMT function is compiled recursively +as SIMT code, so an explicit decorator is not required. + + +.. _execution-simt-kernels: + +SIMT Kernels +------------ + +A *SIMT kernel* is an entry point executed by each thread in each block in a +grid. Kernels cannot be called directly. Use :func:`launch` to queue a kernel +for execution with explicit grid and block dimensions. + + +Python Subset +------------- + +|SIMT code| supports a subset of the Python language. There is no Python +runtime within SIMT code. Features such as exceptions and coroutines are not +supported today. + +Object Model & Lifetimes +~~~~~~~~~~~~~~~~~~~~~~~~ + +Arrays and pointers are views of memory and may alias or update the same +storage. Their visibility and lifetime depend on their CUDA memory space. +Because a launch is asynchronous with respect to the host, kernel arguments +must remain valid until the kernel completes. + +Control Flow +~~~~~~~~~~~~ + +Supported Python control flow statements include ``if``, ``while``, and +range-based ``for`` loops. + +Current limitations +^^^^^^^^^^^^^^^^^^^ + +The ``step`` of a ``range`` must be strictly positive. Negative-step ranges +such as ``range(10, 0, -1)`` are not supported today. diff --git a/experimental/cuda-lang/docs/source/index.rst b/experimental/cuda-lang/docs/source/index.rst index c15336d..2ac6544 100644 --- a/experimental/cuda-lang/docs/source/index.rst +++ b/experimental/cuda-lang/docs/source/index.rst @@ -71,6 +71,8 @@ memory: :maxdepth: 2 :hidden: + execution + metaprogramming data operations private diff --git a/experimental/cuda-lang/docs/source/metaprogramming.rst b/experimental/cuda-lang/docs/source/metaprogramming.rst new file mode 100644 index 0000000..c2cb417 --- /dev/null +++ b/experimental/cuda-lang/docs/source/metaprogramming.rst @@ -0,0 +1,51 @@ +.. SPDX-FileCopyrightText: Copyright (c) <2026> NVIDIA CORPORATION & AFFILIATES. All rights reserved. +.. +.. SPDX-License-Identifier: Apache-2.0 + +.. currentmodule:: cuda.lang + +.. _metaprogramming-metaprogramming: + +Metaprogramming +=============== + +|cuda lang| uses compile-time values and Python evaluation to specialize +generated SIMT code. + +Compile-Time Constants +---------------------- + +Some facilities require values that are known at compile time. Literals and +expressions derived entirely from compile-time constants remain known to the +compiler. A value that depends on run-time thread state is dynamic. + +Annotate a kernel parameter with ``cl.Constant[T]`` when its launch-time value +must be embedded in the compiled kernel. A distinct kernel is compiled for each +unique value of a constant-embedded parameter. + +Static Evaluation +----------------- + +``cl.static_eval`` evaluates an expression with Python semantics during +compilation. It can inspect compile-time properties of dynamic values and +select among symbolic expressions, but it cannot perform run-time GPU +operations. + +Static Iteration +---------------- + +``cl.static_iter`` evaluates an iterable during compilation and expands the +loop body once for each item. + +Static Assertions +----------------- + +``cl.static_assert`` checks a condition during compilation. +``cl.ensure_constant`` requires a value to be known at compile time and returns +that value. + +Static Functions +---------------- + +The ``@cl.static_def`` decorator marks a reusable helper function to be +evaluated with static-evaluation semantics. diff --git a/experimental/cuda-lang/docs/source/operations.rst b/experimental/cuda-lang/docs/source/operations.rst index f2fc7c4..e5ff37f 100644 --- a/experimental/cuda-lang/docs/source/operations.rst +++ b/experimental/cuda-lang/docs/source/operations.rst @@ -73,6 +73,8 @@ Type Casts bitcast +.. _operations-simt-model: + SIMT Model ---------- .. autosummary::