Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 13 additions & 10 deletions docs/snippets/example/050_kernel.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -43,38 +43,41 @@ TEMPLATE_LIST_TEST_CASE("tutorial kernel intro vector add", "[docs]", docs::test
onHost::concepts::Device auto device = selector.makeDevice(0);
onHost::Queue queue = device.makeQueue();

// BEGIN-TUTORIAL-allocateBuffers
// When you use `makeIdxMap`, your algorithm is not subject to any restrictions regarding data size, such as the
// requirement that the data must be a power of two or a multiple of 100.
size_t numElements = 293u;
auto lhsBuffer = onHost::alloc<int>(device, numElements);
// Allocate buffers on the compute device.
auto rhsBuffer = onHost::allocLike(device, lhsBuffer);
auto resultBuffer = onHost::allocLike(device, lhsBuffer);
// END-TUTORIAL-allocateBuffers

std::vector<int> lhs(numElements);
std::vector<int> rhs(numElements);
std::iota(lhs.begin(), lhs.end(), 0);
std::iota(rhs.begin(), rhs.end(), 1000);
std::vector<int> result(lhs.size(), -1);

// BEGIN-TUTORIAL-copyToDevice
// Copy input data to the device.
onHost::memcpy(queue, lhsBuffer, lhs);
onHost::memcpy(queue, rhsBuffer, rhs);
onHost::memset(queue, resultBuffer, 0x00);
// END-TUTORIAL-copyToDevice

// BEGIN-TUTORIAL-kernelFrameSpec
// BEGIN-TUTORIAL-kernelLaunch
// Let alpaka calculate a well-functioning `frameSpec` for you.
// This assumes that you are using `onAcc::makeIdxMap` in the kernel.
onHost::concepts::FrameSpec auto frameSpec = onHost::getFrameSpec<int>(device, Vec{numElements});
// END-TUTORIAL-kernelFrameSpec

// BEGIN-TUTORIAL-kernelLaunch
// Create a kernel object and enqueue it along with the `frameSpec´ and kernel arguments.
// Depending on how many tasks are still in the queue, the kernel may be executed immediately or after a delay.
queue.enqueue(frameSpec, KernelBundle{VectorAddKernel{}, resultBuffer, lhsBuffer, rhsBuffer});
// If you use a non-blocking queue, the kernel runs asynchronously with respect to the host.
// To synchronize the kernel, you must call `onHost::wait(queue)`.
// onHost::wait(queue);
// END-TUTORIAL-kernelLaunch

// BEGIN-TUTORIAL-copyFromDevice
// Copy the result back and wait for completion before reading it.
onHost::memcpy(queue, result, resultBuffer);
onHost::wait(queue);
// END-TUTORIAL-copyFromDevice


for(size_t i = 0; i < result.size(); ++i)
{
Expand Down
10 changes: 10 additions & 0 deletions docs/source/basic/terms.rst
Original file line number Diff line number Diff line change
Expand Up @@ -211,6 +211,16 @@ Go to the `IBuffer Interface definition <https://alpaka3.readthedocs.io/en/lates
Kernel
------

.. _executor:

Executor
--------

.. _thread_spec:

Thread Spec
-----------

.. _frame:

Frame, Frame Extents and Frame Spec
Expand Down
181 changes: 62 additions & 119 deletions docs/source/tutorial/kernel.rst
Original file line number Diff line number Diff line change
Expand Up @@ -2,145 +2,88 @@ Kernel
======

After selecting a device, creating a queue, and allocating memory, the next step is to launch work on the device.
In *alpaka*, the simplest kernel is usually just a small function object plus a host-side launch with a ``FrameSpec``.
If you know CUDA, a frame is understood as a logical index distribution, not as an exact block/grid launch description.
In *alpaka*, :ref:`kernels <kernel>` are implemented as function objects, known as *functors*.
A *functor* is a C++ ``class`` or ``struct`` that contains at least the member function ``void operator() const``, which implements the algorithm to be executed on the :ref:`device`.
A *functor* can be extended with additional functionality and is limited only by the requirement that it must be `trivially copyable <https://en.cppreference.com/cpp/language/classes#Trivially_copyable_class>`__.
In this tutorial, we will focus on a very simple :ref:`kernel`.

What matters early is that a ``FrameSpec`` is **not** required to describe the whole problem size.
It describes the maximum parallelism that alpaka can make available to the kernel at one time.
The real number of thread blocks and threads per block is derived by alpaka based on the device kind and API of the queue.
It does not guarantee how many physical thread blocks the backend will use or how large those thread blocks are.
The actual problem processed within the kernel can be much larger.
The kernel then uses ``makeIdxMap`` to walk over the complete data index range.
The :ref:`kernel` is implemented using the same principle as CUDA, HIP, or SYCL. The function specifies which part of a problem a single hardware thread processes based on its ID. To execute the algorithm, the :ref:`kernel` is launched in parallel. Therefore, the same algorithm is executed n times, each time with a different thread ID.

What a Beginner Kernel Looks Like
---------------------------------
A :ref:`FrameSpec <frame>` is used to describe the parallelism of a :ref:`kernel` launch.

Most first kernels in alpaka end up looking almost the same:
Writing the Kernel
------------------

- The kernel is a function object with ``ALPAKA_FN_ACC void operator() const``.
- The first argument is the accelerator handle ``acc`` followinfg the concepts ``onAcc::conepsts::Acc``.
- Output buffers use ``IMdSpan`` and input buffers use ``IDataSource``.
- Thread group mapping and work distribution is expressed with ``onAcc::makeIdxMap(...)``.
- The kernel body only talks about data indices, not about raw block and thread IDs.
An *alpaka* :ref:`kernel` must meet the following requirements:

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelStructure
:end-before: END-TUTORIAL-kernelStructure
:dedent:
- The kernel must be a function object. Therefore, it must implement the function call operator with the following signature: ``ALPAKA_FN_ACC void operator()(onAcc::conepsts::Acc auto acc, ...) const`` or ``constexpr void operator()(onAcc::conepsts::Acc auto acc, ...) const``.
- The *functor* must be `trivially copyable <https://en.cppreference.com/cpp/language/classes#Trivially_copyable_class>`__.
- The first argument is the accelerator handle ``acc``, which follow the concepts ``onAcc::conepsts::Acc``.
- The :doc:`Device Memory <memoryAllocation>` handles are passed via the arguments [#f1]_. Typically, output buffers use ``IMdSpan`` and input buffers use ``IDataSource``.

This most important rule in *alpaka* is: write the kernel in terms of the data that needs to be processed.
``makeIdxMap`` distributes that work over chosen thread groups based on the APi, deviceKind and executor.
That keeps the code portable across CPUs and GPUs and is usually better than manual thread index arithmetic.
.. [#f1] We need to pass a handle instead of copying the memory directly, because the kernel must be trivially copyable. The simplest handle is a raw pointer. We strongly advise against using raw pointers, as they have many drawbacks and are a major source of errors.

Launching the Kernel
--------------------

On the host side, the pattern is straightforward:

1. Allocate buffers on the compute device.

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-allocateBuffers
:end-before: END-TUTORIAL-allocateBuffers
:dedent:

2. Copy input data to the device.

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-copyToDevice
:end-before: END-TUTORIAL-copyToDevice
:dedent:

3. Choose a frame specification.

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelFrameSpec
:end-before: END-TUTORIAL-kernelFrameSpec
:dedent:

4. Enqueue the kernel.
.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelStructure
:end-before: END-TUTORIAL-kernelStructure
:dedent:
Comment on lines -36 to +30

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not against to show some code around the kernel launch. But I would keep it one code block, because it is much better readable. The only problem is, we cannot define a code region in a code region, without that the outer code region will display the BEGIN and END mark.

// BEGIN-TUTORIAL-kernelaround
// Copy input data to the device.
onHost::memcpy(queue, lhsBuffer, lhs);
onHost::memcpy(queue, rhsBuffer, rhs);
// ...

// BEGIN-TUTORIAL-kernelLaunch
// Let alpaka calculate a well-functioning `frameSpec` for you.
// This assumes that you are using `onAcc::makeIdxMap` in the kernel.
onHost::concepts::FrameSpec auto frameSpec = onHost::getFrameSpec<int>(device, Vec{numElements});

// Create a kernel object and enqueue it along with the `frameSpec´ and kernel arguments.
// Depending on how many tasks are still in the queue, the kernel may be executed immediately or after a delay.
queue.enqueue(frameSpec, KernelBundle{VectorAddKernel{}, resultBuffer, lhsBuffer, rhsBuffer});
// If you use a non-blocking queue, the kernel runs asynchronously with respect to the host.
// To synchronize the kernel, you must call `onHost::wait(queue)`.
// onHost::wait(queue);
// END-TUTORIAL-kernelLaunch

// Copy the result back and wait for completion before reading it.
onHost::memcpy(queue, result, resultBuffer);
onHost::wait(queue);
// BEGIN-TUTORIAL-kernelaround

If we display the code region BEGIN-TUTORIAL-kernelaround - BEGIN-TUTORIAL-kernelaround, the code will also show the comments BEGIN-TUTORIAL-kernelLaunch and BEGIN-TUTORIAL-kernelLaunch` in the source code.

I can bring back it. But this takes a little bit work. Also at the end of the page, there is the whole source code to provide the context and we already assume people read this code to understand the context on the pages before. Therefore I think it is not necessary.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's why it is split into sections. It is a workaround for the fact that we can not handle nested includes.
This can be revisited later and is IMO not mandatory.

@SimeonEhrig SimeonEhrig May 11, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the technical perspective, that makes sense. But if we want to have the this documentation, I suggest simply to copy the code and set different markers. The maintenance overhead is small.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ok kept yours


.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelLaunch
:end-before: END-TUTORIAL-kernelLaunch
:dedent:
*alpaka* offers many different ways to write a :ref:`kernel`, but it also provides several functions to help you create a kernel easily and securely.
``onAcc::makeIdxMap`` is the most important helper function for writing a kernel.
It distributes a range among the workers. The second argument defines who a worker is.

5. Copy the result back and wait for completion before reading it.
For the simplest kernel, we use the worker group ``onAcc::worker::threadsInGrid``, which means that the range is distributed across all global threads [#f2]_.
In the example, the range is the extents of ``out``.
We access all elements of ``out`` in parallel and write the sum of ``lhs`` and ``rhs`` to data position ``i``.
The parallelism is defined by the :ref:`FrameSpec <frame>` when the :ref:`kernel` starts.
Within the :ref:`kernel`, we work with the parallelism defined by the :ref:`ThreadSpec <thread_spec>`, which is derived from the :ref:`FrameSpec <frame>`.
The :ref:`tutorial_launch_kernel` section explains how this derivation works.

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-copyFromDevice
:end-before: END-TUTORIAL-copyFromDevice
:dedent:
The function call ``onAcc::makeIdxMap(acc, onAcc::worker::threadsInGrid, IdxRange{out.getExtents()})`` provides us with certain guarantees:

- All elements in ``out`` are processed independently of the parallelism specified via :ref:`FrameSpec <frame>` [#f3]_.
- This ensures that no out-of-range memory access happens.
- Memory access is optimized based on the used :ref:`device` and :ref:`executor`. For example, a blocking access pattern is used on the CPU, while a grid-stride loop is used on the GPU.

The queue can be non-blocking, so ``alpaka::onHost::wait(queue)`` is the point where the host knows the device work is finished.
Without that synchronization, reading the result on the host can race with the running kernel.
``onAcc::makeIdxMap`` offers many more features needed for performance optimization.
Once you've finished the tutorial, check out the :doc:`chunked` section to discover the full potential of this function.

What ``FrameSpec`` Means
------------------------
.. [#f2] The total number of threads is calculated by multiplying the number of threads by the number of blocks in a :ref:`thread_spec`.
.. [#f3] It is also possible to configure a :ref:`FrameSpec <frame>` with the values `{1,1}`, which means that all elements in ``onAcc::makeIdxMap`` are processed sequentially.

``FrameSpec`` is the maximum parallelism exposed to the kernel.
It is not a promise that the total problem size is exactly equal to ``frameCount * frameExtent``.
It is also not a promise that the launch will use exactly ``frameCount`` thread blocks of size ``frameExtent``.
If the frame extent is given as a compile-time :doc:`CVec <vector>`, that extent is also available as compile-time information
inside the kernel.
.. _tutorial_launch_kernel:

Start with the following points in mind:

- Chose a reasonable parallel launch shape, e.g. the expected problem size device by a frame extent, often 256 elements per frame but at least one frame.
- In the kernel describes the full valid data range with ``makeIdxMap()`` and ``IdxRange{...}``.
- If the problem is larger than the immediate launch shape, the workers simply iterate until the whole range is covered.

That is why a kernel in this example can process a vector of length ``293`` even if the frame extent is something like ``128`` or ``256``.
The frame specification limits the available parallelism per launch shape.
It does not limit the logical size of the problem.

Choosing the Correct Frame Specification
----------------------------------------

For a first implementation, frame selection should be boring.
The host chooses how much work is grouped into one frame, and the kernel then iterates over the valid data indices assigned to it.

.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelFrameSpec
:end-before: END-TUTORIAL-kernelFrameSpec
:dedent:

Rules of thumb:

- ``onHost::getFrameSpec<T>(device, extents)`` is the easiest way to get a reasonable first frame specification.
- Start with simple sizes. For 1D kernels, something around ``128`` to ``256`` elements per frame is usually a reasonable first try.
- When you have multiple dimensions, prefer more work in the fastest varying dimension, which is usually ``x``.
- Use a compile-time ``CVec`` frame extent when the kernel benefits from knowing the frame size at compile time.
Comment on lines -104 to -121

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed the whole section, because it is not easy to explain how to select a good Frame Spec without explaining a lot of details. In my opinion, it is super frustrating that the tutorial tell you, that you have to select a good Frame Spec but does not explain how. And I think onHost::getFrameSpec does a good job at this point.

Launching the Kernel
--------------------

If you have seen CUDA-style beginner code, this is one of the major differences in style.
You do not start by hand-writing a global-index formula and hoping the launch exactly matches the problem.
Instead, you choose a sensible frame shape and let ``makeIdxMap`` carry that parallelism across the full problem range.
On the host side, we first need to create a :ref:`FrameSpec <frame>`.
A :ref:`FrameSpec <frame>` describes the maximum possible parallelism.
A :ref:`FrameSpec <frame>` contains the number of :ref:`Frames <frame>` and the :ref:`Frame Extents <frame>`.
When the :ref:`kernel` starts, the :ref:`FrameSpec <frame>` is reduced to a :ref:`thread_spec`.
Therefore, the number of :ref:`Frame Extents <frame>` is reduced to the number of threads, and the number of :ref:`Frames <frame>` is reduced to the number of blocks.
This reduction depends on the :ref:`device` and the :ref:`executor`.
The numbers from the :ref:`Frame Extents <frame>` can be directly transferred to the :ref:`thread_spec`, but the number of blocks and/or threads may also be smaller than in the :ref:`FrameSpec <frame>`.

In practice, choose the frame from the data layout first and only tune it later if profiling gives you a reason.
The same ``FrameSpec`` can run on different backends, but it may not be equally good everywhere.
A shape that feels natural for CUDA or HIP will run correctly on CPU backends, just with different performance characteristics.
Let's take, for example, a :ref:`Frame Spec <frame>` with the dimensions {10, 32} (number of Frames, Frame extents):

Once you are comfortable with this basic launch style, the next important alpaka step is :doc:`chunked`, where frames are treated as reusable tiles of work.
- On a GPU, :ref:`thread_spec` could take the value ``{10, 32}`` (blocks, threads), since we can launch 10 blocks and each GPU core has 32 threads.
- On a CPU, the :ref:`thread_spec` could be ``{10, 1}`` (blocks, threads), since we can run 10 blocks on 10 or fewer cores, but each core can have only one thread.

How ``makeIdxMap`` Helps
------------------------
``onAcc::makeIdxMap`` allows the use of any number of *blocks* and *threads*.
Therefore, any :ref:`FrameSpec <frame>` will work.
The size of the :ref:`FrameSpec <frame>` is only relevant for performance.
To begin with, you should use the ``onHost::getFrameSpec`` function to create a :ref:`FrameSpec <frame>`.
This function assumes that you are using the ``onAcc::makeIdxMap`` function in your :ref:`kernel` and returns a well-functioning :ref:`FrameSpec <frame>` depending on the :ref:`device` and :ref:`executor`.

``makeIdxMap`` is the beginner-friendly way to iterate over the part of the problem assigned to the running workers.
Conceptually, it gives you the portable version of the "grid-stride loop" idea that CUDA users often learn early:
all threads within a group cooperate to cover the whole range, and the loop only yields valid indices.
.. literalinclude:: ../../snippets/example/050_kernel.cpp
:language: cpp
:start-after: BEGIN-TUTORIAL-kernelLaunch
:end-before: END-TUTORIAL-kernelLaunch
:dedent:

The object that describes that iteration space in the tutorial examples is ``IdxRange``.
``IdxRange{out.getExtents()}`` means "the full valid index range of this output object".
For a vector, that is all indices from the first element to the last element.
For a matrix or image, it is the full multidimensional box of valid coordinates.
In the advanced area, section :doc:`chunked`, you'll learn how ``onAcc::makeIdxMap`` works in detail.
Once you understand this, you can manually set a :ref:`FrameSpec <frame>` that might work better than the :ref:`FrameSpec <frame>` returned by ``onHost::getFrameSpec``.

Typical Beginner Mistakes
-------------------------
Expand All @@ -149,7 +92,7 @@ Typical Beginner Mistakes
- Forgetting to copy the result back to the host after the kernel.
- Forgetting to wait before reading host-side results from a non-blocking queue.
- Choosing a one-dimensional frame for naturally multidimensional code and then reimplementing manual index arithmetic in the kernel.
- Writing the kernel in terms of raw thread IDs even though the algorithm is just "process every element once".
- Calculate your own thread ID and use it to access the data, rather than accessing the data via ``onAcc::makeIdxMap``.

Complete Source File
--------------------
Expand Down