Aligned Load & Store

The exercise

You’ll implement three kernels that compute the same memory-bound map, out[i] = a[i] * 2 + 1, over a 1M-element float32 buffer. They differ only in how they touch memory:

  1. scalar_kernel: one element per thread with no vectorization. This is the baseline.
  2. unaligned_kernel: vectorized by SIMD_WIDTH, but the access alignment is under-stated (scalar alignment), so the compiler emits scalar loads.
  3. aligned_kernel: the same vectorized kernel, with the alignment communicated, so the compiler emits a single vectorized load/store per chunk.

All three produce identical output. The point of the puzzle is everything that happens after you confirm that.

def scalar_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """One element per thread. No vectorization, so alignment is irrelevant.

    This is the baseline: each thread issues a scalar load and a scalar store.
    """
    var i = block_dim.x * block_idx.x + thread_idx.x
    # FILL ME IN (1-2 lines): guard `i < size`, then `output[i] = a[i] * SCALE + BIAS`.


def unaligned_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.

    The data is naturally 16-byte aligned, but `load`/`store` are told only the
    scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
    the access is 16-byte aligned, so it falls back to scalar
    `ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
    `.v4` form. The alignment trap: correct results, but lost bandwidth.
    """

    # Each thread owns one SIMD_WIDTH-wide chunk.
    var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
    # FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load a SIMD_WIDTH-wide vector with `a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))` and store `v * SCALE + BIAS` into `output` with the same `alignment=SCALAR_ALIGN`.


def aligned_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """Same vectorized kernel, but the access alignment is communicated.

    Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
    float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
    load and `st.global.v4.f32` store per chunk. `load` and `store` already
    default to this alignment, so the vectorized form is what you get unless you
    understate it. Identical output to the unaligned kernel — only the codegen
    (and the bandwidth) changes.
    """

    var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
    # FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load with the *correct* alignment via `a.load[width=SIMD_WIDTH](Coord(base))` (`alignment` defaults to `VEC_ALIGN`) and store `v * SCALE + BIAS` into `output` with `alignment=VEC_ALIGN`.


View full file: problems/p35/p35.mojo

Key API

The file defines two alignment constants up front:

comptime VEC_ALIGN = align_of[SIMD[dtype, SIMD_WIDTH]]()   # 16 bytes for float32x4
comptime SCALAR_ALIGN = align_of[dtype]()                  # 4 bytes

The vectorized kernels take TileTensor arguments and call load and store, whose alignment parameter you set explicitly:

# Under-stated: compiler can't prove 16-byte alignment -> scalar codegen
var v = a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))
output.store[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](
    Coord(base), v * SCALE + BIAS
)

# Aligned: 16-byte alignment -> ld.global.nc.v4.f32 / st.global.v4.f32
# `alignment` defaults to `VEC_ALIGN`, so the load needn't pass it
var v = a.load[width=SIMD_WIDTH](Coord(base))
output.store[width=SIMD_WIDTH, alignment=VEC_ALIGN](
    Coord(base), v * SCALE + BIAS
)

On GPU, alignment defaults to align_of[SIMD[dtype, width]](), so a SIMD_WIDTH-wide access is already described as 16-byte aligned unless you understate it. The aligned kernel passes alignment=VEC_ALIGN on the store to keep the pairing visible next to the unaligned one.

Running it

pixi run mojo -I . solutions/p35/p35.mojo --scalar
pixi run mojo -I . solutions/p35/p35.mojo --unaligned
pixi run mojo -I . solutions/p35/p35.mojo --aligned

--scalar runs scalar_kernel. Each command prints <kernel> kernel: passed, confirming all three are correct. The performance difference is the subject of the next section.

Tips
  • The data is already aligned. You are not moving the pointer or padding the buffer. Device allocations come back aligned, and each thread’s base is a multiple of SIMD_WIDTH, so every access is on a 16-byte boundary. The only thing that changes between the unaligned and aligned kernels is the alignment value you pass to load/store.
  • Guard the tail: each vectorized thread handles SIMD_WIDTH elements, so guard with if base + SIMD_WIDTH <= size: to avoid reading past the end.
  • load[w] already means load[w, alignment=align_of[SIMD[dtype, w]]()] on GPU. Understating the alignment is the deliberate mistake this puzzle asks you to make.
  • Don’t expect a wall-clock gap on a tiny input or on a non-NVIDIA GPU. The codegen difference instead shows up in Nsight Compute’s instruction/sector metrics. We’ll cover this in the next section.

Solution

Complete solution

The three kernels are identical in result and differ only in how they touch memory.

def scalar_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """One element per thread. No vectorization, so alignment is irrelevant.

    This is the baseline: each thread issues a scalar load and a scalar store.
    """
    var size = Int(size_dev)
    var i = block_dim.x * block_idx.x + thread_idx.x
    if i < size:
        output[i] = a[i] * SCALE + BIAS


def unaligned_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.

    The data is naturally 16-byte aligned, but `load`/`store` are told only the
    scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
    the access is 16-byte aligned, so it falls back to scalar
    `ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
    `.v4` form. The alignment trap: correct results, but lost bandwidth.
    """
    var size = Int(size_dev)

    # Each thread owns one SIMD_WIDTH-wide chunk.
    var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
    if base + SIMD_WIDTH <= size:
        var v = a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))
        output.store[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](
            Coord(base), v * SCALE + BIAS
        )


def aligned_kernel[
    Engine: TensorEngine,
](
    output: TileTensor[
        mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
    ],
    a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
    size_dev: Int32,
) where (Engine.element_size == 1):
    """Same vectorized kernel, but the access alignment is communicated.

    Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
    float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
    load and `st.global.v4.f32` store per chunk. `load` and `store` already
    default to this alignment, so the vectorized form is what you get unless you
    understate it. Identical output to the unaligned kernel — only the codegen
    (and the bandwidth) changes.
    """
    var size = Int(size_dev)

    var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
    if base + SIMD_WIDTH <= size:
        # `alignment` defaults to `VEC_ALIGN` for a SIMD_WIDTH-wide load.
        var v = a.load[width=SIMD_WIDTH](Coord(base))
        output.store[width=SIMD_WIDTH, alignment=VEC_ALIGN](
            Coord(base), v * SCALE + BIAS
        )


What separates them:

  • The scalar kernel issues one scalar load and one scalar store per thread. Alignment is irrelevant because there is no vector to align.
  • The unaligned kernel asks for a SIMD_WIDTH-wide load but declares only SCALAR_ALIGN (4 bytes). The compiler cannot prove the 16-byte alignment a .v4 instruction requires, so it lowers the vector access to a sequence of scalar ld.global.nc.f32 / st.global.f32 instructions. Correct, but it issues SIMD_WIDTH× the memory instructions of the aligned kernel.
  • The aligned kernel passes VEC_ALIGN (16 bytes). Now the compiler emits a single ld.global.nc.v4.f32 load and st.global.v4.f32 store per chunk, one quarter of the memory instructions for the same work.

The takeaway: state the alignment explicitly. The data was aligned in all three kernels; only the aligned kernel informed the compiler, and only the aligned kernel got the vectorized instruction. The benchmark and profile section demonstrates this in numbers.