Aligned Load & Store
The exercise
You’ll implement three kernels that compute the same memory-bound map,
out[i] = a[i] * 2 + 1, over a 1M-element float32 buffer. They differ only in
how they touch memory:
scalar_kernel: one element per thread with no vectorization. This is the baseline.unaligned_kernel: vectorized bySIMD_WIDTH, but the access alignment is under-stated (scalar alignment), so the compiler emits scalar loads.aligned_kernel: the same vectorized kernel, with the alignment communicated, so the compiler emits a single vectorized load/store per chunk.
All three produce identical output. The point of the puzzle is everything that happens after you confirm that.
def scalar_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""One element per thread. No vectorization, so alignment is irrelevant.
This is the baseline: each thread issues a scalar load and a scalar store.
"""
var i = block_dim.x * block_idx.x + thread_idx.x
# FILL ME IN (1-2 lines): guard `i < size`, then `output[i] = a[i] * SCALE + BIAS`.
def unaligned_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.
The data is naturally 16-byte aligned, but `load`/`store` are told only the
scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
the access is 16-byte aligned, so it falls back to scalar
`ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
`.v4` form. The alignment trap: correct results, but lost bandwidth.
"""
var a_lt = a.to_layout_tensor()
var out_lt = output.to_layout_tensor()
# Each thread owns one SIMD_WIDTH-wide chunk.
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
# FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load a SIMD_WIDTH-wide vector with `load[width=SIMD_WIDTH, load_alignment=SCALAR_ALIGN](Index(base))` and store `v * SCALE + BIAS` with `store_alignment=SCALAR_ALIGN`.
def aligned_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""Same vectorized kernel, but the access alignment is communicated.
Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
load and `st.global.v4.f32` store per chunk. `aligned_load` is the
convenience wrapper that picks this alignment for you. Identical output to
the unaligned kernel — only the codegen (and the bandwidth) changes.
"""
var a_lt = a.to_layout_tensor()
var out_lt = output.to_layout_tensor()
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
# FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load with the *correct* alignment via `a_lt.aligned_load[width=SIMD_WIDTH](Index(base))` and store `v * SCALE + BIAS` with `store_alignment=VEC_ALIGN`.
View full file: problems/p35/p35.mojo
Key API
The file defines two alignment constants up front:
comptime VEC_ALIGN = align_of[SIMD[dtype, SIMD_WIDTH]]() # 16 bytes for float32x4
comptime SCALAR_ALIGN = align_of[dtype]() # 4 bytes
The vectorized kernels take TileTensor arguments, convert them with
to_layout_tensor(), and then use LayoutTensor’s load and store, whose
alignment you set explicitly:
# Under-stated: compiler can't prove 16-byte alignment -> scalar codegen
var v = a_lt.load[width=SIMD_WIDTH, load_alignment=SCALAR_ALIGN](Index(base))
out_lt.store[width=SIMD_WIDTH, store_alignment=SCALAR_ALIGN](
Index(base), v * SCALE + BIAS
)
# Aligned: 16-byte alignment -> ld.global.nc.v4.f32 / st.global.v4.f32
# `aligned_load[w]` == `load[w, load_alignment=VEC_ALIGN]`
var v = a_lt.aligned_load[width=SIMD_WIDTH](Index(base))
out_lt.store[width=SIMD_WIDTH, store_alignment=VEC_ALIGN](
Index(base), v * SCALE + BIAS
)
aligned_load[w] is the convenience wrapper: it picks
align_of[SIMD[dtype, w]]() for you. aligned_store exists too, but only in a
2D (m, n) form—there is no coordinate-list overload—so this rank-1 kernel
passes store_alignment=VEC_ALIGN explicitly.
Running it
pixi run mojo solutions/p35/p35.mojo --scalar
pixi run mojo solutions/p35/p35.mojo --unaligned
pixi run mojo solutions/p35/p35.mojo --aligned
--scalar runs scalar_kernel. Each command prints <kernel> kernel: passed,
confirming all three are correct. The performance difference is the subject of
the next section.
Tips
- The data is already aligned. You are not moving the pointer or padding the
buffer. Device allocations come back aligned, and each thread’s
baseis a multiple ofSIMD_WIDTH, so every access is on a 16-byte boundary. The only thing that changes between the unaligned and aligned kernels is the alignment value you pass toload/store. - Guard the tail: each vectorized thread handles
SIMD_WIDTHelements, so guard withif base + SIMD_WIDTH <= size:to avoid reading past the end. aligned_load[w]is the same asload[w, load_alignment=align_of[SIMD[dtype, w]]()]. Use whichever reads more clearly.- Don’t expect a wall-clock gap on a tiny input or on a non-NVIDIA GPU. The codegen difference instead shows up in Nsight Compute’s instruction/sector metrics. We’ll cover this in the next section.
Solution
Complete solution
The three kernels are identical in result and differ only in how they touch memory.
def scalar_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""One element per thread. No vectorization, so alignment is irrelevant.
This is the baseline: each thread issues a scalar load and a scalar store.
"""
var size = Int(size_dev)
var i = block_dim.x * block_idx.x + thread_idx.x
if i < size:
output[i] = a[i] * SCALE + BIAS
def unaligned_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.
The data is naturally 16-byte aligned, but `load`/`store` are told only the
scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
the access is 16-byte aligned, so it falls back to scalar
`ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
`.v4` form. The alignment trap: correct results, but lost bandwidth.
"""
var size = Int(size_dev)
var a_lt = a.to_layout_tensor()
var out_lt = output.to_layout_tensor()
# Each thread owns one SIMD_WIDTH-wide chunk.
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
if base + SIMD_WIDTH <= size:
var v = a_lt.load[width=SIMD_WIDTH, load_alignment=SCALAR_ALIGN](
Index(base)
)
out_lt.store[width=SIMD_WIDTH, store_alignment=SCALAR_ALIGN](
Index(base), v * SCALE + BIAS
)
def aligned_kernel(
output: TileTensor[mut=True, dtype, LayoutType, MutAnyOrigin],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin],
size_dev: Int32,
):
"""Same vectorized kernel, but the access alignment is communicated.
Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
load and `st.global.v4.f32` store per chunk. `aligned_load` is the
convenience wrapper that picks this alignment for you. Identical output to
the unaligned kernel — only the codegen (and the bandwidth) changes.
"""
var size = Int(size_dev)
var a_lt = a.to_layout_tensor()
var out_lt = output.to_layout_tensor()
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
if base + SIMD_WIDTH <= size:
# `aligned_load[w]` == `load[w, load_alignment=VEC_ALIGN]`.
var v = a_lt.aligned_load[width=SIMD_WIDTH](Index(base))
out_lt.store[width=SIMD_WIDTH, store_alignment=VEC_ALIGN](
Index(base), v * SCALE + BIAS
)
What separates them:
- The scalar kernel issues one scalar load and one scalar store per thread. Alignment is irrelevant because there is no vector to align.
- The unaligned kernel asks for a
SIMD_WIDTH-wide load but declares onlySCALAR_ALIGN(4 bytes). The compiler cannot prove the 16-byte alignment a.v4instruction requires, so it lowers the vector access to a sequence of scalarld.global.nc.f32/st.global.f32instructions. Correct, but it issuesSIMD_WIDTH× the memory instructions of the aligned kernel. - The aligned kernel passes
VEC_ALIGN(16 bytes). Now the compiler emits a singleld.global.nc.v4.f32load andst.global.v4.f32store per chunk, one quarter of the memory instructions for the same work.
The takeaway: state the alignment explicitly. The data was aligned in all three kernels; only the aligned kernel informed the compiler, and only the aligned kernel got the vectorized instruction. The benchmark and profile section demonstrates this in numbers.