Aligned Load & Store
The exercise
You’ll implement three kernels that compute the same memory-bound map,
out[i] = a[i] * 2 + 1, over a 1M-element float32 buffer. They differ only in
how they touch memory:
scalar_kernel: one element per thread with no vectorization. This is the baseline.unaligned_kernel: vectorized bySIMD_WIDTH, but the access alignment is under-stated (scalar alignment), so the compiler emits scalar loads.aligned_kernel: the same vectorized kernel, with the alignment communicated, so the compiler emits a single vectorized load/store per chunk.
All three produce identical output. The point of the puzzle is everything that happens after you confirm that.
def scalar_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""One element per thread. No vectorization, so alignment is irrelevant.
This is the baseline: each thread issues a scalar load and a scalar store.
"""
var i = block_dim.x * block_idx.x + thread_idx.x
# FILL ME IN (1-2 lines): guard `i < size`, then `output[i] = a[i] * SCALE + BIAS`.
def unaligned_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.
The data is naturally 16-byte aligned, but `load`/`store` are told only the
scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
the access is 16-byte aligned, so it falls back to scalar
`ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
`.v4` form. The alignment trap: correct results, but lost bandwidth.
"""
# Each thread owns one SIMD_WIDTH-wide chunk.
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
# FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load a SIMD_WIDTH-wide vector with `a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))` and store `v * SCALE + BIAS` into `output` with the same `alignment=SCALAR_ALIGN`.
def aligned_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""Same vectorized kernel, but the access alignment is communicated.
Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
load and `st.global.v4.f32` store per chunk. `load` and `store` already
default to this alignment, so the vectorized form is what you get unless you
understate it. Identical output to the unaligned kernel — only the codegen
(and the bandwidth) changes.
"""
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
# FILL ME IN (~4 lines): guard `base + SIMD_WIDTH <= size`, then load with the *correct* alignment via `a.load[width=SIMD_WIDTH](Coord(base))` (`alignment` defaults to `VEC_ALIGN`) and store `v * SCALE + BIAS` into `output` with `alignment=VEC_ALIGN`.
View full file: problems/p35/p35.mojo
Key API
The file defines two alignment constants up front:
comptime VEC_ALIGN = align_of[SIMD[dtype, SIMD_WIDTH]]() # 16 bytes for float32x4
comptime SCALAR_ALIGN = align_of[dtype]() # 4 bytes
The vectorized kernels take TileTensor arguments and call load and store,
whose alignment parameter you set explicitly:
# Under-stated: compiler can't prove 16-byte alignment -> scalar codegen
var v = a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))
output.store[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](
Coord(base), v * SCALE + BIAS
)
# Aligned: 16-byte alignment -> ld.global.nc.v4.f32 / st.global.v4.f32
# `alignment` defaults to `VEC_ALIGN`, so the load needn't pass it
var v = a.load[width=SIMD_WIDTH](Coord(base))
output.store[width=SIMD_WIDTH, alignment=VEC_ALIGN](
Coord(base), v * SCALE + BIAS
)
On GPU, alignment defaults to align_of[SIMD[dtype, width]](), so a
SIMD_WIDTH-wide access is already described as 16-byte aligned unless you
understate it. The aligned kernel passes alignment=VEC_ALIGN on the store to
keep the pairing visible next to the unaligned one.
Running it
pixi run mojo -I . solutions/p35/p35.mojo --scalar
pixi run mojo -I . solutions/p35/p35.mojo --unaligned
pixi run mojo -I . solutions/p35/p35.mojo --aligned
--scalar runs scalar_kernel. Each command prints <kernel> kernel: passed,
confirming all three are correct. The performance difference is the subject of
the next section.
Tips
- The data is already aligned. You are not moving the pointer or padding the
buffer. Device allocations come back aligned, and each thread’s
baseis a multiple ofSIMD_WIDTH, so every access is on a 16-byte boundary. The only thing that changes between the unaligned and aligned kernels is the alignment value you pass toload/store. - Guard the tail: each vectorized thread handles
SIMD_WIDTHelements, so guard withif base + SIMD_WIDTH <= size:to avoid reading past the end. load[w]already meansload[w, alignment=align_of[SIMD[dtype, w]]()]on GPU. Understating the alignment is the deliberate mistake this puzzle asks you to make.- Don’t expect a wall-clock gap on a tiny input or on a non-NVIDIA GPU. The codegen difference instead shows up in Nsight Compute’s instruction/sector metrics. We’ll cover this in the next section.
Solution
Complete solution
The three kernels are identical in result and differ only in how they touch memory.
def scalar_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""One element per thread. No vectorization, so alignment is irrelevant.
This is the baseline: each thread issues a scalar load and a scalar store.
"""
var size = Int(size_dev)
var i = block_dim.x * block_idx.x + thread_idx.x
if i < size:
output[i] = a[i] * SCALE + BIAS
def unaligned_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""Vectorized by SIMD_WIDTH, but the access alignment is *under-stated*.
The data is naturally 16-byte aligned, but `load`/`store` are told only the
scalar alignment (`align_of[dtype]()` == 4 bytes). The compiler cannot prove
the access is 16-byte aligned, so it falls back to scalar
`ld.global.nc.f32` / `st.global.f32` instructions instead of the vectorized
`.v4` form. The alignment trap: correct results, but lost bandwidth.
"""
var size = Int(size_dev)
# Each thread owns one SIMD_WIDTH-wide chunk.
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
if base + SIMD_WIDTH <= size:
var v = a.load[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](Coord(base))
output.store[width=SIMD_WIDTH, alignment=SCALAR_ALIGN](
Coord(base), v * SCALE + BIAS
)
def aligned_kernel[
Engine: TensorEngine,
](
output: TileTensor[
mut=True, dtype, LayoutType, MutAnyOrigin, Engine=Engine
],
a: TileTensor[mut=False, dtype, LayoutType, ImmutAnyOrigin, Engine=Engine],
size_dev: Int32,
) where (Engine.element_size == 1):
"""Same vectorized kernel, but the access alignment is communicated.
Passing `VEC_ALIGN` (`align_of[SIMD[dtype, SIMD_WIDTH]]()` == 16 bytes for
float32x4) lets the compiler emit a single vectorized `ld.global.nc.v4.f32`
load and `st.global.v4.f32` store per chunk. `load` and `store` already
default to this alignment, so the vectorized form is what you get unless you
understate it. Identical output to the unaligned kernel — only the codegen
(and the bandwidth) changes.
"""
var size = Int(size_dev)
var base = (block_dim.x * block_idx.x + thread_idx.x) * SIMD_WIDTH
if base + SIMD_WIDTH <= size:
# `alignment` defaults to `VEC_ALIGN` for a SIMD_WIDTH-wide load.
var v = a.load[width=SIMD_WIDTH](Coord(base))
output.store[width=SIMD_WIDTH, alignment=VEC_ALIGN](
Coord(base), v * SCALE + BIAS
)
What separates them:
- The scalar kernel issues one scalar load and one scalar store per thread. Alignment is irrelevant because there is no vector to align.
- The unaligned kernel asks for a
SIMD_WIDTH-wide load but declares onlySCALAR_ALIGN(4 bytes). The compiler cannot prove the 16-byte alignment a.v4instruction requires, so it lowers the vector access to a sequence of scalarld.global.nc.f32/st.global.f32instructions. Correct, but it issuesSIMD_WIDTH× the memory instructions of the aligned kernel. - The aligned kernel passes
VEC_ALIGN(16 bytes). Now the compiler emits a singleld.global.nc.v4.f32load andst.global.v4.f32store per chunk, one quarter of the memory instructions for the same work.
The takeaway: state the alignment explicitly. The data was aligned in all three kernels; only the aligned kernel informed the compiler, and only the aligned kernel got the vectorized instruction. The benchmark and profile section demonstrates this in numbers.