How to Use This Book

Each puzzle maintains a consistent structure to support systematic skill development:

  • Overview: Problem definition and key concepts for each challenge
  • Configuration: Technical setup and memory organization details
  • Code to Complete: Implementation framework in problems/pXX/ with clearly marked sections to fill in
  • Tips: Strategic hints available when needed, without revealing complete solutions
  • Solution: Comprehensive implementation analysis, including performance considerations and conceptual explanations

The puzzles increase in complexity systematically, building new concepts on established foundations. Working through them sequentially is recommended, as advanced puzzles assume familiarity with concepts from earlier challenges.

Running the code

All puzzles integrate with a testing framework that validates implementations against expected results. Each puzzle provides specific execution instructions and solution verification procedures.

Prerequisites

System requirements

Make sure your system meets our system requirements.

Compatible GPU

You’ll need a compatible GPU to run the puzzles. After setup, you can verify your GPU compatibility using the gpu-specs command (see Setting up your environment).

Operating System

Note: Here is some documentation how to setup GPU support in your OS for

Windows WSL2 for Linux with NVIDIA

To setup NVIDIA GPU support on Windows Subsystem for Linux (WSL2) e.g. Ubuntu please follow the NVIDIA CUDA on WSL Guide.

The important information is to install the NVIDIA Windows CUDA Driver for Windows because they fully support WSL2. Once a Windows NVIDIA GPU driver is installed on the system, CUDA becomes available within WSL 2. The CUDA driver installed on Windows host will be stubbed inside the WSL 2 as libcuda.so, therefore users must not install any NVIDIA GPU Linux driver within WSL 2.

Once you have installed the drivers please test the installation

Verify from Windows: Open PowerShell (not WSL)

nvidia-smi

Verify from inside WSL: (first start WSL e.g. via wsl -d Ubuntu)

ls -l /usr/lib/wsl/lib/nvidia-smi
/usr/lib/wsl/lib/nvidia-smi

Check setup from Pixi optionally install missing requirements e.g. for cuda-gdb debugging

pixi run nvidia-smi
pixi run setup-cuda-gdb
pixi run mojo debug --help
pixi run cuda-gdb --version

For WSL you can install VSCode as your Editor

Note: All 35 puzzles work on WSL and Linux with a supported NVIDIA GPU. Some puzzles require a minimum compute capability, and the debugging and profiling puzzles require the corresponding NVIDIA tools. See the GPU support matrix.

Linux native with NVIDIA

Check GPU + Ubuntu version (Supported Ubuntu LTS: 20.04, 22.04, 24.04)

lspci | grep -i nvidia
lsb_release -a

Install NVIDIA driver (mandatory)

sudo ubuntu-drivers devices
sudo ubuntu-drivers autoinstall
sudo reboot

For Linux you can install VSCode as your Editor

  • Install VS Code in Linux via VS Code APT repository

Import Microsoft GPG key

wget -qO- https://packages.microsoft.com/keys/microsoft.asc \
  | gpg --dearmor \
  | sudo tee /usr/share/keyrings/packages.microsoft.gpg > /dev/null

Add VS Code APT repository

echo "deb [arch=amd64 signed-by=/usr/share/keyrings/packages.microsoft.gpg] \
https://packages.microsoft.com/repos/code stable main" \
| sudo tee /etc/apt/sources.list.d/vscode.list

Install VS Code and verify installation

sudo apt update
sudo apt install code
code --version

Note: All 35 puzzles work on Linux with a supported NVIDIA GPU. Some puzzles require a minimum compute capability; puzzle 34 needs SM90 (Hopper) or newer. See the GPU support matrix.

macOS Apple Silicon

For osx-arm64 users, you’ll need:

  • macOS 15.0 or later for optimal compatibility. Run pixi run -e apple check-macos and if it fails you’d need to upgrade.
  • Xcode 16 or later (minimum required). Use xcodebuild -version to check.

If xcrun -sdk macosx metal outputs cannot execute tool 'metal' due to missing Metal toolchain proceed by running

xcodebuild -downloadComponent MetalToolchain

and then xcrun -sdk macosx metal, should give you the no input files error.

Note: Puzzles 1-8, 11-19, 23-28, and 35 work on macOS (24 of the 35). The remainder need NVIDIA-specific tooling, hardware, or PyTorch GPU support. See the GPU support matrix. We’re working to enable more. Please stay tuned!

Programming knowledge

Basic knowledge of:

  • Programming fundamentals (variables, loops, conditionals, functions)
  • Parallel computing concepts (threads, synchronization, race conditions)
  • Basic familiarity with Mojo (language basics parts and intro to pointers section)
  • GPU programming fundamentals is helpful!

No prior GPU programming experience is necessary! We’ll build that knowledge through the puzzles.

Let’s begin our journey into the exciting world of GPU computing with MojoπŸ”₯!

Setting up your environment

  1. Clone the GitHub repository and navigate to the repository:

    # Clone the stable branch, which matches this book
    git clone --branch stable https://github.com/modular/mojo-gpu-puzzles
    cd mojo-gpu-puzzles
    

    The stable branch is what puzzles.modular.com is built from, and it is pinned to the current MAX release. The repository’s default branch, main, tracks nightly builds instead, so cloning it gives you puzzle code that may not compile against the release toolchain these instructions install. If you want to contribute a change, see Development.

  2. Install a package manager to run the MojoπŸ”₯ programs:

pixi

pixi is the recommended option for this project because:

  • Easy access to Modular’s MAX/Mojo packages
  • Handles GPU dependencies
  • Full conda + PyPI ecosystem support

Note: Some puzzles only work with pixi

Install:

curl -fsSL https://pixi.sh/install.sh | sh

Update:

pixi self-update

Option 2: uv

Install:

curl -fsSL https://astral.sh/uv/install.sh | sh

Update:

uv self update

Create a virtual environment:

uv venv && source .venv/bin/activate
  1. Verify setup and run your first puzzle:
# Check your GPU specifications
pixi run gpu-specs

# Run your first puzzle
# This fails waiting for your implementation! follow the content
pixi run p01
# Check your GPU specifications
pixi run -e amd gpu-specs

# Run your first puzzle
# This fails waiting for your implementation! follow the content
pixi run -e amd p01
# Check your GPU specifications
pixi run -e apple gpu-specs

# Run your first puzzle
# This fails waiting for your implementation! follow the content
pixi run -e apple p01
# Install GPU-specific dependencies
uv pip install -e ".[nvidia]"  # For NVIDIA GPUs
# OR
uv pip install -e ".[amd]"     # For AMD GPUs

# Check your GPU specifications
uv run poe gpu-specs

# Run your first puzzle
# This fails waiting for your implementation! follow the content
uv run poe p01

Working with puzzles

Project structure

  • problems/: Where you implement your solutions (this is where you work!)
  • solutions/: Reference solutions for comparison and learning that we use throughout the book

These links point at the stable branch, matching the clone instructions above. Browsing main instead shows nightly code that can differ from what this book describes.

Workflow

  1. Navigate to problems/pXX/ to find the puzzle template
  2. Implement your solution in the provided framework
  3. Test your implementation: pixi run pXX or uv run poe pXX (remember to include your platform with -e platform such as -e amd)
  4. Compare with solutions/pXX/ to learn different approaches

Guarded output buffers

Most puzzles allocate their output inside a larger buffer whose margins hold NaN. When the puzzle finishes it checks those margins, and fails if your kernel wrote into one:

❌ Write detected outside of output buffer

A kernel can write every expected value into output and still write past the end of it, which comparing values alone cannot detect. This message means the indexing ran outside the output region, most often a missing bounds check on the last block, where the thread count exceeds the size of the data.

The margins extend a fixed number of elements past each end, so a write far enough beyond them lands outside the guarded region and goes unreported. A clean check is good evidence, not proof.

Nor does the check reach every puzzle. Puzzles 17 to 22 keep their kernels behind a Python driver and allocate through the MAX graph rather than a device context, so they have no guarded buffer, and neither do puzzles 9, 10 and 30 to 32, which drive external tools instead of asking you to write a kernel. Those puzzles never print the message above, whatever your kernel does.

Essential commands

# The NVIDIA environment is the default. On an AMD or Apple GPU, add
# `-e amd` or `-e apple` to every `pixi run` and `pixi shell` below.

# Run puzzles
pixi run pXX             # NVIDIA (default) same as `pixi run -e nvidia pXX`
pixi run -e amd pXX      # AMD GPU
pixi run -e apple pXX    # Apple GPU

# Test solutions
pixi run tests           # Test all solutions
pixi run tests pXX       # Test specific puzzle

# Run manually
pixi run mojo -I . problems/pXX/pXX.mojo   # Your implementation
pixi run mojo -I . solutions/pXX/pXX.mojo  # Reference solution

# Interactive shell
pixi shell               # Enter environment
mojo -I . problems/p01/p01.mojo              # Direct execution
exit                     # Leave shell

# Development
pixi run format         # Format code
pixi task list          # Available commands
# Note: uv is limited and some chapters require pixi
# Install GPU-specific dependencies:
uv pip install -e ".[nvidia]"  # For NVIDIA GPUs
uv pip install -e ".[amd]"     # For AMD GPUs

# Test solutions
uv run poe tests        # Test all solutions
uv run poe tests pXX    # Test specific puzzle

# Run manually
uv run mojo -I . problems/pXX/pXX.mojo   # Your implementation
uv run mojo -I . solutions/pXX/pXX.mojo  # Reference solution

Puzzles 30, 31 and 32 ship no reference solution, so the solutions/pXX/pXX.mojo commands above have nothing to run for them. Work through those three from the book pages and the profiling output instead.

GPU support matrix

The following table shows GPU platform compatibility for each puzzle. Different puzzles require different GPU features and vendor-specific tools.

PuzzleNVIDIA GPUAMD GPUApple GPUNotes
Part I: GPU Fundamentals
1 - Mapβœ…βœ…βœ…Basic GPU kernels
2 - Zipβœ…βœ…βœ…Basic GPU kernels
3 - Guardsβœ…βœ…βœ…Basic GPU kernels
4 - 2D Mapβœ…βœ…βœ…Basic GPU kernels
5 - Broadcastβœ…βœ…βœ…Basic GPU kernels
6 - Blocksβœ…βœ…βœ…Basic GPU kernels
7 - 2D Blocksβœ…βœ…βœ…Basic GPU kernels
8 - Shared Memoryβœ…βœ…βœ…Basic GPU kernels
Part II: Debugging
9 - GPU Debuggerβœ…βŒβŒNVIDIA-specific debugging tools
10 - Sanitizerβœ…βŒβŒNVIDIA-specific debugging tools
Part III: GPU Algorithms
11 - Poolingβœ…βœ…βœ…Basic GPU kernels
12 - Dot Productβœ…βœ…βœ…Basic GPU kernels
13 - 1D Convolutionβœ…βœ…βœ…Basic GPU kernels
14 - Prefix Sumβœ…βœ…βœ…Basic GPU kernels
15 - Axis Sumβœ…βœ…βœ…Basic GPU kernels
16 - Matrix Multiplicationβœ…βœ…βœ…Advanced memory patterns
Part IV: MAX Graph
17 - Custom Opβœ…βœ…βœ…MAX Graph integration
18 - Softmaxβœ…βœ…βœ…MAX Graph integration
19 - Attentionβœ…βœ…βœ…MAX Graph integration
Part V: PyTorch Integration
20 - 1D Convolution Opβœ…βœ…βŒPyTorch integration
21 - Embedding Opβœ…βœ…βŒPyTorch integration
22 - Fusionβœ…βœ…βŒPyTorch integration
Part VI: Functional Patterns
23 - Functionalβœ…βœ…βœ…Advanced Mojo patterns
Part VII: Warp Programming
24 - Warp Sumβœ…βœ…βœ…Warp-level operations
25 - Warp Communicationβœ…βœ…βœ…Warp-level operations
26 - Advanced Warpβœ…βœ…βœ…Warp-level operations
Part VIII: Block Programming
27 - Block Operationsβœ…βœ…βœ…Block-level patterns
Part IX: Memory Systems
28 - Async Memoryβœ…βœ…βœ…Advanced memory operations
29 - Barriersβœ…βŒβŒAdvanced NVIDIA-only synchronization
Part X: Performance Analysis
30 - Profilingβœ…βŒβŒNVIDIA profiling tools (Nsight)
31 - Occupancyβœ…βŒβŒNVIDIA profiling tools
32 - Bank Conflictsβœ…βŒβŒNVIDIA profiling tools
Part XI: Modern GPU Features
33 - Tensor Coresβœ…βŒβŒNVIDIA Tensor Core specific
34 - Clusterβœ…βŒβŒNVIDIA cluster programming
Part XII: Memory Alignment
35 - Memory Alignmentβœ…βœ…βœ…Aligned vectorized load/store

Legend

  • βœ… Supported: Puzzle works on this platform
  • ❌ Not Supported: Puzzle requires platform-specific features

Platform notes

NVIDIA GPUs (Complete Support)

  • All puzzles (1-35) work on NVIDIA GPUs with CUDA support
  • Requires CUDA toolkit and compatible drivers
  • Best learning experience with access to all features

AMD GPUs (Extensive Support)

  • Most puzzles (1-8, 11-28, 35) work with ROCm support, 27 of the 35
  • Missing only: Debugging tools (9-10), barriers (29), profiling (30-32), Tensor Cores (33), cluster programming (34)
  • Excellent for learning GPU programming including advanced algorithms and memory patterns

Apple GPUs (Substantial Support)

  • Fundamental (1-8, 11-19), advanced (23-28), and memory alignment (35) puzzles are supported, 24 of the 35
  • Missing: Debugging tools (9-10), PyTorch integration (20-22), barriers (29), profiling (30-32), Tensor Cores (33), cluster programming (34)
  • Covers everything up to and including warp and block operations and asynchronous memory patterns

Future Support: We’re actively working to expand tooling and platform support for AMD and Apple GPUs. Missing features like debugging tools, profiling capabilities, and advanced GPU operations are planned for future releases. Check back for updates as we continue to broaden cross-platform compatibility.

GPU Resources

Free cloud GPU platforms

If you don’t have local GPU access, several cloud platforms offer free GPU resources for learning and experimentation:

Google Colab

Google Colab provides free GPU access with some limitations for Mojo GPU programming:

Available GPUs:

  • Tesla T4 (older Turing architecture)
  • Tesla V100 (limited availability)

Limitations for Mojo GPU Puzzles:

  • Older GPU architecture: T4 GPUs may have limited compatibility with advanced Mojo GPU features
  • Session limits: 12-hour maximum runtime, then automatic disconnect
  • Tooling-dependent puzzles: Puzzles 9, 10 and 30-32 drive external NVIDIA tools rather than GPU features. compute-sanitizer (9, 10) comes with the environment. Profiling with ncu (30-32) needs GPU performance-counter access, which shared platforms often restrict, and nsys (30, 31) needs a system CUDA installation. Interactive debugging with cuda-gdb (9) needs one too, and expects a terminal rather than a notebook cell
  • Compute capability limits: T4 is compute capability 7.5, so puzzles requiring 8.0 (16, 28, 29, 33) and 9.0 (34) won’t run
  • Package installation restrictions: May require workarounds for Mojo/MAX installation
  • Performance limitations: Shared infrastructure affects consistent benchmarking

Recommended for: Most of the curriculum. On a compute-capability-7.5 GPU every puzzle runs except 16, 28, 29 and 33 (which need 8.0) and 34 (which needs 9.0); the tooling-dependent puzzles above depend on what your environment exposes.

Kaggle Notebooks

Kaggle offers more generous free GPU access:

Available GPUs:

  • Tesla T4 (30 hours per week free)
  • P100 (limited availability)

Advantages over Colab:

  • More generous time limits: 30 hours per week compared to Colab’s daily session limits
  • Better persistence: Notebooks save automatically
  • Consistent environment: More reliable package installation

Limitations for Mojo GPU Puzzles:

  • Same GPU architecture constraints: T4 compatibility issues with advanced features
  • Tooling-dependent puzzles: Puzzles 9, 10 and 30-32 drive external NVIDIA tools rather than GPU features. compute-sanitizer (9, 10) comes with the environment. Profiling with ncu (30-32) needs GPU performance-counter access, which shared platforms often restrict, and nsys (30, 31) needs a system CUDA installation. Interactive debugging with cuda-gdb (9) needs one too, and expects a terminal rather than a notebook cell
  • Mojo installation complexity: Requires manual setup of Mojo environment
  • Compute capability limits: T4 is compute capability 7.5, so puzzles requiring 8.0 (16, 28, 29, 33) and 9.0 (34) won’t run

Recommended for: Extended learning sessions across most of the curriculum. On a compute-capability-7.5 GPU every puzzle runs except 16, 28, 29 and 33 (which need 8.0) and 34 (which needs 9.0); the tooling-dependent puzzles above depend on what your environment exposes.

Recommendations

  • Complete Learning Path: Use NVIDIA GPU for full curriculum access (all 35 puzzles)
  • Comprehensive Learning: AMD GPUs work well for most content (27 of 35 puzzles)
  • Broad Coverage: Apple GPUs cover fundamental through advanced concepts (24 of 35 puzzles)
  • Free Platform Learning: Google Colab/Kaggle cover most of the curriculum. Their T4 GPUs are compute capability 7.5, which rules out only puzzles 16, 28, 29, 33 (8.0) and 34 (9.0); the debugging and profiling puzzles depend on the tooling each platform exposes
  • Debugging & Profiling: NVIDIA GPU required for debugging tools and performance analysis
  • Modern GPU Features: NVIDIA GPU required for Tensor Cores and cluster programming

Development

This section is for contributing changes to the puzzles themselves, not for solving them. Contributions target the main branch, which tracks nightly builds, rather than the stable branch you cloned to work through the book. For the build, test, and pull request workflow, see Development in the README.

Join the community

Subscribe for Updates Modular Forum Discord

Join our vibrant community to discuss GPU programming, share solutions, and get help!