The CUDA Programming Guide is thick and occasionally contradictory

I've read it cover to cover three times. It's 900 pages of NVIDIA's official documentation for CUDA C++, and it will tell you everything about the programming model but occasionally miss the part where your kernel actually crashes. The guide is the canonical reference. It is not a tutorial in the beginner sense. Treat it like an API manual that double as a language specification. It is the documentation set maintained by NVIDIA that describes the CUDA programming model, language extensions, device APIs, runtime versus driver API, optimization techniques, and architecture-specific details. The current version follows CUDA 12.x and covers compute capabilities 3.5 through 9.0+. You can download it from the NVIDIA developer site as a PDF or browse it online. There is no single "latest" link that lasts forever because NVIDIA restructures the URL scheme every major release. Search for "CUDA Toolkit Documentation" and pick the version that matches your installed toolkit. Mismatching the guide version with your toolkit is how you end up reading about cuBLAS 11.x APIs while running cuBLAS 12.x. The document is divided into parts. Part I covers the programming model and architecture overview. Part II is the CUDA C++ language extension. Part III and IV handle the runtime and driver APIs. Part V goes into device side execution, shared memory, and vectorized memory accesses. Part VI covers optimization chapters, profiling tools, and debugging. The later sections get into CUDA graphs, streams, cooperative groups, and the various math libraries. Read them in that order if you have time. I usually skip ahead to the optimization chapters first because they contain the things that will save you from writing garbage code.

How to read it without losing your mind

Most people open the guide and start at chapter one. That is a mistake. The language specification section is dense and full of footnotes about what happens when you break strict aliasing rules or use __restrict__ incorrectly. Instead, go straight to the execution model chapter. It explains grid, block, thread hierarchy, warp scheduling, and occupancy. Once you understand that threads execute in groups of 32 and that divergence costs you cycles, everything else starts making sense. I keep the guide open on a second monitor while coding. When I hit a compiler error about __launch_bounds__ or a linking error about unresolved external symbols from the driver API, I search the PDF rather than posting to a forum. The guide has terrible search indexing because NVIDIA exports it as a scanned-style PDF with poor text layer handling. The web version is better for keyword searches. Use the online reader for quick lookups and the PDF for deep reading sessions.

Practical stuff the guide won't emphasize enough

Memory hierarchy is the thing that breaks most CUDA programs. The guide describes L1, L2, shared memory, constant memory, texture memory, and global memory in separate sections. It does not always make clear how they interact under real workloads. Shared memory bank conflicts are still a real problem on sm_80 and newer architectures even though NVIDIA claims to have reduced their impact. I ran into this recently with a matrix transpose kernel. The naive implementation used one thread per element writing to shared memory with a stride that hit the same bank repeatedly. Performance was 12 GB/s on an A100. Adding a one-element padding to each row eliminated the bank conflicts and pushed it to 890 GB/s. The guide mentions padding as a technique but buries it in an optimization chapter footnote. You will not see the actual performance numbers there. Another thing beginners miss is the difference between kernel launch configuration and actual occupancy. You can launch a kernel with 1024 threads per block and think you are maxing out the GPU. You are not. The scheduler limits concurrent blocks based on register usage, shared memory allocation, and maximum threads per block. If your kernel uses 64 registers per thread, you might only be able to run two warps per SM. The occupancy calculator on the NVIDIA dev site helps, but it assumes ideal conditions. Real world occupancy is often lower because of driver overhead, other processes on the GPU, and dynamic shared memory allocations that the compiler cannot resolve at compile time. Stream concurrency is another area where the guide is accurate but incomplete. You can launch multiple streams simultaneously and expect parallel kernel execution. That works on data-parallel workloads. It breaks when kernels have implicit dependencies through global memory writes that subsequent kernels read. The guide describes cudaStreamWaitEvent and cudaStreamMemOp for managing dependencies but does not warn you about how expensive event synchronization is. A single cudaEventSynchronize call can stall the CPU for hundreds of microseconds if the GPU is busy. Use asynchronous event queries with cudaEventQuery instead of blocking sync calls inside performance-critical loops.

Debugging with the guide as your reference

cuda-gdb and Nsight Compute are the primary debugging tools. The guide dedicates a section to each. cuda-gdb supports thread-level debugging but is slow because it serializes execution. Nsight Compute gives you cycle-accurate metrics but requires a profile run before you can inspect registers. I usually write a small harness that runs the kernel 1000 times with different inputs, captures a profile, then inspects the wave occupancy and memory throughput. If the numbers look wrong, I add __syncthreads() at strategic points to confirm whether a race condition is causing the issue. This takes about 20 minutes instead of spending hours trying to reason through undefined behavior. The guide also covers CUDA_MemCheck, which is essentially a leak sanitizer for GPU memory. It is slow. Something like 10x to 50x slowdown depending on allocation patterns. Use it during development, not in production. If you are shipping code, run it once on a test suite and trust the results.

What the guide leaves out

NVIDIA does not document every edge case. The guide will tell you that variable-length argument lists in device functions are unsupported. It will not tell you that using printf inside a kernel on certain architectures can silently corrupt shared memory if you exceed the printf buffer size. I learned that the hard way on a V100. The buffer overflow did not crash the program. It just produced wrong results that took three days to trace back to a printf call with a format string longer than 256 characters. The workaround was switching to a custom logging kernel that wrote to a ring buffer in shared memory and flushed it to global memory at the end of each kernel launch. Another gap is the interaction between CUDA and newer C++ standards. The guide covers C++11 and C++14 fairly well. C++17 and C++20 features have uneven support depending on your compute capability and compiler version. std::optional and std::variant work on sm_80+ with nvcc 12.0+, but constexpr if statements can cause unexpected compilation failures in template-heavy code because the CUDA compiler handles SFINAE differently than nvcc's host compiler. If you are writing template libraries that target both host and device, keep the device-side code in a separate namespace and compile it with -std=c++14 even if your host code uses C++20.

Where to get it

The guide ships with the CUDA Toolkit installation. If you installed the toolkit, you already have it. Check $CUDA_HOME/doc/html/index.html for the online version or look for the PDF in $CUDA_HOME/doc/. You can also download standalone versions from the NVIDIA CUDA website. Pick the version that matches your toolkit exactly. Using a guide from CUDA 11.8 with toolkit 12.2 will work for most things but will confuse you on changes to cuBLAS handle management and the new cooperative groups API additions. There is no free third-party book that covers CUDA as comprehensively as the official guide. "Professional CUDA C Programming" by Cheng et al. is useful but outdated after CUDA 10. "CUDA by Example" is too basic for production work. The official guide plus Nsight Compute documentation plus the cuBLAS and cuDNN API references is the complete set most engineers end up relying on.

Final practical note

Don't treat the guide as something you read sequentially from front to back. It is a reference. Open it when you need something. The most valuable sections are the execution model, shared memory and cache utilization, and the optimization chapters. Those three parts alone will prevent most performance issues. The rest is lookup material for APIs you are not sure about. Keep a bookmark on the occupancy calculator page. Use it before every kernel launch change. It takes five seconds and will save you from shipping code that runs at half the speed you expected.