GPU Programming & CUDA Development
Production-grade CUDA C/C++ — 11 modules with labs: streams, tiled GEMM, warp primitives, cuBLAS/cuFFT/cuDNN, roofline tuning with Nsight, and a benchmarked industry capstone. NVIDIA pathway Step 3.
4.8
(265 students)
8 Weeks · 46 hours
Next Cohort:
No upcoming sessions scheduled yet.
Can't find a suitable date?
What is this course about?
GPU Programming & CUDA Development
Streams · Libraries · Nsight · Production CUDA
Move from CUDA basics to production GPU programming—streams, NVIDIA libraries, GEMM patterns, and Nsight-driven optimization—with rigorous labs.
What you'll walk away with:
- Asynchronous CUDA patterns with streams and synchronization
- Library fluency across cuBLAS/cuDNN-style acceleration paths
- Nsight profiling skills for bottleneck-driven optimization
- Production habits for robust, measurable GPU code
- Capstone work ready for advanced multi-GPU and AI tracks
Course Features
Post Graduate Diploma
8 Weeks of Content
Hands-on Projects
Community Support
Lifetime Access
Student Reviews
Daniel Cho
Streams + cudaMemcpyAsync week finally killed our pipeline bubbles. Nsight Systems timeline snapshots went straight into our infra postmortem template.
Seoul Vision AI LabAmelia Frost
Chapter 8 warp shuffle material levelled up our reductions without fragile shared-memory sizing. Cooperative groups section paid for itself in one sprint.
UK Quant HPC DeskRavi Nambiar
cuBLAS vs custom kernel trade-off lectures mirrored exactly how our CTO reviews perf PRs. cuDNN bridge helped our DL squad trust my numbers.
Bengaluru Edge MLLaura Benítez
Nsight Compute section taught me to read memory throughput vs compute bounds like a second language. Promotion doc practically wrote itself.
Madrid CAE GPU TeamEthan Park
Chapter 11 capstone forced a defensible benchmark harness—latency percentiles, not vibes. Hiring manager said it was the strongest GPU portfolio they had seen.
Vancouver Inference StartupMei Lin
CUDA-GDB block saved us from a race that only showed on Blackwell-class GPUs. Exactly the production-grade debugging depth I needed.
Shenzhen Robotics OEM
