Deepskilling — Learn. Build. Research.
Deepskilling — Learn. Build. Research.

GPU Programming & CUDA Development

Production-grade CUDA C/C++ — 11 modules with labs: streams, tiled GEMM, warp primitives, cuBLAS/cuFFT/cuDNN, roofline tuning with Nsight, and a benchmarked industry capstone. NVIDIA pathway Step 3.

4.8

(265 students)

8 Weeks · 46 hours

Next Cohort:

No upcoming sessions scheduled yet.

Can't find a suitable date?

What is this course about?

GPU Programming & CUDA Development

Streams · Libraries · Nsight · Production CUDA

Move from CUDA basics to production GPU programming—streams, NVIDIA libraries, GEMM patterns, and Nsight-driven optimization—with rigorous labs.

What you'll walk away with:

  • Asynchronous CUDA patterns with streams and synchronization
  • Library fluency across cuBLAS/cuDNN-style acceleration paths
  • Nsight profiling skills for bottleneck-driven optimization
  • Production habits for robust, measurable GPU code
  • Capstone work ready for advanced multi-GPU and AI tracks
Course Features

Post Graduate Diploma

8 Weeks of Content

Hands-on Projects

Community Support

Lifetime Access

Student Reviews
Daniel Cho
2024-09-14

Streams + cudaMemcpyAsync week finally killed our pipeline bubbles. Nsight Systems timeline snapshots went straight into our infra postmortem template.

Seoul Vision AI Lab

Amelia Frost
2024-09-11

Chapter 8 warp shuffle material levelled up our reductions without fragile shared-memory sizing. Cooperative groups section paid for itself in one sprint.

UK Quant HPC Desk

Ravi Nambiar
2024-09-08

cuBLAS vs custom kernel trade-off lectures mirrored exactly how our CTO reviews perf PRs. cuDNN bridge helped our DL squad trust my numbers.

Bengaluru Edge ML

Laura Benítez
2024-09-05

Nsight Compute section taught me to read memory throughput vs compute bounds like a second language. Promotion doc practically wrote itself.

Madrid CAE GPU Team

Ethan Park
2024-09-02

Chapter 11 capstone forced a defensible benchmark harness—latency percentiles, not vibes. Hiring manager said it was the strongest GPU portfolio they had seen.

Vancouver Inference Startup

Mei Lin
2024-08-29

CUDA-GDB block saved us from a race that only showed on Blackwell-class GPUs. Exactly the production-grade debugging depth I needed.

Shenzhen Robotics OEM

More in NVIDIA

View all NVIDIA courses