GPU Infrastructure, Distributed AI & Research
Architect GPU fleets for the AI era — Hopper/Blackwell internals, MIG/MPS multi-tenancy, Kubernetes & Slurm, NCCL at scale, enterprise AI platforms, and a publication-grade research capstone. NVIDIA pathway capstone (Step 5 of 5).
4.7
(220 students)
4 Weeks · 48 hours
Next Cohort:
No upcoming sessions scheduled yet.
Can't find a suitable date?
What is this course about?
GPU Infrastructure, Distributed AI & Research
Clusters · K8s · Slurm · NCCL · Research Ops
Operate GPU fleets for AI and research—scheduling, multi-node communication, Kubernetes/Slurm patterns, and governed infrastructure at scale.
What you'll walk away with:
- GPU cluster and scheduler fluency for shared research platforms
- NCCL and distributed training communication basics
- K8s/Slurm patterns for GPU workload operations
- Observability and multi-tenancy habits for GPU fleets
- Research-grade infrastructure judgment for lab and enterprise AI
Course Features
Post Graduate Diploma
4 Weeks of Content
Hands-on Projects
Community Support
Lifetime Access
Student Reviews
Saskia van der Berg
Slurm + GPU observability week mirrored our national lab intake training. Queue fairness section finally gave our PI a sane way to explain preemption policies.
Amsterdam Research ComputingMarcus Osei
MIG vs MPS trade-offs chapter stopped us from over-partitioning A100s—saved a negotiation with our Infra team. Containers + isolation block was worth the course fee alone.
Accra Enterprise AI CloudYuki Taneda
NCCL + topology material mapped directly to our multi-node training postmortems. GPUDirect RDMA explanation filled gaps our network team assumed we already had.
Osaka Distributed AI LabHannah McAllister
Kubernetes GPU operator patterns + autoscaling hooks gave our platform team a shared vocabulary with GKE specialists. Production deployment patterns chapter is bookmarked.
Seattle AI Platform Co.Ibrahim Farah
Compiler/research survey in Chapter 8–9 helped our group decide where to collaborate with universities versus build in-house. Rare systems-course clarity.
Riyadh GPU Research ConsortiumElena Rossi
Chapter 10 research capstone format—evaluation, scalability analysis, publication appendix—is now our lab’s template for annual GPU cluster renewal proposals.
Turin HPC & AI Institute
