





Niche HPC GPU skillset reduces applicants but metro Mumbai location increases competition.
Very specialized GPU/HPC infrastructure skills limit cross-industry transferability.
Explicit 7–12 years, mandatory GPU/HPC, Slurm/Kubernetes, and certifications create strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and operate large-scale GPU compute clusters (up to 10,000 GPUs) for AI training and inference ensuring high throughput and low latency.
Own day-to-day cluster operations including firmware, driver, and GPU management; execute changes with minimal downtime.
Lead incident response for high-severity GPU, CUDA, and scheduler issues, coordinating cross-functional teams to restore service and perform root cause analysis.
7–12 years of hands-on experience building and operating HPC/AI GPU clusters at scale with deep CUDA/MIG/vGPU expertise.
Bachelor of Technology (BTech) or BE in Computer Science, Electronics & Communication, or equivalent.
Experience with H100/B series GPUs, Slurm and/or Kubernetes device scheduling, and NCCL/UCX performance tuning.
Not explicitly mentioned: notice period requirements. Preferred certifications include NVIDIA Certified Professional (AI Infrastructure) and Linux certifications (RHCE/LPIC).
Operates effectively in complex, large-scale HPC and AI GPU infrastructure environments with multi-cluster orchestration and scheduling expertise.
Demonstrates strong technical leadership during production outages and complex incident resolution involving GPU hardware, CUDA, and cluster schedulers.
Experienced in performance optimization, capacity planning, and enforcing security/compliance for multi-tenant GPU clusters.