





Tier-1 brand, remote role, metro locations, and a mid-level 5-year requirement amplify competition.
Highly specialized HPC cluster, GPU and MPI skills make cross-industry transferability low.
Explicit 5+ years plus mandatory HPC schedulers, Linux, automation, containers, and MPI make shortlisting strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead the administration and reliability improvements of large-scale AI/HPC compute infrastructure including compute, networking, and storage.
Own day-to-day operations, performance optimization, and incident response for production AI/HPC clusters ensuring system health and user satisfaction.
Develop scalable automation and collaborate cross-functionally to support AI/ML research workloads and evolving user needs.
Bachelor’s degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
Minimum 5 years of experience designing and operating large-scale compute infrastructure.
Proficient with AI/HPC job schedulers (e.g., Slurm, Kubernetes, PBS, LSF) and Linux system administration (CentOS/RHEL/Ubuntu).
Experience with cluster configuration management tools (e.g., Ansible, Terraform), container technologies (Docker, Singularity), Python programming, and AI/HPC workflow performance tuning.
Experienced in managing heterogeneous AI/ML clusters both on-premises and cloud environments with strong operational and incident management skills.
Strong background in AI/HPC workloads including MPI, performance analysis, and tuning to meet SLA targets.
Familiarity with NVIDIA GPUs, CUDA, distributed storage (Lustre, GPFS), and high-performance networking (InfiniBand, RDMA) is a plus.