





Tier-1 employer, Bangalore location, and mid-level experience make applicant competition high.
Role requires specialized HPC/GPU cluster and networking expertise, so backgrounds transfer poorly across industries.
Explicit 5+ years requirement plus mandatory HPC schedulers, Linux, networking, and automation skills imply high strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead operations and systems administration of large-scale AI/HPC GPU-accelerated clusters, ensuring system health, reliability, and user satisfaction.
Manage day-to-day production AI/HPC cluster activities including system upgrades, incident response, performance optimization, and resource utilization.
Develop and maintain scalable automation, collaborate globally, and provide support for researchers running AI/ML workloads.
Bachelor’s degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
Minimum 5 years experience designing and operating large-scale compute infrastructure.
Experience with AI/HPC job schedulers (e.g. Slurm, Kubernetes, PBS, LSF) and administration of CentOS/RHEL or Ubuntu Linux.
Proficient with cluster configuration management tools (e.g. BCM, Terraform, Ansible), container technologies (Docker, Singularity), Python programming, and bash scripting.
Strong expertise in managing distributed HPC/AI ML clusters including on-premises and cloud environments, focused on performance and reliability.
Experience supporting AI/HPC workflows using MPI and tuning AI/ML workloads for efficiency in GPU-based environments.
Operates with leadership in cross-team collaboration and incident management for high-availability GPU cluster infrastructure.