





Tier-1 brand and metro location but niche HPC/GPU specialization and seniority limit competition.
Highly specialized AI/HPC GPU, infrastructure and storage skills limit cross-industry transferability.
Explicit 8+ years, 3+ years GPU/HPC requirement, and specific tech/certifications make filtering strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Manage, operate, and optimize HPE's next-generation AI infrastructure platforms focusing on Private Cloud for AI (PCAI) and AI Factory deployments.
Ensure operational stability, lifecycle management, and continuous improvement of AI/HPC GPU-accelerated environments using tools like HPE Ezmeral, NVIDIA AI Enterprise, Kubernetes, and automation frameworks.
Lead incident management, performance monitoring, automation of provisioning and scaling, security compliance, and provide final escalation for operational issues.
Bachelor’s or Master’s degree in Computer Science, IT, or equivalent.
At least 8 years of IT infrastructure administration experience, including 3+ years specifically in AI/HPC or GPU-based environments.
Hands-on expertise with HPE hardware (DL380a, DL325, Cray XD670), NVIDIA GPUs (L40S/H100/H200), InfiniBand NDR networks, virtualization and container platforms (vSphere, RHEL, Kubernetes, Rancher), and automation tools (Ansible, AWX).
Work Location: Hybrid with requirement to work approx. 2 days per week from HPE office.
Experienced infrastructure administrator with deep hands-on knowledge of enterprise-grade AI/HPC/GPU platforms and private cloud environments.
Proven ability to manage complex, large-scale AI platforms involving HPE Ezmeral, NVIDIA AI Enterprise, container orchestration, and automation for lifecycle management.
Skilled in incident root cause analysis, operational documentation, and leading continuous improvement initiatives in global, hybrid IT service settings.