





Strong Tier-1 brand and metro location increase interest, but senior niche GPU/HPC skillset limits pool.
Specialized GPU, HPC, and LLM infrastructure expertise reduces cross-industry transferability.
Explicit 8+ years and deep GPU/HPC, CUDA, RDMA, and distributed systems requirements create stringent filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead installation, configuration, and performance optimization of AI infrastructure including GPU servers, storage, high-speed networking, and AI software stacks on HPE platforms.
Perform detailed system-level performance characterization, benchmarking, and optimization of AI/ML workloads, including large language models and distributed training environments.
Collaborate with customers, partners, and internal teams to troubleshoot, optimize, and validate AI hardware/software solutions, and produce technical documentation and best-practice guidance.
8+ years of experience in AI/ML infrastructure, HPC, or related performance engineering roles.
Strong expertise with Linux system administration and command-line across multiple enterprise Linux distributions.
Proficiency with AI frameworks (e.g., PyTorch, JAX, Hugging Face), containerized environments (Docker, Kubernetes), and GPU-accelerated systems including distributed GPU clusters.
Work Experience Required: Minimum 8 years; Notice Period: Not explicitly mentioned in the JD.
Senior or principal-level engineer with demonstrated ability to independently evaluate emerging AI technologies and author technical papers and reference architectures.
Deep domain expertise in AI training/inference performance engineering, including benchmarking large-scale foundation and language models.
Experienced collaborator working with customers and cross-functional teams in complex distributed AI environments, capable of technical leadership and mentoring junior engineers.