Principal Engineer – Scale-Up GPU Networking (HPC / AI)
Hewlett Packard Enterprise (HPE)Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design and implement GPU-aware scale-up networking for high-bandwidth, low-latency intra-node communication including GPU→NIC→GPU data paths.
Develop and optimize GPU ecosystem integration with CUDA, ROCm stacks and enhance runtime communication layers like Libfabric, UCX, OpenMPI for GPU-accelerated workflows.
Lead performance optimization on multi-NIC/NUMA GPU setups, drive upstream open-source contributions, and own complex debugging across driver, runtime, kernel, and user-space boundaries.
Minimum Requirements
10–15+ years experience in high-performance networking, GPU, or kernel-level software engineering.
Proficiency in C/C++, Linux internals, RDMA, PCIe, IOMMU, DMA engine programming, and GPU memory models including CUDA and ROCm ecosystems.
Hands-on experience with MPI, SHMEM, Libfabric, UCX, or equivalent communication stacks and demonstrated ability to drive architecture decisions and mentor engineers.
Hybrid work model requiring average 2 days per week onsite at HPE office.
Ideal Candidate Profile
Experienced senior engineer or architect with deep HPC/AI scale-up networking expertise focused on GPU and multi-accelerator system optimizations.
Proven leader influencing cross-team designs and open-source projects within HPC/AI communication and GPU runtime stacks.
Strong background in complex system performance tuning, debugging across hardware-software layers, and upstream open-source community engagement.
