Principal Engineer – Scale-Up GPU Networking (HPC / AI)
Hewlett Packard Enterprise (HPE)Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Architect and implement GPU-aware networking for high-bandwidth, low-latency intra-node communication with focus on GPU → NIC → GPU data pathways.
Optimize and extend runtime and communication stacks such as Libfabric, UCX, and OpenMPI for GPU-accelerated scale-up workflows including multi-NIC and NUMA performance tuning.
Lead upstream contributions and influence architecture in open-source HPC and AI ecosystem projects, including debugging across driver, runtime, GPU, kernel, and user-space boundaries.
Minimum Requirements
10–15+ years experience in high-performance networking, GPU, or kernel-level software development.
Strong expertise in C/C++, Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, and DMA engines.
Experience with CUDA, ROCm, GPU memory models, GPUDirect Storage (GDS), GPUDirect RDMA (GDR), and communication stacks like MPI, SHMEM, Libfabric, or UCX.
Ability to lead architecture decisions, mentor senior engineers, and own end-to-end delivery. Hybrid work model requiring average 2 days/week in HPE office.
Ideal Candidate Profile
Senior engineer with deep systems-level and HPC/AI networking expertise able to influence cross-functional teams and drive technical strategy.
Experienced in working with complex GPU ecosystems, including CUDA, ROCm, NCCL/RCCL, and advanced NIC architectures (e.g., CXI, RoCE, Infiniband).
Background contributing to or engaging with upstream open-source HPC/AI projects and optimizing multi-accelerator, NUMA-aware system performance.
