Principal Engineer – Scale-Up GPU Networking (HPC / AI)
Hewlett Packard Enterprise (HPE)Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Architect and deliver scale-up GPU networking designs focusing on high-bandwidth, low-latency intra-node communication and GPU-aware networking paths.
Lead integration and optimization with GPU ecosystem technologies such as NVIDIA CUDA, NVLink, AMD ROCm, and related communication stacks including Libfabric, UCX, OpenMPI.
Own debugging, performance tuning, multi-NIC/NUMA optimization, and upstream contributions to open-source HPC/AI networking projects to influence future architecture designs.
Minimum Requirements
10–15+ years of experience in high-performance networking, GPU, or kernel-level software development.
Strong expertise in C/C++, Linux internals, memory management, RDMA, PCIe, IOMMU, ATS, and DMA engines.
Experience with CUDA, ROCm, GPU memory models, P2P, GPUDirect Storage (GDS), and GPUDirect RDMA (GDR).
Hands-on knowledge of MPI, SHMEM, Libfabric, UCX or similar communication stacks.
Ideal Candidate Profile
Proven ability to drive architecture and cross-organization technical decisions and deliver end-to-end solutions in HPC/AI scale-up networking contexts.
Experienced in mentoring senior engineers and influencing multi-team design and upstream open-source contributions.
Deep understanding of HPC system architecture, NUMA tuning, multi-accelerator systems, and NIC architectures such as CXI, RoCE, Infiniband, Slingshot, NVLink Switch.
