





Tier-1 brand plus mid-level experience and metro location, but niche RDMA/InfiniBand skills reduce candidate pool.
Highly specialized RDMA, InfiniBand, HPC networking and cluster automation skills limit cross-industry transferability.
Multiple mandatory domain-specific skills and explicit 5+ years plus tooling and benchmarking requirements increase filter rigidity.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design and implementation of large-scale AI/HPC networking projects involving Ethernet and InfiniBand technologies.
Own support for operational reliability and performance of large-scale AI clusters, including real-time monitoring, logging, and alerting.
Collaborate with customers, partners, and internal teams throughout service lifecycle from design to deployment and ongoing maintenance.
Bachelor's, Master's, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
5+ years professional experience in networking fundamentals, specifically Ethernet or InfiniBand.
Hands-on experience with network switch/router platforms (e.g., Cumulus Linux, SONiC, IOS, JunosOS, EOS).
Proficient in end-to-end InfiniBand/Ethernet cluster deployment, RDMA technologies, automated network provisioning (Ansible, Salt, Python), and troubleshooting network issues.
Experienced in performance optimization and benchmarking of RDMA and Ethernet networks in AI/HPC or distributed storage environments.
Capable of independently diagnosing complex network anomalies and implementing advanced RDMA network optimization strategies.
Familiar with CI/CD pipeline development for network operations and understanding of HPC cluster management and job scheduling systems (e.g., Slurm, PBS).