





Strong Tier-1 brand, metro location, and mid-level experience create high competition despite niche specialization.
Highly specialized GPU, RDMA, and Kubernetes platform expertise limits transferability across industries.
Multiple mandatory Kubernetes, GPU, RDMA, and operator skills plus explicit 5+ years make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate Kubernetes Custom Resource Definitions (CRDs), operators, and AI-aware GPU scheduler for large-scale AI training and inference infrastructure.
Architect and maintain multi-NIC RDMA-enabled Kubernetes GPU clusters with topology-aware workload placement and hybrid cloud bursting capabilities.
Collaborate with ML Platform and AI Research teams to optimize infrastructure components such as KubeRay for distributed Ray workloads and GPU pool management across on-prem and cloud environments.
Bachelor's or Master's degree in Computer Science, Engineering, or related field.
5+ years of experience building distributed systems or infrastructure platforms with deep Kubernetes expertise.
Strong programming skills in Go and/or Python; experience with Kubernetes controller frameworks like Kubebuilder or Operator SDK.
Hands-on experience with Kubernetes networking including Gateway API, multi-NIC RDMA clusters, GPU scheduling, and GPU lifecycle management tools.
Experienced in developing and operating Kubernetes operators and custom schedulers targeting AI and GPU workloads at scale.
Skilled in complex Kubernetes networking design and multi-zone GPU cluster orchestration with hybrid cloud integrations.
Proficient in troubleshooting and performance optimization across GPU driver stacks, RDMA networking, and distributed compute infrastructure for ML workloads.