





Tier-1 brand, metro location, and mid-level seniority, but niche GPU/Kubernetes expertise reduces density.
High because the role requires deep Kubernetes controller, GPU scheduling, and platform engineering experience specific to AI infra.
High due to explicit 5+ years, deep Kubernetes/GPU platform expertise, and mandatory product ownership expectations.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and own control-plane services and platform layers enabling thousands of GPUs for multi-tenant AI training and inference workloads without manual intervention.
Develop features like scheduling, autoscaling, serving platform, RBAC, observability, CLI, and APIs that ML engineers use daily.
Design and ship Kubernetes-native infrastructure components handling complex GPU constraints and ensure platform scalability, reliability, and self-service capabilities.
5+ years in infrastructure or platform software engineering with delivered systems or control planes.
Strong programming skills in Go or Python with experience in building maintainable, debuggable production systems.
Deep Kubernetes knowledge at the controller/operator level including scheduler internals and API machinery.
Working understanding of GPU platform specifics such as MIG, GPU sharing, gang scheduling, topology-aware placement, and interplay of training and serving workloads.
Experienced in building or contributing to AI/ML serving, training, or inference platforms at scale with autoscaling and rollout capabilities.
Familiar with GPU scheduling systems like Kueue, Volcano, Slurm, or custom schedulers in multi-tenant production environments.
Product-oriented engineer who treats platform as a product, focusing on internal user experience, API design, adoption, and enabling self-service operations.