





Strong brand and metro location increase applicant density, but niche GPU/Kubernetes skills moderate competition.
Role requires specialized GPU scheduling and Kubernetes controller expertise, limiting cross-industry transferability.
Explicit 5+ years plus mandatory Kubernetes controller, GPU scheduling, and systems engineering skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and own core platform features managing a large multi-vendor GPU fleet for AI training and inference, including scheduling, scaling, multi-tenancy, RBAC, and observability.
Develop control-plane services, autoscaling controllers, inference-serving platforms, and APIs/CLI for ML teams enabling thousands of GPUs usage without human intervention per job.
Design, implement, and ship end-to-end infrastructure capabilities with a product mindset focusing on usability, performance, and self-service for internal ML users.
5+ years experience building infrastructure or platform software with a record of designing and shipping substantial systems or control planes.
Proficiency in Go or Python for building maintainable, production-debuggable systems.
Deep understanding of Kubernetes controller and internals including scheduler, operators, and API mechanisms.
Working knowledge of GPU-specific platform constraints such as MIG, GPU sharing, gang scheduling, topology-aware placement, and multi-tenant hardware resource contention.
Experienced systems engineer with hands-on expertise in Kubernetes operator/controller development at a deep level and GPU platform scheduling.
Strong product mindset focusing on internal developer experience and platform adoption rather than only operational metrics.
Capability to own infrastructure features end-to-end from design through rollout, documentation, and enablement of self-service for AI workloads.