





Strong Tier-1 backing, Bangalore metro and mid-level role, but niche GPU/scheduler skills moderate competition.
Requires deep Kubernetes controller and GPU platform expertise, limiting transferability across industries.
Explicit 5+ years plus Kubernetes internals and GPU scheduling requirements enforce strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build the platform software that enables large-scale, multi-tenant GPU infrastructure usage for training and inference workflows, including scheduling, scaling, serving, and self-service capabilities.
Develop and own control-plane services such as schedulers, autoscaling controllers, RBAC/quota systems, and APIs/CLIs used daily by ML engineers.
Collaborate end-to-end from design to rollout, ensuring the platform is reliable, maintainable, and measured by adoption and self-service rates.
5+ years experience building infrastructure or platform software involving control planes and scalable systems.
Strong software engineering skills in Go or Python with experience debugging production systems.
In-depth Kubernetes knowledge at the controller and internals level; experience writing operators or controllers, understanding scheduler and API mechanics.
Working understanding of GPU platform constraints including MIG, GPU sharing, gang scheduling, and topology-aware placement relevant to multi-tenant GPU environments.
Experienced in designing and shipping complex infrastructure products rather than simple scripts or glue code, with ability to own features end-to-end.
Familiar with GPU scheduling frameworks (e.g. Kueue, Volcano, Slurm) and multi-tenancy isolation mechanisms integrated into platforms.
Product-oriented mindset focusing on building user-friendly APIs and abstractions for internal ML engineering teams, prioritizing adoption and self-service metrics.