





Metro mid-level role with brand visibility, though niche LLM/GPU infra reduces broad applicant competition.
Highly specialized GPU/LLM serving and SRE expertise limits transferability across industries.
Mandatory 4+ years plus required SRE, Go, and Kubernetes experience narrows candidate pool.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build scalable, multi-tenant AI inference services focused on throughput, GPU utilization, and fault tolerance.
Develop and operate high-scale, reliable distributed systems and enhance platform resiliency using observability, automation, and operational tooling.
Collaborate with platform and product teams to deliver production-grade APIs and lead efforts in incident management and service health improvement.
Minimum 4 years experience building and operating multi-tenant platforms or distributed backend systems.
At least 1 year hands-on production experience with Go / Golang.
At least 1 year experience with Kubernetes.
Role location: Hybrid based in Bengaluru, India.
Experienced in operating high-scale distributed services with a strong focus on SRE principles including observability, incident management, and reliability engineering.
Familiar with cloud-native architectures, microservices, and debugging production performance and scalability issues.
Understanding of AI/ML inference serving architectures and metrics like Time To First Token and GPU utilization is a strong plus.