





Medium: known brand and Bangalore metro increase applicants, but specialized SRE/GPU inference skills reduce candidate pool.
High: role requires specialized SRE, GPU inference, and cloud-native platform skills with limited cross-industry transferability.
Medium: explicit 2+ years requirement plus mandatory Go and Kubernetes skills enforce moderate filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build scalable multi-tenant services for AI inference workloads focusing on throughput, GPU utilization, and fault tolerance.
Develop and operate high-scale distributed systems with reliability, availability, and performance objectives.
Lead on-call rotations, improve platform resiliency via observability, automation, and operational tooling, and contribute to architecture decisions around traffic management and scalability.
2+ years experience building and operating multi-tenant platforms or distributed backend systems.
1+ years hands-on experience with Go/Golang in production systems.
1+ years experience with Kubernetes.
Location: Hybrid role based in Bengaluru, India.
Strong expertise in operating high-scale distributed services with a deep understanding of SRE principles (observability, incident management, reliability engineering).
Experience debugging performance, scalability, and reliability issues in production, with proficiency in infrastructure/inference metrics like TTFT and GPU utilization.
Familiarity with cloud-native architectures, microservices, and distributed systems fundamentals, ideally with exposure to AI/ML inference serving architectures or related optimization systems.