





Popular mid-level SRE role in Bangalore with common cloud skills and attractive startup brand increases competition.
Core cloud, SRE, and IaC skills are broadly transferable across industries.
Explicit 3–5 years requirement plus mandatory cloud/SRE tools (Kubernetes, Terraform) and on-call experience increases strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own reliability, latency, and scalability of cloud-based AI inference service across multiple global regions, including Asia, Europe, and Latin America.
Lead incident management and drive preventative improvements to reduce downtime and on-call disruptions.
Implement and maintain infrastructure as code, CI/CD pipelines, and monitoring systems to optimize performance, cost, and uptime.
Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
3-5+ years experience as Site Reliability Engineer, DevOps, or similar role supporting large-scale public cloud services (AWS, GCP, Azure).
Proficient in Python, Go, or Java programming/scripting; experience with Docker and Kubernetes.
Experience with Infrastructure as Code tools (Terraform, CloudFormation) and monitoring tools like Prometheus and Grafana.
Experience managing hybrid environments bridging public cloud and on-prem infrastructure at scale.
Prior involvement with production ML/AI inference services and GPU-accelerated workload optimization.
Strong operational focus on designing low-latency, highly available distributed systems with automation to minimize on-call toil.