





Strong employer brand, popular SRE role, and Pune metro location increase competition significantly.
Cloud-native SRE skills transfer across industries but require specific platform and observability experience.
Explicit 8+ years plus mandatory cloud, Kubernetes, observability, IaC, and CI/CD experience creates strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own design and implementation of scalable, reliable systems across hybrid or multi-cloud (primarily AWS/EKS/ECS/Lambda).
Drive improvements in uptime, latency, and service health (SLOs, SLIs, SLAs) and lead major incident management including RCA and postmortems.
Build/manage infrastructure automation (Terraform, Helm, Kubernetes) and own observability stack enhancements (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, etc.).
8+ years of experience in production Kubernetes environments (EKS, ECS) and AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC).
Hands-on experience with infrastructure automation tools: Terraform, Helm, Jenkins, GitHub Actions, GitOps (ArgoCD or Flux).
Proficiency in observability frameworks and tools: metrics, logs, traces, Prometheus, Grafana, Loki, Tempo, OpenTelemetry.
Work Experience Required: 8+ years relevant experience in site reliability engineering and cloud-native operations. Notice period: Not explicitly mentioned in the JD.
Experienced senior-level SRE with deep expertise in AWS cloud and container orchestration managing production workloads.
Proficient with infrastructure as code, CI/CD pipelines, and build automation for safe deployments and operational efficiency.
Comfortable leading incident response processes, mentoring juniors, and collaborating with product and security teams to enforce operational readiness and compliance.