Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, observability, and automation frameworks for significant production domains in a multi-account AWS environment.
Act as incident commander for complex, cross-service incidents and lead root cause analysis with permanent fixes.
Mentor junior SRE engineers, set technical patterns, review designs, and raise overall SRE standards in the India team.
Minimum Requirements
4-6 years hands-on SRE, DevOps, or cloud engineering experience with ownership of production services at scale.
Deep experience with AWS services (EC2, ECS/EKS, RDS, S3, IAM, VPC, Lambda) in multi-account setups.
Strong skills in Terraform module development, GitHub Actions or equivalent CI/CD tooling, and Python for automation.
Bachelor's degree in Computer Science, Engineering, or related field, or equivalent demonstrated skills.
Ideal Candidate Profile
Practitioner skilled in SLO-driven operations and incident command with experience running blameless post-incident reviews.
Engineer with expertise building robust observability using OpenTelemetry or equivalent, and automating self-healing and autoscaling workloads.
Experienced in integrating AI into engineering workflows with measured impact and validation steps, plus prior mentorship or tech lead experience in distributed/global teams.
