Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Define and implement Site Reliability Engineering best practices including production support, disaster recovery, automation, and cloud operations.
Lead and mentor a team of SREs, managing incident responses, root cause analysis, and preventative measures to improve system reliability and reduce operational toil.
Establish and manage SLIs, SLOs, SLAs, monitoring/observability strategies, and drive initiatives for fault tolerance and self-healing systems using tools like Gremlin, Chaos Monkey, AWS FIS, Prometheus, Grafana, ELK, and CloudWatch.
Minimum Requirements
Degree in Computer Science (B.Sc., MCA) or Engineering (B.Tech.) or higher.
Minimum 12 years software development experience with at least 5 years leading an SRE team.
Deep expertise in AWS cloud services and advanced knowledge of monitoring and observability tools.
Experience building secure, high-volume transaction web systems in regulated domains like finance or insurance.
Ideal Candidate Profile
Proven leadership experience in SRE with capability to define strategy and operational practices for reliability at scale.
Strong expertise in incident management and implementing automated remediation and fault tolerance systems.
Experience working with senior stakeholders to align reliability goals with business objectives within regulated, mission-critical environments.
