





Tier-1 employer and Hyderabad metro raise competition, but senior specialization moderates applicant density.
Requires deep SRE and production-support expertise, limiting cross-industry portability.
Explicit 10+ years and specific SRE, observability and incident-response requirements make filters stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own production support and site reliability engineering for high-volume or batch-driven systems, including performance tuning and long-running job monitoring.
Lead incident response efforts, TechLines, root cause analysis, and improve monitoring coverage and alerting aligned to SLOs.
Drive automation to reduce operational toil and maintain runbooks, operational guides, and incident playbooks.
Bachelor’s degree in Computer Science, Engineering, or related field.
Minimum 10 years of experience in production support, SRE, or systems engineering.
Experience participating in on-call rotations with market-hours and after-hours escalations.
Hands-on expertise with observability tools (Splunk, Grafana, Moogsoft, or xMatters) and strong communication skills to lead and influence teams.
Experienced in SRE principles including SLIs, SLOs, error budgets, and burn rate alerts to drive service reliability.
Proficient with incident automation and building self-healing systems to reduce alert noise and operational toil.
Familiarity with infrastructure as code (IaC) tools such as Terraform, Ansible, Salt, or CloudFormation is preferred.