Site Reliability Engineer (SRE) EMS Production & Observability
CognizantMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and support EMS production incidents end-to-end including investigation, mitigation, recovery, and root cause analysis.
Design, maintain, and enhance operational dashboards and monitoring systems focused on customer-impacting SLIs/SLOs and reduce alert noise.
Lead proactive reliability and capacity planning for major high-volume events like US Open 2026, including event-specific dashboards, simulations, and readiness reviews.
Minimum Requirements
Experience in Site Reliability Engineering, Production Engineering, DevOps, or Production Support for enterprise-scale applications.
Strong hands-on production troubleshooting, incident management, and performing RCA to drive preventive actions.
Proficiency with cloud infrastructure, containerized environments (Kubernetes), automation (CI/CD), and scripting (Python, Bash, PowerShell).
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Operates with a strong focus on reliability engineering practices including SLIs/SLOs, error budgets, monitoring, and automated remediation.
Experienced in coordinating multi-team incident resolution across infrastructure, application, networking, and platform domains during high-severity outages.
Skilled in capacity planning and proactive event readiness for high-visibility, high-traffic customer-facing platforms requiring operational excellence.
