Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead observability strategy and manage end-to-end monitoring stack (metrics, logs, tracing, synthetics, RUM) to improve reliability and SLA compliance.
Drive incident volume reduction by implementing proactive detection, automated response workflows, alert tuning, and continual service improvement.
Manage L1 SRE engineers, oversee incident lifecycle, lead RCA reviews, and integrate observability with ITSM tools and operational processes.
Minimum Requirements
6+ years experience in Observability, SRE, Platform, or Monitoring roles supporting SaaS or enterprise applications.
Hands-on expertise with telemetry tools (e.g., DataDog), distributed tracing, SLIs/SLOs, and automation using languages like Python.
Experience with AWS cloud and ITSM frameworks (Incident/Problem/Change/RCAs).
Must be physically located in Pune, Maharashtra with ability to work hybrid onsite as per role requirements.
Ideal Candidate Profile
Proven capability to architect and mature observability ecosystems across distributed systems with measurable reliability outcomes.
Experience leading cross-functional teams including SRE engineers and collaborating closely with engineering and operations.
Strong focus on automation and metrics-driven incident management to reduce MTTR and alert noise in complex SaaS environments.
