Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Define and implement operational readiness standards like observability, alerting, SLIs/SLOs, runbooks, and production readiness reviews across multiple consumer-facing domains.
Develop reusable reliability tools and frameworks including instrumentation, synthetic monitoring, load testing, and cross-domain dashboards.
Lead incident analysis, root cause investigations, and coordinate cross-team reliability improvements with an emphasis on increasing team self-sufficiency and operational maturity.
Minimum Requirements
10+ years experience in building and supporting production software systems, including distributed systems.
Proficient in one or more programming languages such as Java, Node.js, TypeScript, or Go for tooling development.
Expertise in observability practices including metrics, distributed tracing, structured logging, SLIs/SLOs, error budgets, and alerting strategies.
Location requirement: Bengaluru, India.
Ideal Candidate Profile
Experienced in leading production incident management and long-term reliability improvement programs at scale.
Capable of coaching and influencing cross-functional engineering teams to elevate operational reliability and maturity without direct authority.
Skilled in partnering with platform engineering and leadership to drive continuous improvements and establish organization-wide reliability standards.
