Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design and implement telemetry, alerting, self-healing, and automation to improve service health and reliability for Azure operational database services.
Participate in on-call rotations to own, investigate, and resolve service issues with broad communication and knowledge sharing.
Own availability, performance, and supportability targets, collaborating with product engineering and management teams and interfacing with customers and support representatives.
Minimum Requirements
6+ years experience in programming (C++, C#, or equivalent), automation/scripting (PowerShell, Python, or similar), and software delivery/production management.
6+ years experience in telemetry-based debugging (KQL preferred) across network, hardware, and distributed systems, with proven ability to fix and optimize code.
Ability to pass Microsoft Cloud background check and meet security screening requirements.
Good verbal and written communication skills.
Ideal Candidate Profile
Experienced in end-to-end system reliability engineering in cloud database or large distributed cloud services.
Strong troubleshooting skills across multiple layers including network, hardware, and distributed services with telemetry-driven analysis.
Able to collaborate effectively with engineering, product teams, and customers to drive improvements and manage service health.
