





Tier-1 employer, mid-level SRE role, metro location and on-call focus increase applicant competition.
Core SRE skills transfer broadly, but Azure/Fabric product specificity raises domain sensitivity to medium.
Mandatory 4+ years, specific SRE tooling, Azure experience and security screening increase shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own incident triage, initial investigation, severity assessment, and response for Microsoft’s Customer Data Integration (CDI) services, ensuring uptime for business-critical data integration products.
Develop and extend AI-driven automation systems including triage agents, routing and classification frameworks, incident summarization, customer communication drafting, and postmortem generation to reduce manual effort.
Measure and improve incident response metrics such as first-time mitigation rate, incident deflection, and time-to-triage while enforcing organizational standards for escalation quality.
4+ years of software engineering experience in site reliability, live site operations, or incident management for cloud services.
Proficiency in one or more programming languages: C#, PowerShell, Python, or KQL/Kusto.
Experience with incident management and observability systems (e.g., ICM, PagerDuty, ServiceNow; Kusto, Geneva, Grafana).
Ability to participate in on-call rotation across time zones in a geographically distributed team.
Must pass Microsoft Cloud background check and security screening.
Experienced in building AI/ML-driven automation or intelligent workflows relevant to incident management and live site operations.
Familiar with live site ecosystem including log traversal, telemetry analysis, and escalation workflows in cloud service environments.
Background with Azure, Power BI, Microsoft Fabric services, and authoring troubleshooting guides for complex incident pattern analysis.