Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design and implement automation to reduce manual operational tasks across infrastructure, cloud, and disaster recovery processes.
Own and lead incident management for production environments, including root cause analysis and driving reliability improvements.
Support and improve monitoring, resilience, and operational efficiency for business-critical platforms in Azure and on-premises environments.
Minimum Requirements
5+ years experience in Site Reliability Engineering, DevOps, cloud, infrastructure or systems engineering.
Proficient in scripting and automation with PowerShell, Python, Bash or similar languages.
Experience with CI/CD and DevOps platforms (e.g., Azure DevOps, Jenkins), configuration management (Ansible), and infrastructure as code (Terraform).
Experience supporting production infrastructure in Azure and Windows Server administration and troubleshooting.
Ideal Candidate Profile
Proactive engineer who identifies repetitive tasks, reliability risks, and monitoring gaps to drive automation and operational improvements.
Strong hands-on technical skills in cloud (Azure), automation, and production incident troubleshooting in highly available environments.
Experience with SRE practices including automation of failover/disaster recovery and integrating observability platforms like Grafana and Prometheus.
