Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own architecture and engineering standards for reliability, performance, and scalability across hybrid cloud and on-premise infrastructure.
Define and implement organisation-wide observability strategy including monitoring, alerting, SLIs/SLOs, and error budgets using tools like Datadog, OpenTelemetry.
Lead complex incident response, root cause analysis, and reduce operational load via automation while mentoring SREs and influencing cross-team reliability goals.
Minimum Requirements
10+ years in SRE, DevOps, Platform Engineering, or Cloud Engineering at a senior/principal level.
Expert hands-on experience with large-scale AWS and on-prem environments, Kubernetes production operations (EKS, EKS Anywhere, EKS Hybrid).
Advanced Infrastructure-as-Code skills including Terraform module design; strong knowledge of GitOps (ArgoCD) and automated deployments.
Experience with observability tools (Datadog, Prometheus, OpenTelemetry) and defining SLI/SLO/error budget practices.
Ideal Candidate Profile
Proven track record at senior/architectural level setting technical direction in hybrid cloud and large-scale production environments.
Experience leading cloud migration projects involving Windows IIS/.NET and Linux applications to Kubernetes environments.
Strong technical influencer capable of cross-team collaboration and communicating complex strategies without direct authority, with hands-on mentoring experience.
