





Senior SRE manager in metro with broad DevOps/Kubernetes requirements yields medium competition.
SRE managerial skills are domain-specific yet reasonably transferable across industries, so medium sensitivity.
Explicit 10+ years plus mandatory AWS, Kubernetes, observability and leadership requirements make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and ensure end-to-end reliability and uptime of enterprise production systems, managing 24x7 incident response operations.
Lead and develop a team (~14 engineers) responsible for incident management, automation, observability, and release stability.
Support reliability, monitoring, and scalability of AI/ML workloads while driving operational efficiency and cloud cost optimization.
Bachelor’s degree in Computer Science, Engineering, or related field.
10+ years of SRE experience including managing production systems and 24x7 operations teams.
Strong hands-on experience with AWS and Kubernetes (EKS preferred).
Experience with monitoring/observability tools like Datadog, CloudWatch, ELK, Prometheus, and knowledge of incident management and RCA processes.
Demonstrated leadership managing and mentoring mid-size SRE or operations teams with focus on building team maturity and ownership.
Proven ability to align reliability initiatives with business goals and drive proactive operational improvements including automation and cost optimization.
Experienced in working with microservices architectures and supporting AI/ML workloads in production environments.