Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead a high-performing Site Reliability Engineering (SRE) pod, owning technical design, production deployments, and resolving complex technical issues.
Architect and implement systemic improvements in AWS and Kubernetes environments focusing on automation and reliability frameworks.
Steer engineering excellence through hands-on coding, rigorous reviews, incident management, and mentoring to enhance team capabilities.
Minimum Requirements
10+ years in SRE, DevOps, or Software Engineering with at least 3 years in a Staff-level, Technical Lead, Manager, or Sr. Manager role.
Expert-level proficiency in AWS ecosystem and Kubernetes (EKS) orchestration managing large-scale distributed cloud systems.
Advanced skills in Infrastructure as Code tools like Terraform/CloudFormation and GitOps practices.
Strong coding ability in Python, Go, or Java and experience with observability tools such as Datadog, Prometheus, or Splunk.
Ideal Candidate Profile
Experienced technical leader capable of bridging long-term reliability strategy with day-to-day execution in a complex cloud environment.
Demonstrates operational discipline with on-call management and a safety-first reliability mindset prioritizing system uptime and MTTR reduction.
Skilled in automation and innovation, including integrating AI/ML for predictive monitoring and proactive incident prevention.
