Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead the architecture and strategy for platform reliability, operational ecosystem, and infrastructure scalability, including setting vision, technical roadmaps, and stakeholder communication.
Design, implement, and maintain Kubernetes platforms, infrastructure automation using Ansible and Terraform, and improve CI/CD pipelines and observability stacks with a focus on reliability and security.
Manage projects using agile practices, track and communicate progress, remove blockers, and drive continuous improvements in processes and architecture.
Minimum Requirements
14+ years experience in building, operating, and optimizing large-scale distributed systems and cloud infrastructure.
Deep hands-on expertise with Kubernetes cluster architecture, networking, scaling, security, and day-2 operations.
Strong experience with Ansible (or equivalent) and Terraform (or equivalent) for infrastructure automation and provisioning.
Proven ownership of CI/CD pipelines, deployment strategies, observability platforms (especially Grafana), and scripting/automation skills (Python, Bash, or equivalent).
Ideal Candidate Profile
Experienced SRE/DevOps leader skilled in formulating multi-year technical roadmaps aligning business and reliability requirements with technical specifications.
Proficient in managing complex infrastructure with a security-first mindset and championing advanced automation, observability, and AI-assisted operations at scale.
Able to balance long-term architectural stability with rapid operational responses, effectively communicating reliability trade-offs to technical and non-technical stakeholders.
