Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, popular SRE role, metro location, and mid-seniority raise candidate competition.
Role requires specialized SRE and cloud infrastructure skills but transferable across industries with relevant platform experience.
Many mandatory SRE skills, cloud/IaC/Kubernetes requirements, and 'seasoned' expertise indicate strict technical filtering.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure reliability, availability, and performance of company systems and infrastructure through monitoring, incident response, and troubleshooting.
Design, develop, and maintain automation tools, infrastructure-as-code, and processes to optimize system management, scalability, and performance.
Lead incident response efforts, conduct root cause analysis, and implement preventive measures to improve system resilience and operational efficiency.
Minimum Requirements
Bachelor's degree or equivalent in Computer Science, Information Technology, or related field.
Seasoned hands-on experience in Site Reliability Engineering or related roles including Linux/Unix systems, networking, cloud platforms (AWS, Azure, Google Cloud), and automation tools.
Proficiency in multiple programming/scripting languages such as Python, Java, Go, Ruby, and experience with infrastructure-as-code and containerization technologies (e.g., Terraform, CloudFormation, Docker, Kubernetes).
Work Experience Required: Seasoned hands-on experience in Site Reliability Engineering or related roles (specific years not explicitly mentioned).
Ideal Candidate Profile
Experienced in designing and maintaining highly available, scalable, and fault-tolerant infrastructure architectures.
Skilled in using performance monitoring and tuning tools (Prometheus, Grafana, New Relic) for optimization and troubleshooting.
Strong in incident management including leading response efforts, root cause analysis, and implementing preventive controls within DevOps and Agile environments.
