Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, implement, and operate highly available, resilient, scalable systems following SRE best practices, including defining and managing SLIs, SLOs, and error budgets.
Lead performance engineering efforts such as load, stress, and endurance testing; conduct capacity planning and performance tuning to support business growth.
Lead and participate in incident management, on-call rotations, and continuous operational improvement, including automation of workflows and monitoring/dashboard development.
Minimum Requirements
Proven hands-on experience in Site Reliability Engineering and Performance Engineering for large-scale distributed systems.
Strong proficiency in cloud platforms (AWS, Azure, or Google Cloud) and container orchestration tools like Docker and Kubernetes.
Experience with monitoring, observability, logging tools (e.g., Prometheus, Grafana, Datadog, Splunk, ELK Stack) and CI/CD pipelines (e.g., Jenkins, GitLab CI/CD, Azure DevOps).
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Expertise in SLI/SLO/Error Budget frameworks with ownership of service reliability and customer outcomes in critical, regulated environments (banking, payments, capital markets).
Experience leading performance optimization and incident response in cloud-native, automated environments utilizing Infrastructure as Code and modern automation tooling (Python, Terraform, Ansible).
Ability to collaborate effectively across Development, QA, DevOps, Security, and Product teams to ensure security, compliance, and continuous reliability improvements.
