Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, scalability, and operational excellence of Cloud-based services, including defining and enforcing reliability standards.
Lead 24x7 on-call response and triage, including error management, incident retrospectives, and maintaining on-call rotation health.
Design and implement observability platforms, automate toil reduction, drive chaos engineering programs, and embed SRE best practices across teams.
Minimum Requirements
5+ years of experience in database scaling, performance, reliability, connection pool optimization, and HA design with failover strategies.
3+ years of hands-on Site Reliability Engineering experience including ownership of SLOs and error budget management.
4+ years of experience with Cloud Platforms, specifically GCP CloudSQL.
Bachelor's Degree in Computer Science or equivalent experience.
Ideal Candidate Profile
Experienced in integrating SRE principles into software development lifecycle with strong ownership of incident and error management processes.
Proficient in cloud-native infrastructure and tools including observability stacks, automation, and orchestration technologies like Kubernetes.
Comfortable leading cross-functional distributed teams and mentoring others on reliability engineering and cloud practices.
