Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Respond to and resolve P1-P4 incidents during day hours in collaboration with NOC teams across North America and India.
Identify root causes of recurring issues and implement permanent fixes to improve platform reliability.
Develop automation and tooling on AWS and Azure to reduce manual operations and technical debt, contributing to projects like backup architecture redesign.
Minimum Requirements
Bachelor's degree in computer science or related field.
At least 3+ years of experience in a technical SaaS environment.
Hands-on experience with Kubernetes, AWS, Azure, Infrastructure as Code (Terraform, Ansible or similar), and CI/CD pipelines (preferably GitLab).
Experience with incident management, on-call rotations, escalations, monitoring and observability tools (e.g., Datadog).
Ideal Candidate Profile
Experienced in managing SaaS platform reliability and operational incident response.
Proficient in cloud infrastructure automation and maintaining high availability in distributed environments.
Capable of driving long-term reliability improvements by reducing manual effort through automation and root cause analysis.
