





Popular mid-level SRE role, metro hybrid location, and moderate brand increase applicant density.
Core SRE skills transfer across industries, reducing background sensitivity.
Multiple mandatory SRE, cloud, observability and IaC skill requirements increase filtering rigor.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Accountable for improving reliability, availability, and performance of Travelopia’s production services using monitoring, automation, and operational engineering.
Define and maintain service health via SLIs, SLOs, dashboards, alerts, and operational reporting, supporting incident and problem management including root cause analysis.
Collaborate with cross-functional teams to improve deployment quality, resilience, recovery capabilities, and operational readiness in a 24/7 support environment.
Experience in Site Reliability Engineering and Cloud Operations in AWS, Azure, or hybrid cloud environments.
Hands-on experience with monitoring and observability tools like Grafana, Prometheus, Datadog, Splunk, CloudWatch, or Azure Monitor.
Proficiency with Infrastructure as Code and automation tools such as Terraform, CloudFormation, Ansible, Azure DevOps, and scripting languages (TypeScript, Python, PowerShell, Bash, or Go).
Work Experience Required: Not explicitly mentioned in the JD.
Demonstrated expertise in diagnosing and resolving issues across applications, APIs, infrastructure, networks, operating systems, and cloud services with incident and problem management experience.
Experience working in a complex, 24/7 operational environment involving DevOps, infrastructure, security, and application teams globally.
Familiarity with CI/CD, container technologies (Docker/Kubernetes), DevSecOps practices, vulnerability management, and continuous operational improvement aligned with SFIA service management practices.