





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
High: Tier-1 brand, remote role, common SRE title, mid-level target band, and broad cloud/infrastructure skills.
Medium: Core SRE skills transfer across industries but enterprise/financial-services preference increases fit sensitivity.
Medium: Mandatory hands-on Kubernetes, RabbitMQ, Postgres and Azure skills but no explicit years requirement.
Act as first responder in a 24/7 incident response rotation supporting critical production incidents, triaging and stabilizing infrastructure-level issues under defined SLAs (P1/P2).
Diagnose and fix infrastructure failures by analyzing system logs (Kubernetes, RabbitMQ, PostgreSQL) and escalate only when issues exceed infrastructure scope or are application-level.
For senior levels, additionally responsible for system hardening, scaling, capacity planning, and proposing improvements to reduce incident frequency over time.
Hands-on experience with Kubernetes, RabbitMQ, and PostgreSQL in an enterprise environment.
Strong working knowledge of Azure; familiarity with GCP or AWS is a plus.
Ability to diagnose unfamiliar systems primarily via logs under time pressure with sound judgment on escalation.
Work Experience Required: Not explicitly mentioned in the JD.
Comfortable working in high-pressure, time-sensitive production incident response scenarios as primary triage owner.
Experienced in infrastructure troubleshooting distinct from application-level issues and capable of independently implementing fixes when appropriate.
Has operational maturity to contribute to system improvements and capacity planning beyond immediate incident response.