





Senior, metro SRE role with specialized reliability skills yields moderate competition.
Reliability, cloud, and JVM SRE skills are transferable across industries but favor engineering backgrounds.
Explicit 8+ years requirement plus mandatory SRE, multi-region, Datadog, AWS, and JVM skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the reliability posture of the entire Metropolis platform, ensuring 99.9%+ uptime through system design, metrics, and operational practices.
Design and implement failover mechanisms, multi-region deployment strategies, and disaster recovery plans for critical external services and databases.
Lead incident management processes, monitoring systems (Datadog), resilience pattern adoption, and reliability standards across the organization.
Minimum 8 years of engineering experience including software, reliability engineering, SRE, or production operations at scale.
Expertise in multi-region architectures, failover automation, chaos engineering, disaster recovery, and monitoring tools like Datadog.
Strong systems design skills for resilient distributed systems handling failures and external dependency outages.
Proficiency in Java and/or Scala with deep knowledge of JVM performance, concurrency, operational aspects.
Experienced in large-scale distributed systems with hands-on multi-region deployment and disaster recovery ownership.
Skilled in establishing reliability standards and influencing cross-functional teams and architecture decisions.
Strong cloud (AWS) production operations background with a focus on monitoring, incident response, and performance optimization.