





Metro location and senior role balanced by niche SRE requirements and specialized stack, medium competition.
Role demands deep SRE, JVM, multi-region AWS, and observability expertise, limiting cross-industry transferability.
Explicit 8+ years, senior SRE skills, multi-region and Datadog requirements create strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the overall reliability of the Metropolis platform, ensuring 99.9%+ uptime by establishing practices, metrics, and systems.
Design and implement failover mechanisms for critical external dependencies and architect multi-region deployment strategies with disaster recovery planning.
Lead incident management including workflows, tooling, post-mortems, and reduce mean time to recovery, while driving organization-wide adoption of resilience and observability best practices.
8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale.
Expertise in reliability engineering practices such as multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery.
Strong experience with observability tools like Datadog for monitoring, alerting, tracing, and logging in high-load environments.
Proficiency in Java and/or Scala with deep understanding of JVM performance and concurrency; production experience operating reliable systems on AWS including multi-region deployments.
Experienced in designing and operating resilient distributed systems with cloud platform expertise and strong systems thinking.
Has led or significantly contributed to reliability engineering or SRE functions, with skills in incident response processes and reducing MTTR in production.
Strong technical communicator who influences architecture decisions and sets company-wide standards for reliability and observability.