





Metro location and visible SRE manager role with broad platform requirements create moderate competition.
High due to preference for regulated financial services, payments, and compliance experience.
Explicit 12+ years and 5+ years management plus mandatory SRE and cloud skills make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead Site Reliability Engineering and Platform Engineering teams to establish and improve reliability, scalability, and operational excellence standards for critical global platforms handling millions of daily transactions.
Own and drive incident management, production operations, automation, and AI-driven operational practices to reduce manual toil and enhance system resilience and observability.
Collaborate with cross-functional and senior leadership teams to influence enterprise-wide reliability strategies, compliance, and service health communications within highly regulated financial services environments.
Bachelor’s degree in Computer Science, Information Systems, IT, or similar preferred.
At least 12 years of experience in Software Engineering, Platform Engineering, DevOps, or SRE; with 5+ years managing SRE, Platform, or Production Support teams.
Strong hands-on expertise in modern SRE practices (SLOs, SLIs, Error Budgets), AWS cloud services (EKS/ECS, EC2, IAM, etc.), Kubernetes, Terraform, CI/CD, and proficiency with scripting languages like Python or Shell.
Work Location: Hybrid model with expectation to work from office minimum three days a week in India.
Experienced leader managing medium to large teams (8-15+ engineers) across multiple workstreams in globally distributed, agile environments supporting business-critical enterprise applications.
Deep domain experience in financial services, payments, fintech or related regulated industries, with knowledge of compliance, audit, and high availability requirements.
Proven capacity to implement and scale automated, self-healing, and AI-assisted reliability and operations engineering practices with measurable impact on operational excellence and incident management.