Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, availability, and scalability of caching systems, specifically operating and troubleshooting Redis at scale (Cluster, Replication, HA, Backup/Recovery, Performance Tuning).
Develop automation tools and scripts using Python to enhance operational efficiency and service lifecycle management.
Lead incident response and implement preventive measures to minimize system outages, collaborating with development teams to integrate reliability best practices.
Minimum Requirements
Strong experience with Redis operations and troubleshooting at scale, including cluster and replication management.
Proven background in Site Reliability Engineering (SRE) focused on service reliability, incident management, observability, and automation.
Proficient in Python programming for automation and tooling development.
Work Experience Required: Not explicitly mentioned in the JD; Location: Pune
Ideal Candidate Profile
Experienced in managing large-scale caching infrastructure such as Redis for mission-critical environments, emphasizing performance and availability.
Demonstrates strong engineering skills with an SRE mindset, focusing on automation, incident resolution, and continuous improvement.
Comfortable working in a financial domain and collaborating cross-functionally with development and operations teams to improve platform resilience and scalability.
