Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, Bangalore location, mid-level SRE title, and broad requirements heighten competition.
Core SRE skills are transferable, but payments and AI/fintech domain specifics increase fit sensitivity.
Explicit years, mandatory SRE experience, coding, cloud, and ML/AIOps requirements create strict screening.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, uptime, and performance of Razorpay’s core payment infrastructure processing $180+ billion annually, focusing on capacity planning, failure-mode analysis, and architectural improvements.
Develop and operate automation and self-healing infrastructure to reduce manual toil and improve system resilience, leveraging AI/ML for observability, anomaly detection, predictive alerting, and AI-assisted incident management.
Lead cross-team stakeholder collaboration, mentor engineers on SRE best practices, and drive adoption of AI-assisted tools to improve incident response efficiency and overall platform reliability.
Minimum Requirements
Bachelor’s degree in Computer Science or related technical field, or equivalent practical experience.
Minimum 5 years experience in software/systems/site reliability engineering; at least 3 years specifically in SRE with scalable, reliable distributed systems.
Proficiency in at least one programming language such as Go, Python, or Java, with experience writing production-quality code and automation.
Experience with AI-assisted development and operations tooling including LLM-based copilots for code generation, debugging, and incident workflows.
Ideal Candidate Profile
Experienced in managing large-scale, high-throughput, low-latency fintech or mission-critical distributed systems with demonstrated impact on availability and operational efficiency.
Skilled in integrating AI/ML technologies into operational tooling, including AIOps, intelligent alerting, anomaly detection, and AI-powered incident response.
Able to influence stakeholders across engineering and product teams, lead complex projects improving system reliability, and mentor others in SRE practices and AI-assisted operations.
