





Tier-1 brand, mid-level SRE title, metro location, and broad skillset increase applicant competition.
Core SRE skills transfer across industries, but banking compliance and enterprise controls increase domain specificity.
Multiple mandatory technical skills, SRE experience, and security controls make shortlisting stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and support scalable, resilient AI/ML data platform applications using Databricks, Snowflake, AWS, and Kubernetes.
Coordinate and run production incident management, including root cause analysis and implementation of production changes.
Develop AI/ML-driven operational solutions to troubleshoot, automate toil reduction, and improve reliability with measurable SLO outcomes.
2+ years applied experience with formal training or certification in security engineering concepts.
Proficiency in site reliability engineering principles, incident management, observability tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk).
Strong skills in Python or PySpark for AI/ML modeling along with working knowledge of enterprise-authorized AI capabilities in SRE workflows.
Work Experience Required: Not explicitly mentioned in the JD.
Experience with production support or SRE roles focused on AWS Cloud, Databricks, and Snowflake technologies (4+ years preferred).
Ability to apply AI-assisted operational recommendations with attention to resiliency, security, risk controls, and auditability.
Experience mentoring teams and partnering globally to drive strategic change in AI/ML data platform reliability.