





Tier-1 brand, popular SRE role, metro location, and broad required skillset increase applicant competition.
SRE skills broadly transferable, but banking compliance and enterprise controls moderately raise industry specificity.
Explicit years, mandatory SRE domain skills, observability and cloud requirements create stringent technical filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and support scalable, resilient AI/ML data platform solutions using Databricks, Snowflake, AWS, and Kubernetes.
Coordinate and lead incident management and root cause analysis to resolve application issues in production.
Develop AI/ML solutions and use enterprise AI capabilities to identify and automate toil reduction and improve service reliability based on measurable SLOs.
2+ years applied experience with formal training or certification on security engineering concepts.
Proficiency in incident management, site reliability principles, and observability tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk.
Strong Python or PySpark skills for AI/ML modeling and experience applying AI capabilities in operational environments with data sensitivity awareness.
Work Experience Required: 2+ years in site reliability or related production support roles.
Experienced in managing production incidents and applying SRE principles to AI/ML data platform environments.
Hands-on with system design, resiliency, testing, operational stability, and disaster recovery in cloud environments (preferably AWS).
Capable of assessing AI-driven operational recommendations critically to maintain security, resiliency, and compliance standards.