





Tier-1 brand, mid-level SRE title, metro hiring, and broad skillset drive high applicant competition.
Specialized SRE and applied-AI operational expertise limits cross-industry transferability.
Explicit 5+ years and numerous mandatory SRE, cloud, and AI ops skills make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Ensure reliability, performance, cost-effectiveness, and operational integrity of high-visibility AI-infused products and cloud platforms by owning SLOs, error budgets, and incident management.
Lead production system admission by gating releases based on reliability checks, observability, and readiness verification, enforcing production standards and environment integrity.
Develop and operate scalable automation, conduct performance testing, chaos engineering, and collaborate with cross-functional teams to optimize platform reliability and cost engineering with applied AI fluency.
Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, or related discipline.
5+ years software/site reliability engineering experience with large-scale distributed cloud-native systems; skilled in Python, Go, Bash, Java, Kubernetes, Terraform, CI/CD, and observability tools.
3+ years site reliability or production engineering experience including defining and owning SLIs, SLOs, incident command, and environment integrity controls.
3+ years cloud-native platform ownership experience on hyperscalers (Azure, AWS, GCP) including AI/ML services and operating AI/agentic workloads in production reliability contexts.
Experienced in operating AI/ML workloads with understanding of reliability failure modes like drift, train/serve skew, and cost anomalies relevant to AI systems.
Demonstrates expertise in designing and enforcing SRE principles with strong incident management, observability, performance testing, chaos engineering, and cloud/AI cost engineering capabilities.
Collaborates effectively with product engineering, security, and risk teams to maintain production standards and enable high-quality, resilient cloud services at scale.