





Tier-1 brand, popular DevOps/SRE title, mid-level experience, metro role, and broad skill requirements.
Strong SRE and applied-AI production requirements make cross-industry transferability limited to experienced platform teams.
Explicit 5+ years SRE experience and extensive mandatory cloud, SRE, and AI production skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own operational reliability, performance, and cost outcomes of high-visibility products and cloud platforms, measured by SLOs and error budgets.
Lead the admission of systems into production enforcing production readiness, observability, and automated reliability checks while ensuring environment integrity and security compliance.
Develop and maintain cloud-native site reliability engineering solutions including AI and agent workload operations, continuous improvement of resilience, and incident management.
Bachelor’s degree in computer science, software engineering, data science, machine learning, or related field.
5+ years in software and site reliability engineering for large-scale distributed cloud-native systems using Python, Go, Kubernetes, Terraform, and observability stacks.
3+ years experience owning SLIs, SLOs, SLAs, error budgets, incident command, production observability, segregation-of-duties controls, and cloud platform ownership with hyperscalers (Azure, AWS, GCP) including AI/ML services.
Experience operating AI/ML workloads in production including understanding of AI failure modes, MLOps/LLMOps, performance and cost engineering, chaos engineering, and load testing.
Experienced senior SRE with demonstrated ownership of operational accountability for cloud-based AI-infused platforms at scale.
Proficient in integrating reliability engineering with production readiness policies, automation, and cross-functional collaboration involving security and risk teams.
Strategic thinker with a preference for incremental, evidence-driven improvements in reliability and performance over large-scale re-engineering interventions.