





High—strong employer brand, mid-level SRE title, and in-demand cloud/Kubernetes skills increase applicant density.
Medium—core SRE and cloud skills transfer across industries, but GenAI and healthcare context increase domain specificity.
High—explicit multi-year requirements, mandatory cloud/Kubernetes/Terraform, AI/ML production experience, and on-call obligations.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design and implementation of AI Ops and RAG-based AI solutions from concept to production, ensuring responsible AI practices.
Own solution architecture and automation of cloud infrastructure using Terraform, Python, and GitHub Actions to maximize reliability and scalability in public cloud environments.
Lead incident response, mentor engineers, and drive improvements in Site Reliability Engineering practices including monitoring and operational efficiency for AI workloads.
Bachelor's degree in Information Systems, Computer Science, Engineering, or related field or equivalent certification.
5+ years of combined software engineering and Site Reliability Engineering experience in public cloud environments such as GCP, AWS, or Azure.
3+ years hands-on experience with Python and Terraform development; 2+ years delivering AI/ML or Generative AI production solutions.
2+ years managing Kubernetes environments (EKS, AKS, GKE, or self-hosted); Available for rotating 24x7 on-call shifts.
Experienced in leading design and production of AI Ops solutions and RAG pipelines for enterprise scale with a focus on responsible AI.
Strong technical proficiency in cloud infrastructure automation, observability, and SRE best practices specifically for AI workloads.
Capable of mentoring, leading incident responses, and interfacing with diverse technical and non-technical stakeholders to improve system reliability and operational efficiency.