





Metro location and broad multi-cloud SRE/ML-Ops demand elevates applicant competition to medium.
Specialized GenAI/ML-Ops SRE skills and regulated-industry experience limit cross-industry transferability.
Mandatory 7+ years and specific SRE/ML-Ops skills and cloud certifications increase shortlisting rigor.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, deploy, and operate scalable AI-enabled platforms and agentic AI workflows with a focus on reliability, performance, and cost optimization.
Lead SRE practices including incident response, capacity planning, automation, observability, and governance for AI/ML production systems.
Collaborate across teams to translate requirements into reliable AI/ML deployments while ensuring security, data privacy, and compliance.
7+ years experience in AI/ML, data engineering, or DevOps with production exposure to AI/ML deployments or reliability projects.
Proficiency in Python for automation and orchestration; experience with cloud platforms (AWS, Azure, or GCP) including security, IAM, and networking.
Experience with containerization (Docker), CI/CD pipelines, infrastructure as code (preferably Terraform), and monitoring tools (Prometheus, Grafana, CloudWatch).
Bachelor’s or Master’s degree in Computer Science, Data Science, AI, or related field.
Experienced in integrating SRE best practices specifically for AI/ML systems at scale and in regulated environments with governance and security controls.
Proven ability to lead and mentor cross-functional, often distributed teams in delivering reliable, secure, and cost-effective AI infrastructure.
Strong strategic and operational understanding of multi-cloud architectures, automation, and AI production workflows with emphasis on observability and incident management.