





Specialized ML-Ops SRE skillset reduces candidate pool despite common DevOps demand at a Tier-2 firm.
Requires specialized ML-Ops SRE and cloud reliability expertise, limiting straightforward cross-industry transferability.
Explicit 7+ years plus mandatory SRE, cloud, ML-Ops, Python, Kubernetes, and IaC skills make filters stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, deploy, and operate scalable AI-enabled platforms and agentic workflows emphasizing reliability, performance, and cost efficiency.
Lead cross-functional teams to implement SRE practices including automation, observability, incident response, capacity planning, and governance for AI/ML solutions.
Optimize AI/ML production systems maintaining service-level objectives, security controls, and compliance while mentoring junior engineers and maintaining technical documentation.
7+ years of experience in AI/ML, data engineering, or DevOps with significant production exposure.
Proficiency in Python and experience with deploying and operating AI/ML solutions including agentic AI workflows and large-scale model deployments.
Strong cloud experience with AWS, Azure, or GCP including security, IAM, networking, and cost optimization; familiarity with containerization tools like Docker and Kubernetes.
Bachelor’s or Master’s degree in Computer Science, Data Science, AI, or related field.
Experienced in blending traditional SRE practices with AI/ML deployments including security, governance, and reliability engineering for complex systems.
Capable of leading and mentoring cross-functional teams while driving technical risk assessments, disaster recovery, and cost/resource optimization in regulated or enterprise environments.
Skilled in multi-cloud and cloud-native technologies with hands-on expertise in automation, orchestration, CI/CD pipelines, and observability tooling for AI workloads.