





Metro Hyderabad and mid-level seniority, but specialized MLOps skills reduce applicant density.
MLOps SRE skills transfer across industries but require specific ML platform and tooling experience.
Explicit 6-8 years requirement plus mandatory cloud, MLOps, IaC, observability, and platform tooling.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead reliability, scalability, and operational excellence for AI/ML and cloud platform infrastructure across multi-cloud environments.
Design, build, and scale enterprise ML platforms (e.g., Dataiku, SageMaker, Databricks, Vertex AI) and AI/ML observability and automation solutions.
Architect and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, ChatOps integrations, and guide platform modernization and technical strategy.
Bachelor's degree in Computer Science, IT, Engineering, Data Science, AI, or related field; Master's preferred.
6-8 years experience in Site Reliability Engineering, Platform Engineering, DevOps, or related fields with enterprise-scale delivery.
Strong hands-on experience with two or more major cloud platforms including AWS, GCP, or Azure.
Proficiency with ML platforms (Dataiku, SageMaker AI, Databricks, Vertex AI), Infrastructure as Code (Terraform, Pulumi, AWS CDK), and programming in Python, Go, or Bash.
Experienced leader in reliability engineering for AI/ML platforms with deep knowledge of enterprise-grade ML workflows and cloud-native technologies.
Skilled in designing observability, anomaly detection, predictive analytics, automated remediation, and ChatOps integrated operational workflows.
Proven ability to assess technical debt, influence platform strategy, implement platform modernization, and mentor technical teams in SRE and MLOps domains.