





Specialized MLOps SRE role at a known employer in Hyderabad with mid-level experience, moderate applicant density.
Requires specialized MLOps, SRE, and ML platform experience, limiting cross-industry transferability.
Explicit 6-8 years plus required multi-cloud, ML platform, IaC, observability, and programming skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead reliability, scalability, and operational excellence for AI/ML and cloud-native platforms across multi-cloud environments with established SLOs, SLIs, and error budgets.
Design, build, and scale enterprise ML platform infrastructure and AI-driven observability solutions including anomaly detection and automated remediation.
Architect and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, ChatOps integrations, and platform automation to enhance engineering productivity and resilience.
Bachelor's degree in Computer Science, IT, Engineering, Data Science, AI or related field; Master's preferred.
6-8 years experience in Site Reliability Engineering, Platform Engineering, DevOps, or related with enterprise-scale delivery.
Hands-on experience with 2+ major cloud platforms (AWS, GCP, Azure) and ML platform technologies such as Databricks, Amazon SageMaker, Dataiku, Google Vertex AI.
Proficient in Infrastructure as Code tools (Terraform, Pulumi, AWS CDK) and CI/CD automation; strong programming/scripting skills in Python, Go, Bash, or similar.
Experienced in architecting and leading AI/ML platform reliability engineering with demonstrated strategic influence on platform modernization.
Strong operational expertise in managing enterprise-scale multi-cloud AI/ML platforms integrating advanced observability and automated remediation.
Capable of driving cross-functional partnerships with Data Science and AI Engineering teams to deliver secure, scalable AI/ML production solutions.