AI/MLOps SRE Lead Engineer
Regeneron Pharmaceuticals, Inc.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead reliability, scalability, and operational excellence for AI/ML and cloud platform infrastructure across multi-cloud environments.
Design, build, and scale enterprise ML platforms (e.g., Dataiku, SageMaker, Databricks, Vertex AI) and AI/ML observability and automation solutions.
Architect and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, ChatOps integrations, and guide platform modernization and technical strategy.
Minimum Requirements
Bachelor's degree in Computer Science, IT, Engineering, Data Science, AI, or related field; Master's preferred.
6-8 years experience in Site Reliability Engineering, Platform Engineering, DevOps, or related fields with enterprise-scale delivery.
Strong hands-on experience with two or more major cloud platforms including AWS, GCP, or Azure.
Proficiency with ML platforms (Dataiku, SageMaker AI, Databricks, Vertex AI), Infrastructure as Code (Terraform, Pulumi, AWS CDK), and programming in Python, Go, or Bash.
Ideal Candidate Profile
Experienced leader in reliability engineering for AI/ML platforms with deep knowledge of enterprise-grade ML workflows and cloud-native technologies.
Skilled in designing observability, anomaly detection, predictive analytics, automated remediation, and ChatOps integrated operational workflows.
Proven ability to assess technical debt, influence platform strategy, implement platform modernization, and mentor technical teams in SRE and MLOps domains.
