





Strong employer brand plus metro location and senior role create moderate applicant density and competition.
Role requires specialized SRE, MLOps, and cloud platform experience, making cross-industry fit moderately sensitive.
Explicit 8+ years, mandated SRE/cloud/AI platform skills and certifications make filters strict and prescriptive.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Operate and improve highly available, scalable AI platforms supporting Generative AI, LLM, RAG, and Agentic AI workloads with focus on Site Reliability Engineering, Cloud Engineering, and Platform Operations.
Implement and support cloud-native infrastructure using Kubernetes, Infrastructure-as-Code, and CI/CD tooling across Azure, AWS, and GCP to ensure platform reliability, performance, and security.
Develop and maintain observability, monitoring, security controls, incident management, and automation frameworks while partnering with AI/ML, Data Engineering, and Platform teams for production readiness and operational excellence.
Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
Minimum 8 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, Infrastructure Engineering, or Production Operations.
Hands-on experience with Kubernetes, containers, cloud-native infrastructure, Infrastructure-as-Code (Terraform, ARM/Bicep, CloudFormation, or equivalent), CI/CD pipelines, and observability platforms.
Experience operating production AI/ML platforms including Generative AI, LLM, RAG, or model-serving workloads; knowledge of IAM, RBAC, encryption, secrets management, and cloud security practices.
Experienced in supporting enterprise-scale AI platforms and cloud-native distributed systems with deep knowledge of SRE practices in regulated environments such as healthcare or financial services.
Strong operational focus with proven ability to implement observability, security, and automation frameworks while collaborating with cross-functional AI and engineering teams.
Skilled in multi-cloud environments (AWS, Azure, GCP), Kubernetes orchestration, and advanced AI platform tooling (MLOps, LLMOps, model serving technologies) with a track record of driving platform reliability and cost optimization initiatives.