





Tier-1 brand, mid-level SRE role, metro location, and broad skillset amplify competition.
Platform SRE skills transfer across industries but require specific OpenShift and AI/ML infrastructure experience.
Explicit 4+ years plus mandatory OpenShift, SRE, Golang/Python, IaC and observability requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design, deployment, and administration of scalable, highly available Red Hat OpenShift platforms supporting AI/ML workloads including Large Language Models (LLMs).
Implement and drive Site Reliability Engineering (SRE) practices to ensure platform reliability, scalability, performance, and operational excellence in AI-focused environments.
Develop automation tools and platform services using Golang and/or Python; manage full cluster lifecycle, CI/CD pipelines, observability, incident resolution, and platform security.
4+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related roles.
Strong hands-on experience with Red Hat OpenShift administration, operations, and troubleshooting.
Proficiency in Golang and/or Python for automation and platform engineering.
Hands-on experience with AI/ML platforms, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and GPU architectures.
Experienced in operating and scaling Kubernetes/OpenShift platforms in AI/ML production environments with a reliability-first and automation-driven mindset.
Skilled in working across global, cross-functional engineering and research teams to align platform capabilities with machine learning workloads.
Capable of managing critical production environments with 16x5 on-call responsibility, focusing on incident management and proactive problem resolution.