





Tier-1 brand, metro location, and broad cross-functional devops/AI requirements increase applicant competition.
Strong requirement for SRE, cloud IaC, and RAG/LLM operational experience limits cross-industry transferability.
Explicit 7+ years, mandatory cloud/SRE/AI stack and IaC/Kubernetes skills make filters highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead architecture and implementation of AI-driven automation and self-healing infrastructure across cloud AI/ML platforms to reduce operational toil and improve MTTR.
Develop advanced diagnostic, RAG-based knowledge tools, conversational AI agents, and automated triage systems integrating with ticketing for Tier-1 and Tier-2 support.
Build and manage observability pipelines, predictive anomaly detection, and executive dashboards to ensure system uptime and SLA compliance across AI ecosystems.
7+ years of hands-on experience in software engineering, SRE, DevOps, or systems operations focused on cloud infrastructure automation and AI/ML operational tooling.
Bachelor’s or Master’s degree in CS, Software Engineering, IT, or related quantitative field.
3+ years experience with cloud ecosystems (AWS or GCP) managing AI/ML compute platforms and native cloud AI services (e.g., AWS Bedrock, SageMaker, GCP Vertex AI).
Proficiency in infrastructure-as-code (Terraform, CloudFormation, or GCP Deployment Manager), Kubernetes/Docker, and advanced scripting (Python, Go, Shell).
Experienced technical architect comfortable bridging software engineering, SRE, and AI/ML operations in high-growth AI product environments.
Proven leader capable of driving complex automation projects independently and influencing cross-functional teams to adopt new operational models.
Strong focus on root-cause analysis and long-term architectural solutions to eliminate recurring incidents and improve operational efficiency.