





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, mid-level generalist SRE, and metro/hybrid role create high applicant competition.
Role requires specialized cloud, Kubernetes, and SRE practices, limiting cross-industry transfer without platform experience.
Explicit 6-8 years plus mandatory Azure, Kubernetes, IaC and observability skills make filters strict.
Define, monitor, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets specifically for AI services.
Maintain and operate scalable, fault-tolerant cloud infrastructure (Azure) including Kubernetes clusters and containerized AI microservices with IaC tools like Terraform.
Lead incident response, root cause analysis, proactive reliability risk mitigation, and build automation for operational tasks including CI/CD pipelines and intelligent alerting systems.
6-8 years of relevant work experience in supporting production-grade, high-availability systems.
Bachelor’s degree in Computer Science, Information Technology, or a related field.
Strong proficiency with Azure cloud services (AKS, Azure DevOps, ARM) and Infrastructure as Code tools such as Terraform.
Experience with containerization (Docker), orchestration (Kubernetes), CI/CD tools, and maintaining observability stacks using Prometheus, Grafana, Datadog, or Open Telemetry.
Experienced in managing multi-tenant, microservices-based architectures with a focus on AI/ML model serving and data workflows.
Demonstrates a reliability-first mindset with hands-on leadership in incident management and production stability for AI services.
Skilled in automation and MLOps practices, partnering effectively with data engineers and data scientists to operationalize AI systems securely and scalably.