





Recognizable brand, popular SRE title, and mid-level metro hiring increase candidate competition.
Observability and SRE skills transfer broadly, but platform-specific virtualization increases domain specificity.
Multiple mandatory observability, infra, and scripting requirements make shortlisting stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, deploy, and maintain observability systems (Grafana, Loki, Prometheus) and AI-driven incident diagnosis workflows.
Ensure high availability, capacity planning, and performance optimization of on-premise cloud platforms (vCF, CloudStack, Proxmox, Kubernetes).
Participate in global on-call rotation, managing incidents using AI tools and coordinate post-incident reviews.
Hands-on expertise with Grafana, Loki, Prometheus, Tempo (or similar tracing systems), and scripting/programming language.
Experience operating Linux-based distributed systems and supporting large scale, highly available architectures.
Familiarity with Kubernetes, CI/CD pipelines, Infrastructure as Code, Configuration Management tools (Ansible/Saltstack), and Atlassian tools (Jira, Confluence, Opsgenie).
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in observability engineering and automation within enterprise-grade IT platform environments.
Comfortable managing complex incident response including participation in on-call rotations and using AI-powered diagnosis tools.
Familiar with on-premise cloud virtualization stacks (vCF, CloudStack, Proxmox) and Kubernetes cluster operations.