





Tier-1 brand, metro location, popular SRE title, and broad multi-skill requirements raise competition.
Core SRE and observability skills transfer across industries, but platform-specific virtualization and storage add some bias.
Strong mandatory tech stack and on-call expectations create moderate filtering without explicit years.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, deploy, and maintain observability systems including Grafana, Loki, Prometheus, with integrated observability pipelines.
Drive reliability improvements through automation, SLIs/SLOs, root cause analysis, and incident management including on-call rotations and post mortem reviews.
Operate and enhance internal cloud platforms (vCF, CloudStack, Proxmox), Kubernetes clusters, and S3 compatible storage ensuring high availability and performance.
Strong hands-on expertise with Grafana, Loki, Tempo (or similar tracing systems), and Prometheus.
Experience with configuration management tools such as Ansible or Saltstack.
Proficiency in operating modern Linux-based distributed systems and supporting large-scale, highly available architectures.
Work Experience Required: Not explicitly mentioned in the JD.
Has deep expertise in observability engineering specifically in automation and scripting for monitoring and incident response.
Comfortable participating and leading in on-call rotations and incident management with AI-driven tools like Ollama, n8n, MCP.
Experience working with enterprise-grade cloud infrastructure platforms, container orchestration (Kubernetes), and SRE culture including SLIs/SLOs and error budgeting.