





Mid-level metro SRE role; Prometheus specialization limits applicants, keeping competition moderate.
Specialized monitoring skills transfer across infra-focused companies but not easily to non-technical sectors.
Explicit 5+ years and mandatory Prometheus/Grafana SME skills create strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and manage the health, availability, and configuration of Prometheus and Grafana monitoring platforms for production environments.
Design, implement, and maintain monitoring and alerting solutions to ensure proactive detection of issues and reduction of alert noise.
Participate in on-call rotations and incident response as the Subject Matter Expert for Prometheus and Grafana platforms.
5+ years experience in SRE, Platform, Monitoring, or Observability roles with L3 responsibilities or equivalent expertise.
Strong hands-on expertise in Prometheus administration including scraping, service discovery, relabeling, recording and alert rules, and exporters ecosystem.
Strong hands-on expertise in Grafana including dashboard design, templating, variables, and query optimization.
Scripting or automation skills using Bash, Python, PowerShell, or equivalent.
Experienced in operating high-availability monitoring platforms in production with knowledge of alerting best practices and platform reliability improvements.
Familiar with cloud-native/Kubernetes monitoring environments using Prometheus, including concepts like federation and remote-write (optional but advantageous).
Comfortable working within a tiered support model and escalation processes, capable of handling critical incident bridge calls as the SME.