





Tier-1 brand, mid-level SRE role, common skillset, and metro hiring increase applicant competition.
Role requires specialized observability and SRE experience, moderately limiting cross-industry transfers.
Mandatory 4+ years plus specific observability, SRE, and tooling experience increases candidate filtering stringency.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Manage and evolve end-to-end observability solutions (metrics, logs, traces) across infrastructure and Kubernetes environments to improve system reliability and production visibility.
Operate and maintain core observability platforms (e.g., Splunk, New Relic, Grafana), including configuration, upgrades, access management, and ensuring platform SLAs.
Lead adoption of instrumentation standards, optimize telemetry pipelines for performance and cost, and drive incident analysis to improve mean time to detect and resolve (MTTD, MTTR).
Bachelor's degree in Computer Science or equivalent practical experience.
4+ years experience in Observability, Site Reliability Engineering, DevOps, or platform engineering roles supporting production systems.
Hands-on experience administering observability/APM platforms like Splunk, New Relic, or Grafana with knowledge of metrics, logs, distributed tracing, and platform configuration.
Experience with Kubernetes, cloud-native environments, telemetry pipeline management, and proficiency in scripting or automation (e.g., Shell, Python).
Experienced in scalable distributed systems observability, including MELT (Metrics, Events, Logs, Traces), SLIs/SLOs, alert tuning, and root cause analysis.
Proven ability to influence and collaborate with cross-functional teams to enhance instrumentation and monitoring standards focused on SLO-driven reliability.
Skilled at optimizing observability pipelines balancing performance, data quality, and cost, with exposure to incident postmortem leadership and continuous improvement.