





Tier-1 brand and metro location increase applicant density despite specialized skillset.
Highly specialized observability and SRE skills limit cross-industry transferability.
Explicit 10+ years, 5+ observability requirement, and a mandatory tech stack create strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build a unified observability stack aligned with OpenTelemetry including metrics, logs, traces, and events with scalable ingestion and processing pipelines.
Develop and enforce instrumentation standards and reusable SDKs for self-service observability across teams.
Own AIOps-ready telemetry pipelines, anomaly detection models, alert correlation engines, SLO/SLI dashboards, and continuously improve alert quality and operational signal architecture.
10+ years overall experience with 5+ years specializing in observability, platform engineering, or SRE with expertise in telemetry and monitoring infrastructure.
Hands-on experience with open-source observability tools like Prometheus, Grafana, OpenTelemetry, Loki, Jaeger, and proficiency in SQL/PromQL/LogQL.
Skilled in designing and operating large-scale telemetry pipelines using tools like Vector or Kafka at production scale.
Strong programming skills in Python and/or Go for automation and building data pipelines supporting ML/AI models.
Experienced in reducing alert noise and improving mean time to detect (MTTD) in production observability environments.
Capable of bridging between data scientists and SRE teams to translate requirements into operational telemetry and signal architectures.
Has deep understanding of distributed systems failure modes and implements comprehensive instrumentation for visibility and data governance.