





Tier-1 brand, metro location, mid-level generalist DevOps role, broad observability skillset increases competition.
Observability and SRE specialization transfers across industries but requires specific tooling expertise.
Multiple mandatory tooling and platform skills plus explicit Python experience create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build, scale, and maintain the enterprise-wide observability and log management ecosystem, focusing on open telemetry pipeline transformation to ingest terabytes of logs, metrics, and traces daily.
Configure and maintain data distribution pipelines using Cribl; manage analytical environments in Sumo Logic and Datadog and configure PagerDuty for incident response.
Automate monitoring agent deployments and pipeline configurations using Ansible; enforce telemetry schema standards; mentor internal teams with technical documentation and enablement sessions.
Minimum 3 years of Python development experience.
Proven expertise in Datadog including AWS integrations and dashboard templating.
Experience applying Site Reliability Engineering (SRE) concepts and working with AWS architecture and cloud-native observability.
Familiarity with OpenShift or Kubernetes, Ansible, Infrastructure-as-Code concepts, and OpenTelemetry.
Experienced in enterprise-scale monitoring and observability platform transformations involving multiple tools (Datadog, Sumo Logic, SignalFX/Splunk).
Strong background working with Infrastructure, Application, and DevOps teams to build relevant metrics and observability solutions.
Ability to manage and optimize telemetry data flows for cost and operational efficiency, with a focus on automation and standardization across large distributed systems.