Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead and develop a globally distributed Observability engineering team responsible for metrics, logging, alerting, and capacity planning platforms for GitLab.com and GitLab Dedicated.
Own reliability, scalability, and cost efficiency of observability systems, improve alerting and telemetry using SLOs and error budgets, and participate in high-severity incident response and on-call rotations.
Collaborate with Site Reliability, Product Engineering, and Infrastructure teams to enhance service observability and apply AI tools to support engineering workflows and incident triage.
Minimum Requirements
Experience leading observability, platform engineering, or site reliability teams in a distributed, asynchronous setting at scale.
Technical expertise with Prometheus metrics systems, logging platforms (e.g., Elasticsearch or cloud-native), alerting design, and capacity forecasting.
Proven experience managing large SaaS platforms with hands-on incident coordination and production on-call participation.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Strategic leader capable of balancing technical tradeoffs and operational priorities across distributed teams and stakeholders.
Strong background in observability tooling, production systems reliability practices, and using SLOs/error budgets to inform investment decisions.
Experienced in operating in asynchronous, global environments and applying AI to improve engineering and incident management workflows.
