Site Reliability Engineer
LSEG (London Stock Exchange Group)Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure reliability, supportability, and usability of LSEG's internal observability platform, focusing on platform health monitoring, issue diagnosis, and service recovery.
Lead incident response, problem management, root cause analysis, and drive continuous improvement to reduce repeat issues and manual work.
Maintain documentation, onboarding guides, and provide enablement on telemetry standards, alerting guidance, SLO practices, and platform workflows for self-service adoption.
Minimum Requirements
Experience supporting production platforms or services in SRE, platform engineering, infrastructure, DevOps, or operations engineering roles.
Proficiency in using observability data (metrics, logs, traces, alerts, dashboards) to investigate issues and operational support processes (incident response, problem management).
Knowledge of automation or infrastructure-as-code practices and solid understanding of cloud, container, Linux, networking, or distributed system environments.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced in platform operations focusing on reliability and automation within observability or telemetry platforms used by multiple engineering teams.
Familiarity with defining/using SLOs, SLIs, error budgets, and supporting monitoring pipeline technologies and GitOps workflows.
Ability to escalate operational improvements, reduce manual work, and empower engineering teams to adopt shared platform standards and self-service capabilities.
