Lead Engineer – Reliability & Observability
Cybrilla TechnologiesMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and scale the Reliability & Observability platform including metrics, logging, tracing, dashboards, and alerting across multi-tenant, multi-site deployments.
Establish SLO frameworks, error-budget reporting, incident management processes, and reliability engineering best practices to improve system reliability.
Drive automation of operational processes, provide technical leadership, and collaborate across product, infrastructure, and security teams to build resilient distributed systems.
Minimum Requirements
7+ years in software engineering, platform engineering, or SRE with experience operating business-critical production systems at scale.
Strong hands-on experience with observability technologies (ELK stack, Grafana, Loki, OpenTelemetry, Jaeger).
Proficiency in AWS, Kubernetes, Linux, networking, distributed systems, and programming/scripting in Go, Java, Python, or similar.
Experience designing production observability platforms, including monitoring, SLOs, alerting, incident management, and automation.
Ideal Candidate Profile
Proven ability to lead technical initiatives and influence engineering teams without formal authority, focusing on platform reliability and scalability.
Experience operating and evolving large-scale, multi-tenant SaaS platforms with distributed infrastructure.
Strong background in designing resilient distributed systems with expertise in observability architectures and automation-based operational excellence.
