Lead Engineer – Reliability & Observability
Cybrilla TechnologiesMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own, scale, and continuously improve the Reliability & Observability platform including metrics, logging, tracing, dashboards, and alerting supporting multi-tenant, multi-site deployments.
Establish and drive reliability engineering standards such as SLO frameworks, error-budget reporting, production readiness, incident management processes, and automation to reduce operational toil.
Provide technical leadership by defining architecture and roadmap, mentoring engineers, and collaborating with product, infrastructure, and security teams to build resilient, self-operating distributed systems.
Minimum Requirements
7+ years of software engineering, platform engineering, or SRE experience with high-scale, business-critical production systems.
Strong hands-on experience with observability tech: ELK stack (Elasticsearch, Logstash, Kibana), Grafana, Loki, OpenTelemetry, Jaeger or equivalents.
Experience with AWS, Kubernetes, Linux, networking, distributed systems, and strong programming/scripting skills in Go, Java, or Python with automation-first approach.
Work Experience Required: 7+ years of relevant engineering experience; no explicit notice period or educational requirements mentioned.
Ideal Candidate Profile
Engineer with deep expertise in building and scaling production observability and reliability platforms in complex, distributed, multi-tenant cloud environments.
Proven ability to independently lead technical initiatives, make architectural decisions, and influence multiple engineering teams without formal authority.
Experience in fintech or regulated environments preferred, with familiarity in large-scale telemetry pipelines, cost optimization, chaos engineering, and self-service developer tools.
