Lead Site Reliability Engineer - Platform Engineering / SRE
ZenotiMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, availability, performance, and scalability of Zenoti's non-database platform components including web applications, APIs/microservices, background processing, and analytics workloads on cloud infrastructure (preferably Azure).
Lead incident response, observability practices, and define SLIs/SLOs/SLAs and error budgets with engineering teams to ensure resilient design and safe delivery.
Drive capacity planning, automation of operational workflows, architectural governance, and mentor engineers while serving as a Subject Matter Expert on reliability and cloud best practices.
Minimum Requirements
12+ years overall Software Engineering experience with at least 4+ years in Site Reliability Engineering or production operations of web and API-based systems.
Strong hands-on experience with public cloud platforms, preferably Azure services including App Service, AKS/VMs, Application Gateway, Load Balancers, Azure Monitor, Key Vault, etc.
Proficiency in monitoring and observability tools (APM, distributed tracing, log aggregation, metrics, dashboards) and solid knowledge of CI/CD pipelines such as Azure DevOps or Jenkins.
Strong automation skills using Python or similar scripting languages and experience with Infrastructure as Code tools like Terraform or ARM/Bicep.
Ideal Candidate Profile
Experienced leader capable of influencing multiple engineering teams and mentoring junior engineers on complex technical issues related to reliability and cloud operations.
Deep understanding of web and API production environments including microservices, containers (Docker/Kubernetes), and analytics/reporting platforms within cloud ecosystems.
Proven ability to independently own end-to-end reliability processes including incident management, root cause analysis, and implementing scalable infrastructure solutions on Azure cloud.
