Director, Site Reliability Engineering (Production Engineering)
Zscaler, Inc.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own end-to-end operational availability, latency, and performance for Tier-0 and Tier-1 internal cloud platforms processing hundreds of billions of daily transactions with 99.99%+ uptime.
Lead SRE teams across India and global HQ to enforce SLAs, SLOs, and error budgets, and execute engineering initiatives to reduce MTTD, MTTA, and MTTR.
Drive capacity planning, implement alert frameworks, and oversee 24/7 automated incident mitigation and operational reliability at scale.
Minimum Requirements
12+ years engineering experience including 5+ years managing distributed engineering managers and teams.
Foundational understanding of AI/ML technologies relevant to operational domain.
Hands-on experience with cloud-native hyper-scale platforms (Kubernetes, multi-cloud), modern observability architectures (OpenTelemetry, Prometheus, Kafka, Elasticsearch), and SRE principles.
Experience working with globally distributed engineering teams and US-based leadership.
Ideal Candidate Profile
Proven ability to lead large SRE teams driving operational rigor in high-scale cloud platforms with near-zero downtime targets.
Strong strategic and hands-on execution skills balancing technical depth with leadership and scaling organizational capabilities across geographies.
Comfortable translating complex technical trade-offs to executives, with experience leveraging AI-driven solutions to improve operational efficiency and decision-making.
