Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and improve the reliability, observability, and scalability of ServiceTitan's Kubernetes-based cloud infrastructure.
Respond to and resolve production incidents, including root cause analysis and remediation to maintain service health.
Develop automation, runbooks, and best practices to reduce manual operational work and improve on-call effectiveness.
Minimum Requirements
10+ years of relevant hands-on experience in Site Reliability Engineering or related roles.
Strong hands-on expertise with Kubernetes systems.
Proficient in cloud platforms (AWS, Azure, or GCP) with solid networking fundamentals (subnetting, IP addressing).
Practical experience with SLIs, SLOs, error budgets, and modern observability tools (e.g., Prometheus, Grafana, Datadog).
Ideal Candidate Profile
Experienced in managing large-scale distributed systems and handling livesite incidents with effective troubleshooting under pressure.
Familiar with CI/CD pipelines (preferably GitHub Actions) and automation to support fast and safe deployments.
Experienced with AI-assisted engineering tools for automation, root cause analysis, and production system reliability enhancement.
