





Metro location and attractive hybrid role balanced by specialized SRE+LLM infra requirements, yielding moderate competition.
Core SRE and cloud skills transfer broadly, but LLM governance and insurance compliance increase domain specificity.
Explicit years, mandatory SRE/DevOps leadership and specific tech (AWS, Kubernetes, Terraform, LLM governance) make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and manage the SRE team responsible for core infrastructure, service reliability, and internal platform development including AI toolchains.
Oversee incident management, service level objectives, cloud infrastructure (AWS) architecture using Terraform, and AI governance for secure model routing.
Drive cloud cost optimization, developer productivity improvements, and align SRE roadmap and team goals with company objectives through hands-on technical leadership.
At least 4 years of engineering management experience in high-growth organizations, specifically leading SRE, DevOps, or Infrastructure teams.
6+ years as an individual contributor with proficiency in Golang, Python, or TypeScript.
Deep expertise with AWS, Kubernetes, Terraform, and experience with AI infrastructure and governance tools (e.g., AI gateways, model routing, audit trails).
Proven experience managing infrastructure costs and implementing large-scale cloud cost optimization initiatives.
Experienced leader with a strong technical background in distributed systems, infrastructure as code, and observability tools like DataDog.
Demonstrated ability to balance operational reliability with evolving infrastructure needs for AI platforms and complex systems.
Strategic thinker capable of technical oversight, recruiting and mentoring engineering teams, and advancing cloud cost and AI governance efficiency.