





Tier-1 brand, metro location, and broad SRE skillset increase applicant density.
Specialized SRE, Kubernetes, GPU and AI infrastructure requirements make cross-industry moves limited.
Explicit 10+ years requirement and many mandatory infra, cloud, and observability skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead technical strategy and roadmap for large-scale SRE initiatives to improve reliability, scalability, and developer efficiency in enterprise systems.
Design and build resilient distributed systems for next-generation AI-powered products, modernizing legacy apps and databases.
Develop automation and observability tools including LLM-aware monitoring and AI-assisted incident response to reduce toil and accelerate mean time to recovery (MTTR).
10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
BS degree in Computer Science or related technical field, or equivalent experience.
Proficiency in Python, Typescript, JavaScript, or Go focusing on automation and distributed systems debugging.
Experience with infrastructure-as-code tools (AWS CDK, CloudFormation, Terraform, CrossPlane) and strong knowledge of Kubernetes, networking, public cloud (AWS, Azure, GCP), and observability at scale including AI workload metrics.
Demonstrated ability to lead complex multi-team technical strategy achieving measurable reliability improvements.
Experience building or operating autonomous/semi-autonomous AI platforms with LLM toolchains and agent orchestration frameworks.
Strong expertise in AI-native platform features, AI-assisted engineering practices, and mentoring engineers to adopt advanced automation and coding agents.