





Tier-1 brand, metro location, broad generalist SRE skills, and visible Staff title increase competition.
Requires deep SRE, Kubernetes, cloud, and GPU/AI platform experience, making backgrounds less transferable.
Explicit 10+ years requirement plus mandatory cloud, Kubernetes, IaC, and AI/GPU infra skills enforce high shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead technical strategy and roadmap for large-scale SRE initiatives improving reliability, scalability, and developer efficiency across enterprise systems.
Design and build resilient distributed systems for next-gen AI-powered enterprise products, transforming legacy apps and databases into scalable architectures.
Develop automation, observability, and AI-assisted monitoring systems to reduce toil, accelerate incident response, and enhance performance and reliability.
10+ years experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
BS degree in Computer Science or related technical field involving coding, or equivalent experience.
Proficiency in Python, Typescript, JavaScript, or Go focused on automation and infrastructure-as-code.
Experience with infrastructure-as-code tools (AWS CDK, CloudFormation, Terraform, CrossPlane) and strong knowledge of Kubernetes, public cloud (AWS/Azure/GCP), systems architecture, and observability at scale including AI workload metrics.
Experienced in driving technical strategy and delivering measurable reliability improvements in complex, multi-team environments.
Hands-on expertise in building or operating autonomous AI platforms with LLM toolchains or agent orchestration frameworks.
Skilled at integrating AI-assisted engineering practices and on-call operations using AI-native platform features.