





High due to Tier-1 brand, metro location, popular SRE role, and broad AI/cloud skill requirements.
High because role demands niche AI/GPU infrastructure, LLM-aware monitoring, and deep SRE/platform experience.
High because of explicit 10+ years requirement, mandatory SRE/platform/cloud expertise, and specialized AI/GPU infrastructure skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead technical strategy and roadmap for multi-functional Site Reliability Engineering initiatives to enhance reliability, scalability, and developer efficiency across enterprise systems.
Design and develop resilient distributed systems and transform legacy applications into modern scalable architectures for AI-powered products and services.
Develop automation, observability, and AI-assisted monitoring/incident response pipelines to improve performance, reduce toil, and evolve on-call operations.
10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
BS degree in Computer Science or related technical field involving coding, or equivalent experience.
Proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with experience in automation and infrastructure-as-code tools (e.g., AWS CDK, Terraform).
Deep expertise in systems architecture, networking, Kubernetes, public cloud services (AWS, Azure, or GCP), and observability implementations at scale, including AI workload telemetry.
Experienced leader capable of steering complex, multi-team technical strategies and delivering measurable improvements in reliability and scalability.
Proven ability to build or operate autonomous or semi-autonomous AI platforms, including LLM toolchains and agent orchestration frameworks.
Strong background in applying AI-first engineering practices and mentoring teams in agentic development workflows and AI-assisted operations.