





Tier-1 brand, mid-level SRE title, and metro location increase candidate competition.
Highly specialized SRE and applied-AI operations skills limit cross-industry transferability.
Multiple explicit technical mandates and a 5+ years requirement make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own reliability, performance, and cost outcomes of high-visibility AI-infused products and platforms operating at scale, measured via SLOs and error budgets.
Lead technical efforts for observability, performance and resilience testing, environment integrity, and admission of systems into production with gating on reliability checks and readiness verification.
Develop and operate production observability tools, run performance testing, manage incident response and postmortems, and continually improve operational integrity with automation and collaboration across cross-functional teams.
Bachelor’s degree in computer science, software engineering, data science, machine learning, or related discipline.
5+ years software engineering and site reliability experience with large-scale, distributed, cloud-native systems including skills in Python, Go, Kubernetes, Terraform, CI/CD, and observability stacks.
3+ years experience in site reliability/production engineering owning SLIs, SLOs, error budgets, incident management, environment integrity, and security controls.
3+ years cloud-native platform engineering experience on Azure, AWS, or GCP including AI/ML services, container orchestration, infrastructure-as-code, plus prior experience with operating AI/ML and agentic workloads in production.
Experienced SRE with proven ability in defining and operating reliability metrics (SLAs, SLOs, error budgets) on large-scale cloud-native AI products.
Hands-on engineer who actively codes and automates monitoring, observability, chaos testing, and cost optimization for AI/agentic workloads.
Strong cross-functional collaborator balancing production standards, operational readiness, and continuous incremental reliability improvements within fast-paced, complex environments.