





Tier-1 brand, mid-level SRE role with broad cloud and applied-AI skill requirements increases competition.
Requires specialized SRE and applied-AI production experience, making hires less transferable across industries.
Multiple mandatory SRE/cloud/AI experience requirements and specific tech stack increase shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own reliability, performance, and operational integrity of high-visibility AI-infused cloud-native products/platforms, maintaining SLOs within budget and managing incident trends to improve reliability.
Lead production admission of systems through readiness verification, resiliency testing, observability design, and automated reliability checks while collaborating cross-functionally with engineering, security, and risk teams.
Develop and operate production observability tooling and automation, manage environment integrity, cost engineering, and chaos testing to ensure resilient, scalable, and cost-effective operations of AI/ML workloads.
Bachelor’s degree in computer science, software engineering, data science, machine learning, or related discipline.
5+ years in software/site reliability engineering operating large-scale, distributed, cloud-native systems; 3+ years specifically in site reliability or production engineering with defined SLIs/SLOs and incident management.
3+ years experience with cloud-native engineering on Azure, AWS, or GCP including AI/ML services, container orchestration, infrastructure-as-code, and multi-environment management.
Prior experience operating AI/ML and agentic workloads in production with understanding of their reliability failure modes, and experience with load/performance testing, chaos engineering, capacity planning, and FinOps tooling.
Experienced in modern SRE practices with a strong engineering background in cloud platform ownership, observability, and AI-infused workload operations at scale.
Operates with a focus on measurable reliability outcomes (e.g., SLOs, error budgets) and pragmatic incremental improvements over big-bang fixes.
Able to collaborate deeply across product, engineering, security, risk, and platform teams to enforce production standards and maintain environment integrity with segregation-of-duties controls.