Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead technical architecture and reliability standards for AI runtime systems within Ad Platforms, focusing on production readiness and resilience.
Define and enforce latency SLOs, failover strategies across multiple AI model providers, and design scalable backend services handling high-volume AI traffic.
Drive operational excellence through incident analysis, systemic risk identification, and partnering with cross-functional teams to harden AI-powered services before revenue-impacting scale.
Minimum Requirements
8+ years backend software engineering experience building distributed systems at scale.
Strong proficiency in Python or similar backend language.
Experience operating AI-powered/model-dependent services in production and familiarity with LLM provider constraints.
Bachelor’s degree in Computer Science, Engineering, or related field; Master’s preferred.
Ideal Candidate Profile
Expert in designing resilient microservices architectures with strong retry semantics, circuit breakers, and graceful degradation patterns.
Experienced with multi-cloud environments (AWS, Azure, GCP) and multi-provider AI failover strategies (Azure, OpenAI, Bedrock).
Demonstrated leadership in defining architectural reliability standards, concurrency modeling, and post-incident reliability improvements for AI or distributed backend systems.
