





Tier-1 employer and metro location create moderate applicant competition.
Role requires specialized LLM/inference and GPU fleet expertise, limiting easy cross-industry transfers.
Explicit 9+ years, 2+ years management, and deep LLM/GPU/ML infra requirements raise filter strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and manage AI engineering teams responsible for scalable AI platform development and infrastructure.
Architect and optimize production-grade AI systems handling millions of low-latency real-time model inferences with multi-region inference and multi-GPU fleet management.
Oversee end-to-end execution including inference optimization, distributed training infrastructure, autonomous AI workflows, and cross-team collaboration for production impact.
9+ years of software engineering experience, including 2+ years in engineering management.
Bachelor's or Master's degree in Computer Science or related field.
Proven hands-on expertise with modern LLM inference stacks and production low-latency model serving at scale.
Experience with distributed training frameworks, GPU fleet management in production (Kubernetes/GKE), and proficiency in Python and systems-level programming languages (C++/Go/Rust).
Experienced leader skilled in managing high-performing engineering teams and owning delivery end to end in AI/ML infrastructure contexts.
Technical depth in GPU-based inference optimization, distributed ML systems, and building scalable AI platforms serving millions of users.
Strong background in MLOps, LLMOps, and automating AI workflows with autonomy, able to partner across product, data science, and platform teams.