





Tier-1 employer, mid-level role in metro location attracts many qualified applicants.
Highly specialized ML-inference and GPU systems expertise limits cross-industry transferability.
Explicit years, leadership and specialized ML runtime/C++/CUDA requirements enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and grow a team to build an on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, owning execution, technical direction, delivery quality, and roadmap alignment.
Drive cross-team and industry alignment to build strategy and strengthen the AI ecosystem across NVIDIA platforms.
Provide technical leadership on inference runtimes and execution stacks, optimize AI models and pipelines for performance on current and next-gen GPU architectures, and establish team processes for debugging, performance optimization, and production readiness.
5+ years overall industry experience with 2+ years in engineering leadership.
Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field.
Proven experience leading engineering teams in systems software, AI infrastructure, or inference runtimes.
Strong technical foundation in C++, debugging, data structures, algorithms, machine learning systems, and extensive experience with AI inference pipelines and frameworks (Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT).
Experience setting technical vision and scaling high-performing engineering teams in fast-paced, rapidly growing product organizations.
Deep expertise in AI inference runtime internals including scheduling, memory management, quantization, and hardware-aware optimizations.
Background in lower-level systems or GPU programming (CUDA, high-performance systems), and contributions to open-source inference runtimes or model tooling.