





Tier-1 brand, mid-level managerial role, and metro location increase applicant competition despite niche specialization.
Requires deep GPU, inference runtime, and ML systems expertise, limiting cross-industry transferability.
Explicit 5+ years and 2+ leadership plus mandatory deep inference, C++, and CUDA expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and grow a team responsible for the on-device AI inference platform supporting RTX, RTX Pro, and DGX-class GPUs, managing execution, technical direction, and delivery quality.
Drive cross-functional alignment with software, research, architecture, product teams, industry partners, and open-source communities to build AI ecosystem strategy across RTX and DGX platforms.
Provide technical leadership for architecture and evolution of inference runtimes across multiple frameworks and AI workloads, including optimization and deployment on resource-constrained devices.
Minimum 5+ years of industry experience with at least 2 years in engineering leadership roles.
Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field.
Strong technical expertise in C++ software development, debugging, data structures, algorithms, machine learning systems, and AI inference pipelines involving frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT.
Experience with inference runtime internals such as scheduling, memory management, quantization, hardware-aware optimization; and strong cross-team collaboration skills.
Experienced leader of high-performing teams in systems software or AI infrastructure with a proven record of technical vision and delivering scalable AI inference platforms.
Strong understanding of modern machine learning and deep generative AI technologies, with practical experience contributing to open-source inference runtimes or tooling.
Technical depth in GPU programming or lower-level systems (e.g., CUDA, Vulkan, DirectX) with demonstrated ability to manage fast-paced, complex projects involving distributed teams.