





Tier-1 brand plus mid-level manager role with AI specialization results in medium competition.
Specialized GPU inference, runtime internals, and ML optimization skills reduce cross-industry transferability.
Explicit 5+ years, 2+ leadership, and specialized GPU/ML stack requirements create strict shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and grow a team developing an on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, owning execution, technical direction, delivery quality, and roadmap.
Drive cross-functional alignment and strategy across software, research, architecture, product teams, industry partners, and open-source communities for AI ecosystem advancement on RTX and DGX platforms.
Provide technical leadership for architecture and evolution of inference runtimes and execution stacks across AI frameworks supporting diverse workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
Minimum 5 years of industry experience with at least 2 years in engineering leadership roles.
Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field.
Strong expertise in C++ software development, debugging, data structures, algorithms, and machine learning systems.
Proven experience leading engineering teams in systems software, AI infrastructure, or inference runtimes, with deep understanding of AI inference pipelines and frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT.
Experienced in building and scaling high-performing engineering teams and setting technical vision during rapid growth phases.
Strong system-level technical depth balancing architecture, performance optimization, and roadmap delivery in a fast-paced, cross-functional environment.
Background in modern machine learning, deep neural networks, generative AI, GPU programming (including CUDA), and open-source AI inference runtime contributions.