





Tier-1 brand, metro location, and mid-level experience make applicant competition high.
Requires deep ML inference and GPU systems expertise, making cross-industry transfers difficult.
Explicit years, leadership, and specialized inference/runtime/C++ requirements create high shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and grow engineering team to develop an on-device AI inference platform targeting RTX, RTX Pro, and DGX GPUs, accountable for execution, technical direction, delivery quality, and roadmap alignment.
Drive cross-functional collaboration with internal teams and external partners to establish strategy and strengthen AI ecosystem across relevant platforms.
Provide technical leadership on inference runtimes and execution stacks across multiple AI frameworks; optimize AI models and runtimes for performance, memory efficiency, and deployment on resource-constrained devices.
Minimum 5 years of industry experience with at least 2 years in engineering leadership roles.
Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field required.
Strong proficiency in C++ development, debugging, data structures, algorithms, and machine learning systems.
Extensive experience with AI inference pipelines and deep learning frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Proven experience building and scaling high-performing engineering teams focused on systems software or AI infrastructure.
Deep understanding of inference backend internals including scheduling, memory management, quantization, and hardware-aware optimization to improve local AI model deployment.
Experience working cross-functionally in multinational product organizations and contributing to open-source AI inference runtimes or performance tooling.