





Tier-1 brand, mid-level 5+ years role, and metro location increase competition.
Requires specialized GPU inference and runtime expertise, limiting cross-industry transferability.
Explicit 5+ years plus mandatory CUDA, inference runtime, and model optimization skills enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own development and optimization of local AI inference stack for RTX, RTX Pro, and DGX GPUs focusing on performance, scalability, and stability.
Architect and develop modern inference runtimes and execution stacks supporting frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT across varied AI workloads.
Perform end-to-end system-level optimization, debugging, and accuracy trade-off analysis to ensure efficient deployment of large AI models on local and edge devices.
5+ years of experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills including strong knowledge of data structures, algorithms, and machine learning.
Proven experience working with AI inferencing pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep understanding of inference backends, runtime internals, scheduling, memory management, hardware-aware optimizations, and model optimization techniques.
Experience delivering end-to-end AI inference products within multinational and distributed engineering teams.
Proficiency in low-level system or GPU programming including CUDA and high-performance system development.
Contributions or hands-on work with open-source inference runtimes, model tooling, or performance infrastructure, demonstrating deep technical involvement in local AI execution.