





Tier-1 brand, mid-level (5+ yrs), metro location increase competition.
High: requires specialized GPU, CUDA, and on-device inference expertise limiting cross-industry transferability.
Explicit 5+ years and specialized GPU, CUDA, TensorRT, inference runtime skills required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Architect and develop modern AI inference runtimes and execution stacks across multiple frameworks and AI workloads.
Perform end-to-end optimization, system-level debugging, and develop infrastructure for performance and accuracy analysis to ensure production readiness of AI models and inference backends.
5+ years of experience with Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field or equivalent experience.
Proficient in C++ programming and debugging with solid understanding of data structures, algorithms, and machine learning.
Experience working with AI inferencing pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep knowledge of inference backend internals including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
Experienced in delivering end-to-end AI inference products with distributed multinational teams in product companies.
Strong background in lower-level system/GPU programming and CUDA for high-performance system development.
Contributed to open-source inference runtimes, model tooling, or performance infrastructure and hands-on experience with AI frameworks and APIs such as Llama.cpp, PyTorch, TensorRT, Vulkan, and DirectX, vLLM.