





Tier-1 brand plus mid-level seniority and desirable AI/GPU skills create moderate applicant competition.
Role requires specialized GPU, inference runtime, and ML optimization skills that transfer poorly outside ML systems.
Explicit 5+ years requirement plus mandatory C++, CUDA, and inference/runtime expertise increases filter strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize on-device AI inference software for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability on diverse hardware.
Collaborate with software, research, architecture, and product teams to align technical requirements and advance the AI ecosystem on RTX and DGX PCs.
Design modern inference runtimes and execution stacks leveraging frameworks like llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT, including end-to-end model optimization and performance tuning.
5+ years experience in software development or related field with Bachelor’s, Master’s, or PhD in CS, Software Engineering, Mathematics, or equivalent experience.
Strong proficiency in C++ programming, debugging, data structures, algorithms, and machine learning fundamentals.
Experience developing and optimizing AI inference pipelines with ML/DL frameworks such as llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.
Deep knowledge of inference backend internals including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimizations.
Demonstrated ability to deliver end-to-end AI inference products in complex, distributed, multinational engineering environments.
Expertise in low-level system and GPU programming including CUDA and building high-performance systems.
Active contributions to open-source AI inference runtimes, model tooling, or performance infrastructure and hands-on experience with frameworks like llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM.