





Strong Tier-1 brand, mid-level seniority, and metro location increase candidate competition.
Highly specialized GPU, inference runtime, and ML optimization skills limit cross-industry transferability.
Explicit 5+ years plus required C++, CUDA, and specialized inference/runtime experience makes filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize local AI inference software stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Design and build modern inference runtimes supporting diverse AI workloads using frameworks like llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX.
Lead end-to-end AI model optimization, system debugging, performance tuning, and readiness for production deployment on resource-constrained platforms.
5+ years experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Strong expertise in C++ programming and debugging with solid understanding of data structures, algorithms, and machine learning principles.
Proven experience developing and optimizing AI inference pipelines using ML/DL frameworks such as llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.
In-depth knowledge of inference backends and runtime internals including scheduling, memory management, quantization, and hardware-aware optimization.
Experienced in delivering complex, high-performance system software in multinational or fast-paced environments with distributed teams.
Strong background in GPU programming, CUDA, and developing efficient inference runtimes and AI model tooling.
Demonstrated ability to collaborate across software, research, architecture, and product teams to align technical requirements and strategic priorities.