





Strong Tier-1 brand and metro location with mid-level seniority increase competition moderately.
Highly domain-specific GPU, CUDA, and inference runtime expertise reduces cross-industry transferability.
Mandatory 5+ years plus C++, CUDA, and inference/runtime expertise enforces strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize local AI inference software stacks for RTX, RTX Pro, and DGX GPUs focusing on high performance, low latency, memory efficiency, and stability.
Design and implement inference runtimes and execution stacks supporting various AI workloads (LLM, vision-language, TTS, ASR, diffusion) using frameworks like llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, TensorRT-RTX.
Perform system-level debugging, performance tuning, model optimization (quantization, pruning, sparsity, distillation), and develop infrastructure for accuracy and performance evaluation to ensure production readiness on constrained platforms.
5+ years of experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills with strong foundation in data structures, algorithms, and machine learning.
Proven experience with AI inference pipelines and ML/DL frameworks such as llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.
Deep understanding of inference backend internals including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
Experienced in delivering end-to-end AI inference products at scale, preferably within large multinational technology companies with distributed teams.
Skilled in low-level system and GPU programming including CUDA and development of high-performance systems.
Has contributed to open-source inference runtimes, model tooling, or performance infrastructure, and is familiar with frameworks/APIs like Vulkan, DirectX, and TensorRT.