





Tier-1 employer increases interest, but niche GPU and inference runtime specialization reduces generalist competition.
Requires specialized GPU, CUDA and inference runtime expertise, limiting cross-industry transferability.
Explicit 5+ years plus mandatory C++, GPU/CUDA, and inference runtime expertise enforces strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize high-performance local AI inference software stacks for RTX, RTX Pro, and DGX GPUs, focusing on latency, memory efficiency, and scalability.
Architect and build inference runtimes supporting frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT across diverse AI workloads (LLMs, vision-language, TTS, ASR, diffusion).
Lead system-level debugging, performance optimization, and quality trade-off analyses to ensure production readiness and accelerate deployment of new AI models and backends on resource-constrained platforms.
5+ years of experience in Computer Science or related field with Bachelor's, Master's, or PhD, or equivalent experience.
Proficiency in C++ programming and debugging with strong understanding of data structures, algorithms, and machine learning.
Experience with AI inferencing pipelines and ML/DL frameworks including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep knowledge of inference backend internals such as scheduling, memory management, quantization, and hardware-aware optimizations.
Experienced in building and delivering end-to-end AI inference products in multinational or distributed team environments.
Demonstrated sophistication in system/GPU programming including CUDA and high-performance systems development.
Active contributor or strong familiarity with open-source inference runtimes, model tooling, or performance infrastructure relevant to local AI deployment.