





Tier-1 brand, mid-level posting, and metro location increase competition despite specialized GPU inference requirements.
Requires niche GPU and ML inference systems expertise, limiting cross-industry transferability.
Mandatory 5+ years plus C++, CUDA, and inference runtime expertise creates stringent screening.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize local AI inference software stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Architect and build modern inference runtimes covering frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT targeting multiple AI workloads (LLMs, vision-language, TTS, ASR, diffusion).
Lead system-level debugging, performance optimization, model optimization (quantization, pruning, sparsity, distillation), and performance–accuracy trade-off analysis to ensure production readiness on resource-constrained platforms.
5+ years of professional experience in software engineering or related fields with a Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics or equivalent.
Strong C++ programming and debugging skills with deep understanding of data structures, algorithms, and machine learning.
Proven experience working with AI inference pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Expertise in inference backends and runtime internals including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
Experienced in building and delivering high-performance local AI inference systems on GPU platforms with a focus on low latency and efficient memory use.
Familiar with end-to-end product development in multinational or geographically distributed engineering teams.
Technical proficiency in GPU/system programming including CUDA and contributions to open-source inference runtimes or model tooling is advantageous.