





Tier-1 brand and metro location, but niche GPU inference specialization reduces applicant density.
Role requires GPU, CUDA, and inference runtime expertise, limiting transferability across industries.
Explicit minimum experience plus mandatory C++, GPU/CUDA, and inference runtime expertise create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and optimize high-performance local AI inference stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Develop and architect modern AI inference runtimes and execution stacks supporting frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT across multiple AI workloads.
Perform system-level debugging, performance and accuracy optimization, and establish engineering guidelines for production readiness of new models and inference backends.
2+ years of experience with Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills; strong understanding of data structures, algorithms, and machine learning.
Proven experience with AI inference pipelines and ML/DL frameworks including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Not explicitly mentioned: notice period or specific location requirements.
Experienced in building and optimizing AI inference backends and runtimes with deep knowledge of scheduling, memory management, quantization, and hardware-aware optimization.
Has contributed to or has strong familiarity with modern machine learning techniques and large open-source AI projects related to inference runtimes or model tooling.
Comfortable working in cross-functional teams with software, research, architecture, and product groups, delivering scalable AI solutions for resource-constrained GPU platforms.