





Tier-1 employer but specialized GPU inference skills reduce applicant density despite mid-level experience and metro location.
Highly domain-specific GPU and inference runtime expertise limits easy transfer across industries.
Explicit 5+ years plus mandatory C++, GPU programming, and inference-runtime expertise create stringent filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own development and optimization of local AI inference stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Collaborate cross-functionally with software, research, architecture, and product teams to align AI strategies and technical needs on local execution.
Architect and develop inference runtimes covering multiple AI frameworks and workloads and lead system-level debugging, performance optimization, and production readiness.
Minimum 5 years experience with a Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills; strong understanding of data structures, algorithms, and machine learning.
Proven experience with AI inferencing pipelines and ML/DL frameworks including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT.
Deep knowledge of inference backends and runtime internals such as scheduling, memory management, quantization, and hardware-aware optimization.
Experienced in end-to-end product delivery within multinational, geographically distributed teams.
Familiarity with low-level system/GPU programming including CUDA and high-performance systems development.
Contributions to open-source inference runtimes, model tooling, or performance infrastructure and hands-on work with AI and graphics frameworks like Llama.cpp, PyTorch, TensorRT, Vulkan, and DirectX.