





Tier-1 brand and metro location, but niche ML inference/GPU specialization reduces candidate density.
Requires deep GPU, runtime, and ML inference expertise, limiting cross-industry transferability.
Explicit 2+ years, mandatory C++, ML inference and GPU/CUDA skills create stringent screening.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize local AI inference software stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability across hardware architectures.
Architect and build inference runtimes covering frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT targeting LLMs, vision-language, TTS, ASR, and diffusion workloads.
Conduct end-to-end optimization of AI models, pipelines, and runtimes including quantization and pruning; perform system-level debugging, performance–accuracy trade-offs, and contribute to production readiness onboarding of new models and backends.
Minimum 2 years experience in relevant fields such as Computer Science, Software Engineering, Mathematics or equivalent experience.
Strong proficiency in C++ programming including debugging and understanding of data structures and algorithms.
Experience with AI inferencing pipelines and frameworks including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, or TensorRT.
Deep understanding of inference runtime internals such as scheduling, memory management, quantization, and hardware optimization.
Has hands-on knowledge or contributions related to machine learning inference runtimes, model tooling, or performance infrastructure in open-source projects.
Experience delivering end-to-end AI software products within multinational, distributed engineering teams.
Skills in lower-level system/GPU programming, CUDA, Vulkan, DirectX, and building high-performance AI applications using modern ML frameworks.