





Tier-1 employer plus metro location raise competition, while specialized ML/GPU focus moderates density.
Requires GPU, CUDA, and inference-runtime expertise, making background transferability across industries low.
Explicit 2+ years plus mandatory C++, CUDA, and inference-runtime expertise enforce strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize high-performance AI inference software for RTX and DGX GPU platforms focusing on local, on-device execution.
Collaborate with cross-functional teams to build and scale inference runtimes and execution stacks supporting frameworks like Llama.cpp, PyTorch, TensorRT across AI workloads including LLMs, vision-language, TTS, ASR, and diffusion.
Lead system-level debugging, performance optimization, and model deployment readiness including model optimization techniques such as quantization, pruning, and distillation for resource-constrained devices.
2+ years of software development experience with a degree in Computer Science, Software Engineering, Mathematics, or equivalent experience.
Strong proficiency in C++ programming with solid knowledge of data structures, algorithms, and machine learning concepts.
Experience with AI inference pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, or TensorRT.
Deep understanding of inference backends and runtime internals including scheduling, memory management, graph execution, quantization, and hardware-aware optimization.
Experience delivering end-to-end AI inference software products in a multinational or distributed team environment.
Hands-on expertise in system-level or GPU programming including CUDA and performance optimization of low-level systems.
Contributions to open-source AI inference runtimes, model tooling, or performance infrastructure showing practical expertise with frameworks like Llama.cpp, PyTorch, Vulkan, or DirectX.