





Tier-1 brand, mid-level ML systems role, metro Pune, and broad GPU/AI skillset increase applicant density.
Strong GPU, CUDA, and inference-runtime specialization limits cross-industry transferability.
Multiple mandatory skills (C++, CUDA, TensorRT, inference runtimes) and explicit 5+ years make filters stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs focusing on performance, stability, and scalability.
Collaborate with software, research, architecture, and product teams to align AI on-device strategies and technical requirements.
Architect and develop inference runtimes and execution stacks for various AI workloads including LLMs, vision-language, TTS, ASR, and diffusion models with end-to-end model and runtime optimization.
5+ years of experience in Computer Science, Software Engineering, Mathematics, or related field with Bachelor's or higher degree or equivalent experience.
Strong C++ programming and debugging skills with solid understanding of data structures, algorithms, and machine learning.
Proven experience with AI inference pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep knowledge of inference backend internals including scheduling, memory management, quantization, and hardware-aware optimizations.
Expertise in developing high-performance AI inference runtimes and optimization for resource-constrained GPU architectures.
Experience delivering end-to-end AI software products in collaboration with distributed, multinational teams.
Background in low-level system/GPU programming (CUDA, Vulkan, DirectX) and contributions to open-source inference runtimes or performance infrastructure.