





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, mid-level role, and Pune metro location produce high competition despite technical specialization.
Role demands GPU and inference-runtime expertise, making cross-industry transfers difficult.
Requires explicit C++/CUDA, ML inference runtime, and production optimization experience, driving high shortlisting strictness.
Build and optimize local AI inference software stack for RTX and DGX GPU systems focusing on performance, stability, and scalability across hardware architectures.
Architect and develop inference runtimes supporting frameworks like Llama.cpp, PyTorch, WinML, DXCGC, and TensorRT for workloads including LLMs, vision-language, TTS, ASR, and diffusion AI.
Perform end-to-end optimization and system-level debugging of AI models and pipelines including model optimization techniques (quantization, pruning, sparsity, distillation) for resource-constrained platforms.
2+ years experience with Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills with strong understanding of data structures, algorithms, and machine learning.
Proven experience working with AI inferencing pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Notice period: Not explicitly mentioned in the JD.
Deep expertise and interest in inference backends, runtime internals, scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
Experience delivering end-to-end AI products in multinational, distributed teams and contributing to major open-source projects related to inference runtimes or model tooling.
Proficiency in low-level system/GPU programming (CUDA), high-performance systems development, and building applications with Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM frameworks.