





Tier-1 employer, mid-level (3-5yrs) role in metro Hyderabad increases applicant density.
Requires specialized on-device ML, quantization, and hardware accelerator expertise, limiting cross-industry transferability.
Explicit 3-5 years plus mandatory ML inference, C++ and hardware-acceleration skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead development and commercialization of Qualcomm AI Runtime (QAIRT) SDK on Qualcomm SoCs to enable on-device AI inferencing.
Optimize performance of large Generative AI models (LLMs, LVMs) on heterogeneous Qualcomm chipsets leveraging extensive hardware acceleration.
Deploy and maintain large C/C++ software stacks for edge-based GenAI solutions aiming for high speed and low power consumption.
Bachelor’s/Master’s/PhD in Computer Science, Engineering, or related field.
3-5 years of relevant software development experience, including AI inferencing or related domains.
Strong programming skills in C/C++, Python, and OS concepts; experience optimizing AI algorithms for hardware accelerators (CPU/GPU/NPU).
Deep understanding of Generative AI models (LLM, LVM), quantization, and floating/fixed point representations.
Experienced in developing and optimizing AI inferencing software on heterogeneous embedded systems, especially Qualcomm SoCs.
Knowledgeable about SIMD processor architecture, kernel development, and parallel computing (OpenCL, CUDA) is preferred.
Familiar with AI/ML frameworks like PyTorch, TFLite, ONNX Runtime, and GenAI model traits (self-attention, cross-attention, kv caching).