Engineer, Staff-Machine Learning-Embedded,C++
Qualcomm IncorporatedMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand and metro location but highly specialized embedded GenAI skillset reduces applicant density.
On-device GenAI, SoC optimizations and kernel work demand highly specialized semiconductor and embedded ML experience.
Requires 6+ years and specialized embedded ML, SIMD/kernel, and C++ expertise creating strict filters.
Job Description
Structured overview of role & requirementsAbout This Role
Lead development and commercialization of the Qualcomm AI Runtime (QAIRT) SDK on Qualcomm SoCs, focusing on AI inferencing performance optimization.
Deploy and optimize large-scale C/C++ software stacks for Generative AI models including LLMs and LVMs on heterogeneous Qualcomm hardware for edge inference.
Collaborate across diverse teams to implement cutting-edge GenAI advancements on-device without cloud dependency.
Minimum Requirements
Bachelor’s degree in Engineering, Computer Science or related field with 6+ years of relevant software development experience (Master’s or PhD with reduced experience acceptable).
Strong proficiency in C/C++ programming, design patterns, and operating system concepts.
Experience working with Generative AI models (LLM, LVM) and concepts like self-attention, cross attention, key-value caching, quantization, and AI hardware accelerator optimization (CPU/GPU/NPU).
Good scripting skills in Python and excellent analytical and debugging abilities.
Ideal Candidate Profile
Experienced software engineer with a strong background in embedded and heterogeneous computing environments focused on AI inferencing and deployment at the edge.
Demonstrated expertise in deploying and optimizing large C/C++ software stacks in performance-critical settings, especially with Generative AI models and AI accelerators.
Comfortable in a collaborative global environment, capable of integrating advancements in LLM/Transformer architectures and low-level system/kernel details.
