





Tier-1 employer, metro location, mid-level role but highly specialized ML kernel/GPU skillset reduces candidate pool.
Specialized ML kernel and accelerator HW-SW optimization skills limit transferability across industries.
Explicit 6–12 years and mandatory deep skills in C++, ML kernels, frameworks, and HW-SW optimization.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and develop software features for AI frameworks supporting Intel AI accelerators and GPUs, focusing on both hardware-agnostic and hardware-aware ML kernel development.
Enhance and optimize deep learning training and inference capabilities within the software stack to improve performance of AI workloads.
Engage with open-source AI communities for development, integration, and upstream adoption of software improvements.
6 to 12 years of relevant work experience in AI software development or related fields.
BTech or MS/MTech degree in Computer Science, Electronics and Communication, or related disciplines.
Proficiency in advanced C++ (C++14/17), intermediate Python skills, and experience with parallel programming.
Hands-on experience in machine learning kernel development (e.g., GEMM, Convolution, Flash Attention) and practical knowledge of deep learning models/LLMs for Vision and NLP.
Experienced contributor with deep understanding of computer architecture and HW-SW optimization techniques relevant to AI frameworks.
Proven ability to debug complex multilayer software systems and work on large production AI frameworks such as PyTorch, TensorFlow, or JAX.
Experience with CUTLASS or Triton kernel development for large language models and knowledge of compiler algorithms for heterogeneous systems.