





Tier-1 brand, Bangalore location, popular ML domain but specialized kernel skills reduce competition.
Highly specialized ML kernel and HW-SW optimization work limits transferability across industries.
Explicit 6–12 years plus mandatory ML kernel, advanced C++ and HW-SW optimization skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and develop software features for AI frameworks focusing on ML kernels and both hardware-agnostic and hardware-aware implementations.
Enhance and extend deep learning training and inference capabilities within Intel's AI software stack, optimizing performance of deep learning workloads.
Engage with open-source AI communities, contribute to development, and manage upstream adoption in Intel's AI frameworks.
6 to 12 years of overall work experience in relevant fields.
BTech or MS/MTech degree in Computer Science, Electronics and Communication Engineering, or related fields.
Proficiency in advanced C++ (C++14/17), intermediate Python, and parallel programming.
Hands-on experience in ML kernel development (e.g., GEMM, Convolution, Flash attention) and in at least one AI framework like PyTorch, Tensorflow, or JAX.
Experienced in debugging complex multi-layered software systems and integrating software in large open-source AI frameworks.
Strong understanding of computer architecture and hardware-software optimization techniques relevant to AI accelerators and GPUs.
Experience working in cross-geographical teams and production-level AI framework/platform development, preferably with knowledge of CUTLASS or Triton kernels and compiler optimizations for heterogeneous systems.