





Tier-1 brand and Bangalore metro increase applicant density, but niche ML kernel expertise limits competition.
Role needs niche ML kernel and HW-SW optimization skills, so cross-industry transferability is low.
Explicit 6-12 years plus specialized ML kernel, C++ and HW-SW optimization requirements make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and develop software features for AI frameworks that are both hardware-agnostic and hardware-aware, focusing on machine learning kernel development.
Optimize and enhance deep learning training and inference capabilities within Intel's AI software stack for data center AI accelerators and next-generation GPUs.
Engage with the open-source community by participating in development, adopting upstream software, and enabling software stack optimization for deep learning workloads.
6 to 12 years of relevant experience with a BTech or MS/MTech in CS, ECE, or related fields.
Proficiency in Advanced C++ (C++14/17) and intermediate Python skills with knowledge of parallel programming.
Hands-on experience with at least one major AI framework such as PyTorch, TensorFlow, or JAX, including development of ML kernels (e.g., GEMM, Convolution, Flash attention).
This role requires on-site presence in Bangalore, India, and knowledge of deep learning models for vision and NLP tasks.
Experienced in working on AI frameworks or platforms that have reached production environment, demonstrating capability to debug complex multilayered software systems.
Strong understanding of computer architecture and hardware-software optimization techniques specific to AI workload acceleration.
Familiarity with open-source community engagement and kernel integration in large language models, with preferable experience in CUTLASS or Triton kernel development and compiler optimization algorithms for heterogeneous systems.