





Senior, highly specialized ML runtime and accelerator role reduces applicant density despite Bangalore location.
Deep expertise in NN operators, CUDA runtimes, and custom accelerators makes skills highly industry-specific.
Explicit senior years plus mandatory CUDA, PyTorch, runtime and systems expertise create strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, optimize, and implement neural network operators for high-performance execution on next-generation AI accelerator hardware.
Develop new neural network operators using CUDA/custom runtime APIs and own runtime to NN layer interfaces and execution model.
Drive runtime-level and operator fusion optimizations to maximize performance, scalability, and hardware utilization in AI workloads.
8+ years of experience in systems software, runtime, or performance engineering.
Strong expertise with PyTorch, TensorFlow, JAX or similar frameworks and NN operator/kernel development and optimization.
Hands-on experience with C/C++ and CUDA or similar low-level programming for runtime systems.
Deep understanding of memory hierarchy, data movement, and parallel execution principles.
Proven ability to lead design and performance optimization of NN operators on AI accelerators or GPU/NPU hardware.
Experience interfacing runtime systems with neural network layers and optimizing execution models.
Background in optimizing large-scale deep learning workloads and working with compiler-runtime interactions (e.g., XLA, MLIR).