





Tier-1 brand, mid-level role, and metro Bangalore location increase candidate competition.
Highly specialized CPU and ML kernel skills limit cross-industry transferability.
Mandatory low-level C/C++ and CPU architecture/ML kernel skills enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Drive CPU software-hardware co-design for next-generation QMX CPU architectures focused on machine learning workloads.
Identify and characterize critical ML workloads; generate simulation traces and perform bottleneck analysis to optimize CPU performance for ML.
Develop highly optimized ML kernels and libraries (GEMM, convolution, attention) for QMX architecture and collaborate with architecture teams to influence CPU design.
Bachelor's degree in Engineering, Computer Science, or related field with 2+ years experience OR Master's with 1+ year OR PhD in relevant field.
Proficiency in C/C++ programming mandatory; experience with computer architecture, systems programming, and ML fundamentals.
Experience or knowledge of performance profiling, benchmarking, and optimization.
Work location: Bangalore or relevant location; Work Experience Required: 2+ years software engineering or related; use of simulators like QEMU preferred but not mandatory.
Experienced in low-level CPU performance optimization, kernel level tuning, and software-hardware co-design for ML workloads.
Familiar with CPU architectural features such as SIMD/vector extensions (e.g. QMX) and memory hierarchy tuning.
Capable of working across stack from ML models to architectural feedback, impacting next-gen CPU designs in collaboration with architecture teams.