





Medium — Tier-1 employer and Bangalore metro increase density, while niche ML/kernel specialization reduces it.
High — deep CPU architecture and ML-kernel skills are highly domain-specific and not easily transferable.
High — explicit years requirement, mandatory C/C++, and specialized CPU/ML kernel expertise required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own CPU software–hardware co-design for next-generation QMX architectures focusing on ML workload characterization, simulation, kernel optimization, and architectural insights.
Develop highly optimized ML kernels and libraries (e.g., GEMM, convolution, attention) integrated with open-source ML frameworks targeting CPU platforms.
Drive measurable ML workload performance improvements via bottleneck analysis, benchmarking, and collaboration with CPU architecture teams to influence future CPU features.
Bachelor's in Engineering, Computer Science, Information Systems, or related field with 4+ years experience OR Master's with 3+ years OR PhD with 2+ years in Software Engineering or related fields.
Mandatory proficiency in C/C++ programming.
Experience with performance profiling, benchmarking, and optimization.
Work Experience Required: 2+ years experience specifically with programming languages such as C, C++, Java, Python.
Experienced in CPU architecture, systems programming, and machine learning fundamentals with focus on ML workload execution on CPUs.
Skilled in simulation and trace generation using tools like QEMU or equivalent for architectural exploration.
Capable of software–hardware co-design collaboration driving architectural enhancements and optimization of ML kernels on CPU architectures including SIMD/vector extensions (NEON, SVE, QMX).