





Tier-1 brand and metro location increase applicants, but niche GPU/performance focus narrows the pool.
High because GPU deep learning performance skills are specialized and not broadly transferable across industries.
Moderate strictness due to explicit language requirements and preferred GPU/performance expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and innovate hardware architectures focused on improving parallel computing performance and energy efficiency in AI workloads.
Benchmark, profile, and analyze AI workloads in both single-node and multi-node configurations to optimize performance.
Collaborate closely with architecture teams and product management to influence product development and stay updated on deep learning trends.
B.Tech. or M.Tech. degree in Computer Science, Electrical Engineering, Mathematics, or relevant discipline.
1+ years of experience programming in C, C++, and Python.
Experience or familiarity with parallel computing and GPU-related technologies is relevant but not explicitly stated as mandatory.
Notice period or location constraints: Not explicitly mentioned in the JD.
Experienced in parallel programming and GPU computing environments, able to benchmark and analyze complex AI workloads effectively.
Skilled at working collaboratively with cross-functional teams including architecture and product management to drive product improvements.
Has familiarity with transformer-based model architectures and tools for architecture simulation, performance modeling, and profiling.