





Strong Tier-1 brand and metro location increase competition, balanced by niche GPU/ML/compiler specialization.
Highly specialized GPU, compiler and ML systems expertise makes skills less transferable across industries.
Explicit degree-and-years minima plus specialized GPU, compiler and ML requirements make screening highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Optimize and accelerate AI workloads on GPU platforms focusing on performance, memory efficiency, and scalability, including LLMs and computer vision pipelines.
Analyze GPU architecture/microarchitecture to identify performance bottlenecks and propose optimizations across the AI stack from model graphs to GPU kernel level.
Contribute to architectural trade-offs and influence GPU roadmap with data-driven insights by collaborating with hardware, software, and ML teams.
Bachelor's degree in Engineering, Computer Science, or related field with 6+ years of Systems Engineering or related experience; OR Master's degree with 5+ years; OR PhD with 4+ years.
Strong understanding of CPU/GPU architecture.
Proficient in Python and C++ programming; experience with ML frameworks such as TensorFlow or PyTorch.
Work Experience Required: 4+ to 6+ years depending on degree qualification as stated; Notice period: Not explicitly mentioned in the JD.
Experienced in performance optimization of AI workloads on GPU platforms, with hands-on skills from kernel to model level integration.
Able to work cross-functionally with multiple teams including hardware, software, and machine learning experts to solve complex bottlenecks.
Familiar with AI inference runtimes, deployment stacks, and modern AI model architectures like transformers and MoE (Mixture of Experts).