





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand and Bangalore metro increase density, but seniority and niche performance specialization limit competition.
Highly specialized GPU cluster and AI performance expertise limits cross-industry transferability.
Explicit 12+ years and specialized GPU, C++/Python, and systems performance skills make shortlisting highly strict.
Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks.
Design and execute performance studies to establish baselines, diagnose regressions, and identify bottlenecks, then drive optimizations through to deployment.
Collaborate with deep learning engineers, platform teams, and GPU architects to validate and deliver performance improvements and influence system/software design decisions.
Bachelor's degree or higher in Computer Science, Computer Engineering, or related field (or equivalent experience).
12+ years of experience programming in C++ and Python with ability to create reliable analysis and automation workflows.
Experience in operating systems, computer architecture, distributed systems, and performance engineering including benchmarking, profiling, and optimization of complex software or systems.
Ability to communicate technical analysis and prioritize impactful work across teams.
Experienced with large-scale AI clusters or distributed training/inference workloads.
Familiarity with CUDA, GPU computing systems, and GPU performance analysis.
Hands-on with deep learning frameworks such as PyTorch or JAX/XLA and advanced system-level workload characterization and optimization.