





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand and metro locations increase applicant density despite niche GPU/ML infra specialization.
Deep GPU, CUDA, distributed training and systems performance expertise required, limiting cross-industry transferability.
Explicit 12+ years and mandatory C++, Python, systems, profiling, and GPU performance expertise create stringent filters.
Analyze and optimize end-to-end performance of large-scale AI workloads spanning compute, network, storage, and software.
Establish performance baselines, diagnose bottlenecks, and design performance evaluation methodologies and benchmarks.
Collaborate with cross-functional teams to validate and implement performance improvements impacting scalable DGX Cloud systems.
Bachelor's degree or higher in Computer Science, Computer Engineering, or related field (or equivalent experience).
12+ years of professional experience with strong programming skills in C++ and Python relevant to analysis and automation workflows.
Solid understanding of operating systems, computer architecture, distributed systems, and experience in performance engineering, benchmarking, profiling, and optimization.
Work Experience Required: Minimum 12 years in relevant fields.
Experience analyzing performance on large-scale AI clusters or distributed training and inference workloads.
Proficient in GPU computing, CUDA, and GPU performance analysis techniques.
Familiarity with deep learning frameworks such as PyTorch or JAX/XLA and skilled in system-level workload characterization and optimization.