Software Engineer | Solution Validation, Benchmarking & Release Engineer | 10+ years
CiscoMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 employer and metro location increase applicant competition despite niche AI-cluster skills.
Role requires specialized GPU, NCCL, and AI-cluster validation expertise, limiting cross-industry transferability.
Explicit seniority and mandatory specialized GPU, benchmarking, and validation skills make shortlisting highly strict.
Job Description
Structured overview of role & requirementsAbout This Role
Own end-to-end validation, testing, benchmarking, and release qualification of AI cluster and Cisco network-switch solutions including GPU debugging and performance profiling.
Develop and execute test plans, automation, and acceptance criteria for AI clusters involving CVIS/NVIS, NCCL, resilience, and interoperability test suites.
Drive defect triage, evidence collection, RA certification, and packaging of acceptance reports to ensure solution readiness and release gate qualification.
Minimum Requirements
Bachelor's degree with 8+ years or Master's degree with 6+ years of related experience.
Experience building and executing validation for GPU or distributed AI systems, including test strategy, automation (Python), Linux, performance analysis, and defect triage.
Ability to translate system requirements into measurable tests, acceptance criteria, and release gates.
Experience with technical writing and cross-functional communication.
Ideal Candidate Profile
Experienced in AI infrastructure validation with hands-on skills in CVIS/NVIS, NVIDIA GPU platforms, and NCCL preferred.
Proficient in workload modeling, benchmark engineering, observability, and CI-based testing within complex networked system environments.
Capable of handling multi-plane, ToR, compute, storage, and software integration validation, including Cisco networking and AI cluster reference architectures.
