





Tier-1 employer, metro location, and mid-level technical role increase applicant competition.
Highly specialized GPU and ML inference skills limit cross-industry transferability.
Explicit 5+ years plus required C++, CUDA, inference runtime and quantization expertise increases filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize high-performance local AI inference software stacks for NVIDIA RTX and DGX GPU systems, focusing on low latency, memory efficiency, and scalability across various hardware.
Design, build, and improve inference runtimes and execution stacks supporting multiple AI workloads (LLM, vision-language, TTS, ASR, diffusion) using frameworks like llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Lead end-to-end optimization, debugging, and performance tuning of AI models and inference backends, creating infrastructure for performance/accuracy analysis and establishing engineering guidelines for production readiness.
Minimum 5 years of experience in software engineering or related fields with a Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or equivalent experience.
Strong expertise in C++ programming, debugging, and computer science fundamentals (data structures, algorithms).
Proven experience developing and optimizing AI inference pipelines using frameworks such as llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.
Deep understanding of inference runtime internals including scheduling, memory management, quantization, and hardware-aware optimization techniques.
Experienced in working with GPU-based AI systems focusing on local inference optimization on resource-constrained platforms.
Has a track record of delivering scalable inference software products in collaborative, multidisciplinary engineering environments.
Proficient with low-level system and GPU programming concepts, including CUDA and high-performance system design, preferably with open-source contributions in AI runtimes or performance tooling.