





Tier-1 brand and mid-level title increase applicants; GPU/inference specialization limits overall competition.
Strong GPU, CUDA, and ML-inference runtime requirements restrict cross-industry transferability.
Explicit 5+ years plus mandatory CUDA, C++, and inference-runtime expertise makes filters highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own development and optimization of high-performance local AI inference stacks for RTX and DGX GPUs, focusing on performance, stability, and scalability across architectures.
Design and implement inference runtimes and execution stacks supporting multiple ML frameworks (e.g., Llama.cpp, vLLM, PyTorch, TensorRT) for various AI workloads including LLMs, vision-language, TTS, ASR, and diffusion.
Lead system-level debugging, performance tuning, and develop infrastructure for performance and accuracy analysis to ensure production readiness on resource-constrained devices.
5+ years of professional experience in software engineering or equivalent with a Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field.
Proficient in C++ programming with strong debugging skills and sound understanding of data structures, algorithms, and machine learning concepts.
Proven experience working with AI inference pipelines using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep knowledge of inference backends and runtime internals including scheduling, memory management, quantization, and hardware-aware optimizations.
Experienced working in multinational or geographically distributed teams delivering end-to-end AI software products.
Strong systems programming skills including GPU programming using CUDA, Vulkan, DirectX and developing high-performance systems.
Demonstrated contributions to open-source projects related to ML inference runtimes, model tooling, or performance infrastructure, evidencing expertise in modern ML and generative AI techniques.