





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 employer, mid-level role and metro location increase density, but GPU/LLM specialization reduces competition.
Highly domain-specific GPU, CUDA, and inference-backend skills limit cross-industry transferability.
Explicit 4+ years requirement plus mandatory C++, CUDA, and inference/runtime expertise creates strict technical filters.
Develop and optimize high-performance local AI inference stacks for RTX, RTX Pro, and DGX GPUs addressing latency, memory efficiency, and scalability.
Architect and implement modern inference runtimes supporting frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT for diverse AI workloads including LLMs, vision-language, TTS, ASR, and diffusion.
Lead end-to-end performance optimizations, system-level debugging, and infrastructure development for performance/accuracy analysis to ensure production readiness of AI models on resource-constrained platforms.
4+ years of experience and Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
Excellent C++ programming and debugging skills, strong understanding of data structures, algorithms, and machine learning.
Proven experience with AI inferencing pipelines using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Strong knowledge of inference backends and runtime internals including scheduling, memory management, quantization, and hardware-aware optimization.
Experience contributing to or building open-source AI inference runtimes, model tooling, or performance infrastructure reflecting deep system and GPU programming skills.
Track record of delivering end-to-end local AI products in multinational, distributed engineering environments.
Strong proficiency in low-level GPU programming (CUDA), and building high-performance AI systems integrating frameworks like PyTorch, TensorRT, Vulkan, DirectX, and vLLM.