





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, mid-level experience band, metro location, and visible engineering role increases applicant competition.
Role demands GPU, low-level C++ and inference-runtime expertise, making cross-industry transfers difficult.
Explicit 4+ years plus specialist C++, inference runtime, and GPU/CUDA skill requirements tighten shortlisting.
Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs prioritizing performance, stability, and scalability.
Develop modern inference runtimes and execution stacks for frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX across diverse AI workloads.
Perform end-to-end optimization of AI models, data pipelines, and inference runtimes; apply model optimization techniques to enable efficient local and edge deployment.
4+ years experience with a Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or related field (or equivalent experience).
Excellent C++ programming and debugging skills with strong understanding of data structures, algorithms, and machine learning.
Proven experience with AI inference pipelines and ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
Deep knowledge of inference backends and runtime internals including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
Has strong integration experience working cross-functionally with software, research, architecture, and product teams in a complex AI hardware/software environment.
Experienced delivering end-to-end local AI inference products with strong domain knowledge of GPU systems and ML frameworks.
Demonstrates proficiency in system-level debugging and performance analysis, with a focus on production readiness and scalability of AI inference on resource-constrained platforms.