





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, metro location, mid-level role but niche LLM/GPU specialization limits applicant density.
Requires specialized LLM/GPU systems skills, making cross-industry transferability limited.
Multiple explicit years, mandatory ML inference, GPU, and systems-level skills create stringent shortlisting filters.
Own production-grade large language model (LLM) inference services, handling release engineering, capacity planning, cost optimization, and incident response.
Optimize LLM inference performance by tuning runtimes and servers to reduce latency and increase throughput on heterogeneous GPU fleets.
Develop benchmarking, monitoring, and observability tools; apply new inference optimizations and collaborate with cross-functional teams to meet performance and availability SLOs.
5+ years of strong development experience, including deploying and operating LLM inference services in production.
Proficient in Python and Go or Rust for systems-level implementation and debugging.
Experience with ML frameworks and runtimes such as PyTorch, vLLM, SGLang, and/or TensorRT.
Strong knowledge of GPU architecture and performance profiling; CUDA/kernel programming is a strong plus.
Experienced in performance optimization and systems programming specifically for AI/ML workloads with demonstrated measurable production improvements.
Skilled in root-cause analysis across model, runtime, networking, and infrastructure components to identify bottlenecks.
Comfortable applying autonomous AI coding agents with advanced prompting and human-in-the-loop code review to accelerate software delivery.