





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Mid-level Bangalore role with niche LLM inference skillset and moderate startup brand creates medium competition.
Deep CUDA/kernel, Triton, and LLM inference expertise makes skills highly domain-specific and less transferable.
Explicit 3-6 years plus mandatory CUDA, Triton, quantization, and specific serving-stack experience enforces high strictness.
Own and optimize large language model serving stack on NVIDIA GPU clusters ensuring low-latency and high-quality voice AI inference.
Drive performance engineering across serving engine, model optimization (quantization, distillation, pruning), and custom kernel development to reduce cost per million tokens.
Manage production inference operations including Kubernetes orchestration, monitoring latency metrics, load testing, capacity planning, and incident response.
3 to 6 years experience in ML infrastructure, model serving, or performance engineering; at least 2 years focused on LLM inference in production.
Hands-on experience with at least two serving engines among vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, including understanding of schedulers and KV cache designs.
Demonstrated quantization expertise (FP8, INT8, INT4) on real models with calibration and accuracy impact measurement.
Proficiency in Python, CUDA fundamentals, multi-GPU serving, Kubernetes/Docker in production, and observability tools like Prometheus/Grafana.
Deep technical expertise with at least two domains among serving stack tuning, model quantization/distillation/pruning, and custom kernel engineering.
Experienced in operating large-scale LLM inference platforms with real-time latency constraints and cost optimization focus relevant to voice AI.
Familiar with complex distributed GPU environments and proficient in profiling and eliminating performance bottlenecks.