





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Mid-senior Bengaluru role with moderate startup brand and in-demand LLM/kernel skills draws medium competition.
Niche LLM inference, CUDA/Triton kernel, and MoE expertise make the role hard to transfer across industries.
Explicit 6–10 years plus mandatory LLM inference, CUDA/Triton kernel, and MoE experience enforces high shortlisting strictness.
Own and optimize large language model serving stack and infrastructure across vLLM, SGLang, and NVIDIA Dynamo for real-time voice AI under strict latency and cost constraints.
Lead model optimization efforts including quantization, distillation, pruning, and speculative decoding to maintain accuracy while reducing inference cost.
Profile and tune low-level CUDA/Triton kernels and maintain production-grade inference platform with scalable deployments, on-call responsibility, and continuous evaluation.
6 to 10 years total experience, with 3+ years owning large-scale LLM or speech inference production systems.
Proven experience owning end-to-end inference platform including architecture, SLOs, capacity, cost management, deployment safety, and on-call support.
Proficient in writing or substantially tuning custom CUDA or Triton kernels deployed in production.
Experience with MoE model serving at scale, expert parallelism, routing, and distributed multi-node system profiling and debugging.
Strong expertise in deep profiling and performance optimization of GPU inference workloads in production, with demonstrated record fixing utilization issues.
Familiarity with managing and implementing multi-layer inference optimizations, balancing latency, accuracy, and cost under real-time constraints.
Experience contributing to or leading technical direction for open-source large model serving or kernel projects (e.g. vLLM, SGLang) or equivalent highly technical platforms.