





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 employer, Bengaluru metro, and mid-level (5+ years) requirement increase applicant competition.
Specialized LLM inference, GPU architecture, and systems programming skills reduce cross-industry transferability.
Explicit 5+ years and specialized LLM, GPU, runtime, and systems programming requirements create strict filters.
Own end-to-end production LLM inference pipeline including release engineering, capacity planning, cost optimization, and incident response.
Optimize inference performance to reduce latency and increase throughput on real production traffic patterns across heterogeneous GPU fleets.
Build benchmarking tools and improve reliability through monitoring, tracing, alerting, and incident management.
5+ years of software development experience with production deployment and operation of LLM inference services.
Strong coding skills in Python and either Go or Rust for systems-level implementation and debugging.
Experience with ML frameworks and runtimes including PyTorch, vLLM, and related tools such as SGLang or TensorRT.
Knowledge of GPU architecture and performance optimization; CUDA/kernel programming skills are a strong plus.
Experienced in AI/ML systems programming focused on performance optimization for LLM inference workloads with demonstrated impact (e.g., doubling throughput, reducing latency).
Skilled in root-cause analysis across models, runtimes, networking, and infrastructure to identify and resolve bottlenecks.
Familiarity with applying autonomous AI coding agents to accelerate development and improve code quality in software delivery pipelines.