





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 employer, metro location, mid-level role with broad specialized skills increases applicant competition.
Specialized LLM inference, GPU, and MLOps expertise limits cross-industry transferability.
Multiple mandatory technical skills, explicit years, and specialized LLM/GPU requirements make filtering stringent.
Own production-grade LLM inference services including release engineering, capacity planning, cost management, and incident response.
Optimize inference performance to reduce latency and increase throughput across real production traffic patterns on heterogeneous GPU fleets.
Build benchmarking tools and improve reliability, observability, and monitoring of inference systems while shipping research-driven optimizations.
5+ years of development experience including deploying and operating LLM inference services in production.
Proficiency in Python and either Go or Rust for systems-level programming and debugging.
Experience with ML frameworks/runtimes such as PyTorch, vLLM, SGLang, TensorRT.
3+ years hands-on experience in performance optimization and systems programming for AI/ML workloads.
Experienced in GPU architecture and profiling, familiarity with CUDA/kernel programming considered a strong plus.
Proven ability to deliver measurable production improvements like doubling throughput and lowering latency.
Skilled in root cause analysis across model, runtime, networking, and infrastructure layers and applying autonomous AI coding agents to accelerate delivery.