





Niche LLM inference specialization with metro Bengaluru and mid-level seniority yields moderate competition.
High because role demands niche LLM serving, GPU optimization, and AI infrastructure expertise.
High due to explicit 6+ years requirement and specialized LLM inference, GPU and Kubernetes expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design, development, and deployment of production-grade, multi-tenant large language model (LLM) inference systems on DigitalOcean's GPU cloud.
Collaborate directly with AI startup founders and tech leads to troubleshoot latency, GPU utilization, and inference code optimizations in high-throughput production environments.
Own full lifecycle of distributed AI inference infrastructure including architecture, debugging, profiling, and internal tooling to reduce inference latency and improve hardware efficiency.
6+ years experience in AI/ML systems with expertise in cluster-scale AI inference serving and distributed systems.
Expert-level proficiency in Python or GoLang, Kubernetes experience, and familiarity with gRPC for running high-scale services.
Hands-on experience with inference frameworks such as vLLM, llm-d, SGLang, TensorRT-LLM, or Modular MAX, including optimization techniques like continuous batching and caching.
Location requirement: Bengaluru, India with ability to travel up to 30% and work overlapping North American business hours (availability until at least noon Eastern Time).
Experienced AI Inference architect or Forward Deployed Engineer with a strong technical consulting background supporting production AI/ML systems at scale.
Skilled at translating complex business latency SLAs into technical solutions and effectively collaborating with external engineering and leadership teams (CTOs, AI leads).
Proven ability to deliver production-ready code and deployment blueprints rather than presentations, comfortable in fast-moving, high-impact environments engaging with GPU and infrastructure vendors.