





Known tech brand, metro Bangalore location, and senior engineer title increase competition density.
Strong bias toward AI-infrastructure and GPU/LLM expertise limits cross-industry transferability.
Requires specialized AI inference, GPU, distributed-systems experience, and specific languages and frameworks.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design, development, and delivery of high-scale, resilient data plane services for AI Inference as a Service.
Architect and optimize distributed inference hosting systems ensuring high availability and performance using techniques like tensor/data parallelism and caching optimizations.
Operate and maintain critical multi-tenant AI inference cloud services with strong focus on observability and SLO adherence.
Strong proficiency in GoLang or Python and experience with gRPC.
Experience designing and maintaining distributed systems, microservices, and high-scale customer-facing cloud services.
Hands-on experience with AI/ML inference engines for large language or multimodal models (e.g., vLLM, SGLang) and distributed inference frameworks.
Work Location: Hybrid role based in Bengaluru, India.
Work Experience Required: Not explicitly mentioned in the JD.
Technical leader with expertise at the intersection of distributed systems and AI hardware optimization for large generative AI models.
Experience working in high-scale cloud environments with customer-facing software products and strong operational discipline.
Demonstrated knowledge of LLM architectures, GPU optimizations, and inference-specific distributed serving frameworks.