





Metro location and broad infra requirements increase competition, but specialized LLM skills narrow applicant pool.
Advanced ML and deployment skills are transferable but require niche LLM experience, so medium sensitivity.
Explicit 7+ years and mandatory Transformer, distributed training, and deployment expertise create rigid shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead end-to-end development, optimization, and deployment of Transformer-based NLP models for real-world applications with a focus on language models like BERT, GPT, and T5.
Design and manage scalable, low-latency inference pipelines achieving 10,000+ requests per second using distributed training and serving technologies such as Kubernetes, Triton, and multi-GPU clusters.
Own ML infrastructure including scalable pipelines (feature store, model registry, CI/CD), containerization, microservices API design, and observability for model performance and zero-downtime updates.
7+ years of professional experience in NLP with at least 3 years dedicated to Transformer models (BERT, RoBERTa, GPT).
Proficient in PyTorch, TensorFlow, Hugging Face transformers, and distributed training frameworks such as DDP, DeepSpeed.
Experience deploying low-latency model APIs with Kubernetes, Docker, FastAPI, and inference servers like Triton.
Location requirement: Gurugram, India (Full-Time, Onsite)
Experienced in architecting and operating large-scale, production-grade NLP models and inference pipelines with focus on throughput and latency optimization.
Skilled in distributed systems, hybrid cloud strategies, and container orchestration to build cost-efficient and highly available ML deployments.
Capable of leading technical mentorship and cross-functional collaboration to translate business needs into state-of-the-art NLP solutions with measurable impact.