





Popular AI role, metro Bengaluru location, and broad technical requirements increase candidate density.
Role requires specialized LLM, quantization, and air-gapped deployment expertise, limiting cross-industry transferability.
Many mandatory LLM, quantization, and offline deployment skills create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, train, optimize, and deploy Large Language Models (LLMs) with end-to-end responsibility including data preprocessing, fine-tuning, model quantization, evaluation, and deployment in offline/on-prem environments.
Develop and maintain Retrieval-Augmented Generation (RAG) pipelines and agentic workflows integrating MCP servers for context-aware AI systems and knowledge retrieval using vector stores and embedding techniques.
Build and manage backend and DevOps solutions for ML workflows, including GPU optimization, local CI/CD, containerization, and working with relational and vector databases in secure/offline settings.
Strong expertise in Python and ML frameworks such as PyTorch or TensorFlow; hands-on experience with Hugging Face Transformers and Datasets.
Proven experience in training, fine-tuning (SFT), and optimization of LLMs including quantization techniques (GGUF, GPTQ, AWQ).
Experience deploying models in offline, on-prem, and air-gapped environments along with GPU optimization and CUDA stack knowledge.
Work Experience Required: Not explicitly mentioned in the JD.
Candidate with deep technical expertise in building scalable LLM-based AI systems and RAG pipelines using vector stores like FAISS, Chroma, Weaviate, or pgvector.
Experience integrating and maintaining MCP (Model Context Protocol) servers to enable dynamic interaction between LLMs and external tools or APIs in secure/offline environments.
Developer experienced in backend Python services and DevOps workflows tailored to ML training and inference, familiar with Docker, Git, Azure DevOps, and managing offline CI/CD pipelines.