





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Remote role plus 4–6 years mid-level demand yields medium competition for specialized LLM talent.
Requires specialized LLM production, vector DB and model-serving skills, so background fit sensitivity is high.
Explicit 4–6 years, mandated LLM production experience and specific tech stack create high shortlisting strictness.
Design, develop and deploy production large language model (LLM) applications including retrieval augmented generation (RAG), multi-step reasoning, and tool-calling systems.
Engineer and optimize backend systems involving vector databases, embedding selection, indexing, filtering, and latency/cost reduction techniques for LLM inference at scale.
Lead AI system evaluation, monitoring (accuracy, groundedness, safety, hallucination), and operate end-to-end model deployment with CI/CD in cloud environments, including mentoring engineers and driving technical standards.
4-6 years engineering experience including at least 2-3 years building and shipping production LLM or Generative AI systems.
Expertise in Python with asynchronous programming and FastAPI; deep experience with LLM application development and transformer internals.
Hands-on experience with vector databases (e.g., Pinecone, Weaviate), LangChain or similar agent frameworks, PyTorch, Hugging Face Transformers, cloud platforms (AWS/Azure/GCP), containers (Docker) and CI/CD.
Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or related field.
Demonstrates strong engineering judgement balancing accuracy, cost, and latency trade-offs in production AI systems.
Experience leading technically complex AI projects, mentoring teams, and improving engineering standards in fast-paced environments.
Pragmatic problem solver comfortable with ambiguous, rapidly evolving AI research and production challenges, with a portfolio or open-source contributions in Generative AI.