





Large enterprise brand and Bangalore location increase applicant density despite specialized agentic-AI SRE skills.
Requires specialized SRE and agentic-AI observability experience, limiting cross-industry transferability.
Mandates specific observability, LangChain/LangSmith/Galileo, and production tier‑3 incident-response experience.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Deploy autonomous AI agents capable of independent planning, execution, and self-healing of infrastructure.
Provide tier-3 production support for critical cloud infrastructure and AI agent pipelines, minimizing downtime.
Lead incident management including live outage troubleshooting, high-priority incident handling, root-cause analysis, and automated post-mortem generation using LLM agents.
Strong experience with observability tools for multi-agent AI systems such as LangChain, LangGraph, LangSmith.
Expertise managing Loki (logs), Grafana (dashboards), Tempo (traces), and Mimir/Prometheus (metrics).
Experience with Galileo and LangSmith for monitoring prompt traces, token costs, latency, and evaluating RAG accuracy, hallucination, and safety.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in operating and supporting autonomous AI agent systems in production environments.
Proficient with multi-tool observability stacks integrating tracing, logging, metrics, and continuous evaluation of AI model outputs.
Capable of managing high-severity incidents involving AI-driven infrastructure and leveraging automated diagnostics and chaos-testing.