





Tier-1 employer, mid-level generalist role, and metro location increase applicant competition.
Requires specialized AIOps, observability, and LLM experience, limiting cross-industry portability.
Multiple mandatory niche LLM/AIOps skills and a 3+ years requirement enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build AIOps capabilities including intelligent incident routing, anomaly detection, operational copilots, ChatOps workflows, and automated remediation.
Integrate LLM-based reasoning, build knowledge-grounding systems, and develop embeddings, retrieval systems, and intent classification for operational intelligence.
Architect automated workflows for incident triage, diagnostics, collaboration, and remediation, with real-time inference on telemetry and integration across observability and cloud platforms.
3+ years experience in system design, platform & reliability engineering, ML engineering, data engineering, or AIOps roles.
Hands-on experience building ML or LLM-based systems using Python, Java, PyTorch, TensorFlow, or modern LLM frameworks.
Experience with automation tools like StackStorm, Rundeck, Airflow, Jenkins, or cloud-native orchestration platforms.
Experience with observability platforms (Datadog, Splunk, Prometheus, Grafana, ELK) and streaming/event-driven systems (Kafka, Kinesis, Pub/Sub).
Experienced in architecting and deploying RAG pipelines, embeddings, intent models, and operational chatbots relevant to AIOps use cases.
Skilled in cloud-native systems including Kubernetes, microservices, and modern deployment patterns supporting scalable, distributed environments.
Proficient with at least one advanced AI coding assistant tool (Claude, Code Codex, GitHub Copilot) and understanding of context engineering and agentic harness frameworks.