





Senior, niche AI-ops SRE role in Hyderabad with specialized skills reduces candidate pool despite metro location.
Highly specialized AI-ops/SRE skillset with LLM operational experience limits transferability across industries.
Explicit 8+ years, senior SRE and AI-ops experience, cloud, Kubernetes, IaC, and SRE fundamentals required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end reliability and availability (99.9%+) of production AI systems including LLM services, RAG pipelines, agentic workflows, and ML inference platforms.
Design, implement and operate high-availability AI infrastructure (cloud-native, Kubernetes, Terraform) with observability, automation, and incident response capabilities.
Lead incident management for AI outages, maintain operational documentation, and partner with AI engineers to ensure production-readiness and continuous improvement.
8+ years professional experience in software engineering, DevOps, or site reliability engineering; 1+ year operating AI/ML systems in production.
Bachelor's degree in Computer Science, Software Engineering, or related field (or equivalent experience).
Strong expertise in SRE principles, cloud platforms (GCP, AWS), Kubernetes, Terraform, CI/CD pipelines, and programming in Python plus a systems language.
Experience maintaining high-availability production systems (99.9%+ uptime), incident leadership, integration with observability tools, and operational knowledge of LLM providers (GCP Vertex AI, OpenAI).
Experienced in large scale AI/ML production systems reliability with a focus on AI-specific operational challenges (model drift, hallucination spikes, embedding pipelines).
Skilled in designing and operating scalable, automated cloud-native infrastructure and observability for AI workloads with strong incident response capabilities.
Comfortable collaborating with AI engineering teams to embed reliability practices and improve operational resilience continuously.