Gen AI Site Reliability Engineer (SRE)- Senior Associate-AI Managed Services - Consult
PwCMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Operate, monitor, and improve reliability and service health for AI platforms across AWS/Bedrock and OpenAI environments.
Lead incident triage, restoration, RCA, and problem management for AI workloads ensuring timely response and resolution.
Implement automation, alert tuning, and continuous reliability improvements while maintaining operational readiness documentation.
Minimum Requirements
4+ years experience in SRE, production support, cloud operations, or equivalent run-state engineering role.
Hands-on experience with monitoring, incident response, and cloud environment operations, preferably AWS and GenAI platforms.
Experience with observability tools (CloudWatch, Datadog, Splunk, Grafana, or OpenTelemetry) and ITIL-aligned incident/problem/change management processes.
Work location: Bangalore or Hyderabad, India (Remote).
Ideal Candidate Profile
Proficient in diagnosing complex, ambiguous production issues using logs, traces, and alerts to drive restoration or escalation.
Experience supporting cloud-based AI services, comfortable managing service health indicators and improving alert quality.
Demonstrated ability to automate routine SRE tasks and maintain detailed runbooks, supporting continuous improvement and stakeholder communication.
