





Tier-1 brand, remote posting, popular SRE role, metro locations, and broad skillset increase competition.
Requires deep cloud-native SRE, Kubernetes, CI/CD, and Kafka expertise, somewhat transferable but domain-specific.
Multiple mandatory platform, cloud, Kubernetes, observability, Kafka, and incident leadership requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and enhance platform reliability including availability, resilience, latency, and operational efficiency.
Lead DevOps automation initiatives, CI/CD pipeline management, and monitoring/alerting stack optimization using tools like Prometheus, Grafana, Azure Monitor, and related technologies.
Drive cloud infrastructure governance, incident troubleshooting for complex issues, capacity and disaster recovery planning, and platform best-practice documentation.
6+ years experience in SRE, platform engineering, DevOps, or advanced production support roles.
Expertise in Kubernetes (especially Azure Kubernetes Service), CI/CD with GitHub Actions, Prometheus/Grafana/AlertManager observability stacks, and Python automation scripting.
Experience with Confluent Kafka ecosystem (Confluent Kafka, Confluent Cloud), Azure Event Hub, and related streaming technologies for enterprise operations.
Location requirement: Onsite in Hyderabad or Bangalore at designated AT&T locations.
Senior to Lead Individual Contributor level with typically 10 to 17 years experience, evidencing leadership in high-severity incident management and reliability improvements.
Strong cross-functional collaboration skills working with architecture, delivery, and operational teams to influence platform standards and resilience patterns.
Experienced in production troubleshooting for Java, React, Spring Boot microservices with emphasis on operational stability, not feature development.