Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)
AT&T Inc.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 employer, metro locations, and senior SRE role increase candidate competition.
Core SRE and cloud skills are transferable, but Kafka and telecom specifics raise domain bias.
Multiple mandatory specialized skills, explicit years, and high-severity incident leadership mandate strict filtering.
Job Description
Structured overview of role & requirementsAbout This Role
Own platform reliability focused on availability, resilience, latency, and operational efficiency across cloud and streaming platforms.
Lead automation initiatives including CI/CD pipeline management, Helm chart automation, and platform/tooling upgrades to enhance deployment and operational processes.
Manage monitoring, alerting, incident troubleshooting, cloud infrastructure governance, capacity and disaster recovery planning, and documentation for high-severity incidents.
Minimum Requirements
6+ years in SRE, platform engineering, DevOps, or advanced production support roles.
Strong hands-on experience with Kubernetes (especially Azure Kubernetes Service), cloud-native platforms, CI/CD engineering with GitHub Actions, and observability tools like Prometheus, Grafana, and AlertManager.
Proficient in Python automation scripting, incident leadership, and production troubleshooting for Java, React, Spring Boot services.
Onsite work required in Hyderabad or Bangalore (designated AT&T locations).
Ideal Candidate Profile
Experienced technologist comfortable operating at enterprise scale with cross-functional influence on reliability standards.
Strong domain knowledge of streaming platforms including Confluent Kafka, Confluent Cloud, Azure Event Hub, and AWS MSK with governance experience.
Skilled in operational troubleshooting rather than feature development, with a focus on automation to reduce toil and accelerate incident resolution.
