





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Strong Tier‑1 brand and metro locations increase competition, but niche SRE specialization limits applicants.
Core SRE skills transfer across industries, but heavy streaming, Azure, and telecom context raises domain specificity.
Explicit 6+ years plus many mandatory cloud, Kubernetes, observability, and Kafka skills increases filtering strictness.
Own platform reliability including availability, resilience, latency, and operational efficiency for event-driven ecosystems.
Lead automation initiatives including GitHub Actions CI/CD pipelines, Helm chart automation, and cloud-native platform upgrades.
Manage and optimize observability stack (Prometheus, AlertManager, Grafana, OpenSearch, etc.), troubleshooting high-complexity production incidents, and lead cloud infrastructure governance and capacity planning.
6+ years experience in SRE, platform engineering, DevOps, or advanced production support roles.
Strong hands-on expertise with Kubernetes (especially Azure Kubernetes Service), CI/CD (GitHub Actions), and observability tools (Prometheus/Grafana/AlertManager).
Hands-on experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS MSK, and Apache Flink.
Onsite work required at Hyderabad or Bangalore AT&T locations.
Senior to Lead individual contributor with 10 to 17 years overall experience in platform reliability for event-driven microservices.
Proficient in Python automation scripting for operational toil reduction and reliability engineering at enterprise scale.
Experienced in governance controls (access management, role enforcement), high-severity incident leadership, and cross-functional collaboration on reliability and scalability patterns.