





Tier-1 brand, remote listing and metro location increase applicants, but specialized SRE/Kubernetes/Kafka skills narrow qualified pool.
Role demands deep platform/SRE cloud and streaming expertise, limiting transferability across industries.
Explicit 10+ years and extensive mandatory platform, cloud, observability, and streaming skill requirements enforce strict shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and drive platform reliability strategy including availability, resilience, latency, and operational efficiency for enterprise integration and event-driven ecosystems.
Serve as final technical escalation point for high-severity incidents and lead post-incident reliability improvements and governance controls.
Provide technical leadership and mentorship to Tier 2 and Tier 3 engineers; lead platform tooling automation and observability stack architecture decisions.
10+ years experience in SRE, platform engineering, DevOps, or advanced production support with demonstrated technical leadership.
Hands-on expertise with Kubernetes (especially Azure Kubernetes Service), cloud-native platform operations, and CI/CD engineering with GitHub Actions at scale.
Deep observability experience with Prometheus, Grafana, AlertManager, and logging tools, plus strong Python scripting skills.
Experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, and cloud governance controls (access management, role enforcement).
Experienced senior individual contributor used to owning end-to-end platform reliability and governance at scale in a large enterprise.
Strong technical leader comfortable with mentoring engineers and driving cross-team architecture and operational decisions.
Domain expertise in cloud-native platforms, streaming event ecosystems, and incident response in production environments.