Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, build, and operate large-scale distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability.
Develop reusable platform services, APIs, and automation frameworks to enable self-service and reduce operational toil across multiple engineering teams.
Provide technical leadership and mentorship while driving architecture and engineering direction across Storage, Compute, Networking, and Platform domains for improved reliability and operational efficiency.
Minimum Requirements
Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience.
10+ years of software engineering, SRE, infrastructure, or distributed-systems experience with demonstrated technical leadership.
Strong proficiency in Go, Python, or equivalent languages with experience building production-grade distributed systems, platform services, APIs, and automation.
Experience with observability and telemetry platforms (e.g., OpenTelemetry, Prometheus), SRE and infrastructure tools (e.g., Kubernetes/OpenShift, VMware) and automation tools (e.g., Terraform, Ansible).
Ideal Candidate Profile
Proven ability to own complex software/platform initiatives across multiple teams or infrastructure domains from architecture through measurable impact.
Deep expertise in distributed systems, event-driven architectures, telemetry, and high-throughput data processing technologies (e.g., Kafka, gRPC).
Experience building self-healing systems, automated remediation, or AI-driven autonomous operations on large-scale on-premises infrastructure including Storage, Compute, Networking, VMware, and Kubernetes/OpenShift.
