





Senior niche SRE role with Kubernetes requirement reduces applicant density despite metro location.
Strong SRE, Kubernetes, and cloud focus make skills less transferable across non-cloud industries.
Explicit 10+ years, Kubernetes must-have and SRE/observability requirements create strict screening.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own reliability and health of cloud-based applications by designing monitoring signals and building scalable, efficient systems across Kubernetes and cloud infrastructure.
Participate in on-call rotations to diagnose and resolve production issues, including incident investigations and root cause analysis with remediation.
Collaborate with product engineering teams to define non-functional requirements and drive best practices in reliability, observability, and automation across the engineering organization.
10+ years of hands-on experience in site reliability or related engineering roles.
Strong hands-on expertise with Kubernetes as a platform and system.
Practical experience defining and monitoring SLI/SLO/error budgets on real production systems.
Experience with cloud platforms (Azure/AWS/GCP), cloud networking fundamentals, and modern observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch).
Experience with CI/CD systems such as GitHub Actions, TeamCity, Azure DevOps, or GitLab CI.
Hands-on experience building web applications on .NET, Python (Flask, FastAPI), or Java (Spring) deployed at scale.
Experience using AI-assisted engineering tools for automation, root cause analysis, and incident management.
Extensive experience operating Kubernetes-based infrastructure and supporting distributed systems with real incident handling and troubleshooting under pressure.
Proven ability to translate reliability principles (SLIs, SLOs) into actionable monitoring and alerting systems.
Comfortable working cross-functionally with cloud networking, infrastructure, and product teams to embed reliability and scalability from design to deployment.