





Tier-1 employer, mid-level generalist SRE role with broad skills and metro appeal.
Core SRE and cloud skills are highly transferable across industries despite optional regulated experience.
Mandatory 5+ years and specific cloud, Terraform, Kubernetes, and observability skills required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Ensure reliability, performance, and operational excellence of enterprise pro-code AI Agent Platforms.
Build and continuously improve observability, monitoring, alerting, and incident response capabilities including defining SLOs and SLAs.
Deploy, maintain, and optimize cloud-native infrastructure as code (e.g., Terraform), focusing on security, cost-efficiency, scalability, and developer experience.
Bachelor’s degree in Computer Science, Data Science, Engineering, or related field.
5+ years experience in Site Reliability, Operations, or Infrastructure Engineering.
Hands-on expertise with telemetry and observability tools (e.g., OpenTelemetry, Grafana, Langfuse).
Experience with Microsoft Azure and/or Google GCP, Terraform, GitHub Actions, Kubernetes, Linux/Unix administration, and scripting (Bash, Python, Typescript).
Experienced in managing cloud-native platform operations with strong focus on observability and incident response in complex, dynamic environments.
Ability to collaborate cross-functionally with architecture, security, and product teams to shape platform roadmap and deliver developer-centric solutions.
Pragmatic engineer with demonstrated expertise in infrastructure automation, CI/CD pipelines, and operating within regulated environments or compliance frameworks (preferred).