





Remote role, mid-level SRE title, metro location, and broad cloud/Kubernetes/Terraform requirements increase applicant competition.
Role requires core SRE/cloud skills transferable across industries but demands specific infrastructure experience.
Explicit 5+ years plus mandatory Terraform/AWS/Kubernetes/observability skills enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the reliability, uptime, and performance of an enterprise Generative AI platform serving tens of thousands of users.
Maintain and enhance infrastructure provisioning, observability, and incident response including CI pipelines, Terraform infrastructure in AWS, and container orchestration (ECS/Kubernetes).
Partner with developers on production readiness reviews, root cause analysis, and security hardening to ensure safe scalable operations.
5+ years experience in SRE, DevOps, or infrastructure engineering with cloud IaC tool expertise.
Proficient in Terraform authoring and AWS infrastructure deployment (including EC2, ECS, ALB).
Experience with monitoring and alerting tools such as CloudWatch, Datadog, or Prometheus/Grafana.
Ability to read and patch application codebases in Node.js, Python, or Go; experience with CI/CD pipelines (GitHub Actions, Jenkins, etc.).
Experienced operating enterprise-scale platforms with 10,000+ users, including ECS/Kubernetes management and PostgreSQL at scale.
Comfortable delivering technical feedback to developers and shaping production readiness and security posture with clear communication.
Familiar with AI/LLM platform operations and using AI coding assistants to enhance productivity.