Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and build CI/CD pipelines for AI systems: automated testing, evaluation gates, controlled releases, and rollback.
Manage AWS cloud infrastructure as code ensuring reproducibility, least privilege, cost control, and environment parity.
Lead reliability practices including SLOs, incident response, capacity management, self-healing loops, and build agent telemetry/observability for nondeterministic AI systems.
Minimum Requirements
5-6 years of work experience in DevOps, SRE, platform, or infrastructure engineering with production accountability.
Strong CI/CD experience using tools like GitHub Actions, GitLab CI or similar.
Proficient with AWS infrastructure management via Infrastructure as Code tools such as Terraform or CDK.
Hands-on experience with observability tools (metrics, logging, tracing, dashboards, alerts) such as OpenTelemetry, Prometheus/Grafana, Datadog, or CloudWatch.
Ideal Candidate Profile
Experienced with production operation and observability of AI or LLM systems, including token/cost accounting and failure mode handling.
Comfortable working on early-stage platform foundational builds rather than mature systems, capable of establishing standards and operational patterns.
Strong software fluency to interpret application code for telemetry wiring and debugging deployments independently.
