





Senior, niche AI-platform role at a known security firm in metro locations drives moderate candidate competition.
Role requires deep platform, SRE, and AI/ML infra experience, making cross-industry transfers moderately constrained.
Multiple explicit mandatory requirements (8+ years, 3+ years AI infra, Terraform, Kubernetes, observability) make screening highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain scalable, secure AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform and IaC best practices.
Own and evolve GitLab CI/CD pipelines for AI platform services to enable automated, reliable, multi-environment deployments.
Architect and lead centralized observability (Prometheus, Grafana), incident response, and platform governance to ensure high availability and operational reliability of AI/ML production systems.
8+ years experience as Platform Engineer / Site Reliability Engineer / DevOps Engineer, with 3+ years supporting AI/ML or data platform infrastructure.
Strong hands-on expertise in AWS services including EKS, Lambda, ECS, VPC, IAM, S3.
Proficient in Infrastructure as Code using Terraform and CI/CD pipeline management with GitLab CI/CD or equivalent.
In-depth experience with observability tools (Prometheus, Grafana, Alertmanager) and incident response for production services.
Technical leader comfortable navigating between strategic platform design and hands-on engineering execution at scale.
Deep experience with cloud-native AI/ML infrastructure, Kubernetes container orchestration, and CI/CD automation.
Data-driven decision-maker focused on improving delivery and reliability metrics (DORA), with strong communication skills to mentor and lead engineering teams.