





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Mid-level role, metro location, and visible MLOps skillset create moderate candidate competition.
Platform and ML-specific experience is required but skills are reasonably transferable across industries, so medium sensitivity.
Explicit 3+ years plus many mandatory infra, cloud and ML-platform skills makes shortlisting highly strict.
Design, build, and operate the AI/ML platform on AWS + Kubernetes, managing clusters, networking, IAM, storage, cost, and reliability.
Own provisioning and evolution of infrastructure using Terraform as code, including CI/CD pipelines from repo to production for data pipelines, ML models, and AI applications.
Implement and maintain platform observability stacks (Prometheus, Loki, Grafana, Datadog) and partner with Data Science and GenAI teams to optimize platform for ML and agentic AI workloads at scale.
3+ years engineering experience in complex technical environments.
Deep production experience with Kubernetes and hands-on AWS experience with at least 3 services including EKS, Lambda, EC2, S3, IAM, VPC, ALB/NLB, RDS, EMR, Glue, Athena, Batch, SageMaker, MWAA/Airflow.
Proficient with Terraform modules, state management, reviews, and drift detection.
Experience with AI/ML platforms in production, including data science tooling, model lifecycle, infrastructure, and self-serve enablement for GenAI and ML teams.
Strong operator of Kubernetes and AWS infrastructure as code with Terraform in production-scale environments.
Experience running containerized AI/ML and agentic AI workloads including autoscaling and GPU management on Kubernetes.
Familiarity with modern Agentic AI stacks, LLM serving, observability tools, and integration of AI agents on large scale platforms.