Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own CI/CD pipelines and production deployment for ML and agentic AI systems on cloud-native platforms (AWS, Azure, or equivalent).
Build and maintain observability, monitoring, and security governance for LLM and agentic AI workloads in regulated environments.
Implement infrastructure-as-code and optimize cost/performance for AI/ML platforms, ensuring reliable operational scale from prototype to production.
Minimum Requirements
5+ years experience in cloud/DevOps/MLOps engineering with AWS, Azure, or GCP.
Hands-on experience with production deployment of ML/GenAI systems including CI/CD, containerization (Docker/Kubernetes), and infrastructure-as-code tools (Terraform).
Proficiency with MLOps tools (e.g., MLflow, SageMaker Pipelines, Azure ML Pipelines) and cloud-native AI platforms (e.g., AWS Bedrock, SageMaker, Azure AI Foundry).
Strong scripting skills in Python and Bash plus security/governance experience in regulated environments (RBAC, secrets management, audit).
Ideal Candidate Profile
Deep operational knowledge of LLM and agentic AI systems including observability, cost monitoring, regression and hallucination testing.
Experience working in regulated sectors such as pharma, life sciences, or financial services is a strong plus.
Ability to collaborate cross-functionally with data engineers, AI engineers, and business stakeholders to scale AI systems securely and reliably into production.
