Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design and operate observability platforms including monitoring, logging, alerting using Datadog, managing SLOs, error budgets, and incident management.
Manage AWS infrastructure (EKS, ECS, EC2, networking, IAM) and operate Kubernetes containerized platforms ensuring system reliability and cloud security.
Monitor and optimize costs related to ML models and API-based AI services, collaborate with security teams on incident management and mitigation.
Minimum Requirements
6-9 years of hands-on experience in SRE, Platform Engineering, Infrastructure or related roles.
Strong hands-on experience with AWS services, especially EKS, ECS, EC2, networking, IAM, and managed services.
Deep expertise in Kubernetes and containerized platforms, and designing/operating observability platforms with Datadog.
Onsite presence required at client office three times a week.
Ideal Candidate Profile
Experienced in applying SRE principles (SLOs, error budgets, incident management) in complex infrastructure environments.
Skilled in cloud security engineering, including managing security incidents and collaborating with security teams.
Knowledgeable in FinOps practices for cost control of ML and AI services and comfortable with operational trade-offs balancing reliability, security, and cost.
