Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure stability, performance, and operational excellence of cloud infrastructure for AI-driven healthcare products.
Monitor and troubleshoot cloud infrastructure, Kubernetes clusters, CI/CD systems, and platform services, including incident response.
Develop automation, Infrastructure-as-Code modules, and maintain operational documentation to improve platform reliability and efficiency.
Minimum Requirements
Bachelor's or Master's degree in Computer Science, IT, Software Engineering, or related technical discipline.
1-3 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure Engineering, or related role.
Proficiency in Linux administration, scripting (Python, Bash, or Shell), and hands-on experience with cloud platforms (AWS, GCP, or Azure).
Experience with containerization (Docker, Kubernetes), Infrastructure-as-Code tools (e.g., Terraform), and monitoring/observability platforms (Prometheus, Grafana, ELK, Datadog).
Ideal Candidate Profile
Has experience supporting and improving highly available, cloud-native production environments with automation and monitoring.
Comfortable collaborating cross-functionally with software engineering, AI, and product teams to enhance system reliability and developer experience.
Familiar with reliability engineering concepts (SLIs, SLOs, error budgets), incident management, distributed systems, and modern CI/CD practices.
