Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Enhance system resilience and performance through monitoring, automation, and reliability best practices.
Implement and contribute to system disaster recovery strategies and architectural design.
Support continuous improvement initiatives focused on system availability and operational excellence.
Minimum Requirements
Bachelor’s degree in Computer Science, Information Technology, Engineering, or comparable experience; advanced degree preferred.
Experience in software development or technology operations with focus on Site Reliability Engineering.
Knowledge of observability tools like Splunk, Elastic Search, Prometheus, Grafana, and containerization technologies such as Kubernetes and Docker.
Experience with public cloud platforms (AWS, Azure, or Google Cloud) and proficiency in Linux/Unix systems, Java, Python, and Bash scripting.
Ideal Candidate Profile
Candidate familiar with cloud-based SRE practices and able to integrate monitoring, logging, and tracing tools effectively.
Experienced in microservices architecture and automation to drive operational efficiency and reliability.
Capable of contributing to architectural decisions and disaster recovery planning within large-scale, enterprise environments.
