Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, metro location, popular DevOps/SRE role with broad tooling increases candidate competition.
Core SRE and observability skills are broadly transferable across industries.
Multiple mandated platform skills and SRE experience make shortlisting stringent.
Job Description
Structured overview of role & requirementsAbout This Role
Innovate, extend, and deploy internal platforms for Gitlab, Github, ADO, and JFrog Artifactory across data centers and SaaS environments.
Manage monitoring, incident response, and observability engineering using tools like Prometheus, Grafana, Dynatrace, OpenTelemetry, and automate operational tasks via Terraform, Python, Helm, and GitOps.
Collaborate with teams to ensure platform reliability, optimize performance, and enhance developer enablement and monitoring for microservices, APIs, infrastructure, and databases.
Minimum Requirements
Experience Required: Not explicitly mentioned in the JD but 2+ years IT experience is a nice-to-have.
Hands-on experience with observability tools such as Prometheus, Grafana, Dynatrace, OpenTelemetry, ELK/OpenSearch, or similar.
Proficient in scripting/automation languages including Python, Bash, or PowerShell and infrastructure automation with Terraform, Ansible, Helm, GitLab/GitHub, CI/CD, GitOps.
Strong understanding of Kubernetes, cloud-native architectures, networking, Linux, incident management, and production support.
Ideal Candidate Profile
Background in SRE, Production, or Platform Engineering roles with strong expertise in observability and telemetry pipelines across microservices and cloud environments.
Operator style focused on automation, monitoring, incident resolution, and scalable platform optimization using a variety of modern DevOps and cloud native tools.
Experience or knowledge of cloud platforms (AWS, Azure, GCP), Kubernetes distributions (EKS, AKS, OpenShift), advanced observability query languages (PromQL, LogQL), and exposure to AI/ML for observability or chaos engineering is advantageous.
