Site Reliability Engineer – Kubernetes
Innovon TechnologiesMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Deploy, configure, scale, and manage containerized applications using Docker and Kubernetes, including Azure Kubernetes Service.
Manage and optimize complex monitoring stacks involving Prometheus, InfluxDB, and Grafana with advanced dashboard scripting.
Implement and maintain Site Reliability Engineering practices including CI/CD pipelines, infrastructure as code, performance tuning, and disaster recovery in a regulated environment.
Minimum Requirements
Strong expertise in Kubernetes, Docker, and container orchestration technologies including service mesh and Azure Kubernetes Service.
Experience managing Prometheus, InfluxDB, and Grafana monitoring tools with proficiency in PromQL, InfluxQL/Flux query languages.
Proficient in scripting languages such as PowerShell, Python, Bash, and C# and knowledge of infrastructure automation tools like Ansible, Terraform, and TeamCity.
Work Experience Required: Not explicitly mentioned in the JD
Ideal Candidate Profile
Experienced in operating and optimizing complex distributed systems in production environments, especially with container orchestration and monitoring.
Comfortable working under pressure during outages with strong problem-solving skills and accountability for service reliability.
Capable of communicating complex technical concepts clearly to non-technical stakeholders at various organizational levels.
