Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms focusing on performance, reliability, capacity, and operational scale.
Lead bare-metal provisioning and lifecycle management in data centers including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.
Develop automation, self-service capabilities, observability solutions, define and operate SLOs/SLIs/error budgets, lead incident response, and partner with cross-functional teams on global platform initiatives.
Minimum Requirements
Bachelor's degree in Computer Science, Engineering, related technical field, or equivalent experience.
10+ years operating production infrastructure or platform services.
Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, Linux systems, and distributed systems troubleshooting.
Experience deploying and operating bare-metal infrastructure in data centers including provisioning, networking, OS lifecycle management, hardware automation, plus proficiency in Python or Go programming.
Ideal Candidate Profile
Experienced in operating HPC, AI, GPU-accelerated, or GPU-enabled Kubernetes/KubeVirt bare-metal compute infrastructure.
Proficient with Infrastructure as Code and automation tools (Terraform, Ansible, Chef, Puppet) and observability tools (Prometheus, Grafana, OpenTelemetry, ELK Stack, Splunk).
Skilled in building secure, integrated operational platforms with APIs, RBAC, secrets management, and incident management systems, capable of leading complex infrastructure projects.
