Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms focusing on performance, capacity, reliability, and scalability.
Lead bare-metal provisioning and lifecycle management in data centers including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and automation of fleet management.
Define and operate SLOs, SLIs, error budgets, alerting, incident-response practices; lead complex incident investigations, corrective actions, and collaborate with multiple teams globally.
Minimum Requirements
Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent experience.
10+ years of experience operating production infrastructure or platform services.
Strong expertise in Kubernetes administration, KubeVirt, Docker, Linux systems, and bare-metal infrastructure provisioning and lifecycle management in data-center environments.
Proficiency in Python, Go, or comparable languages; experience with Infrastructure as Code and automation tools like Terraform, Ansible, Chef, or Puppet; strong understanding of TCP/IP networking and infrastructure security.
Ideal Candidate Profile
Experienced in operating HPC, AI, GPU-accelerated, or GPU-enabled Kubernetes/KubeVirt compute infrastructure.
Familiarity with platforms and hypervisors such as VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.
Exposure to generative AI or agentic workflows for infrastructure diagnostics and operational automation, and to building secure integrated operational platforms using APIs, RBAC, secrets management, and audit controls.
