Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own Tier 3 incident, problem, and change management for HPC storage technologies within Managed Services environments.
Administer and optimize parallel/distributed filesystems (e.g., Lustre, GPFS, Ceph) to enhance performance for AI training and inference workloads.
Build automation for storage provisioning, monitoring, and support large-scale HPC/AI storage clusters including troubleshooting across storage, Linux, network, and I/O layers.
Minimum Requirements
5+ years experience in HPC, AI infrastructure, or large-scale storage engineering.
Bachelor's degree or equivalent in Information Systems or related field, or equivalent specialized experience.
Strong Linux systems administration skills and hands-on experience managing distributed/parallel filesystems with storage tuning for performance-sensitive workloads.
Experience with HPC schedulers (e.g. Slurm), container platforms (e.g. Kubernetes), high-speed interconnects (InfiniBand/RDMA), and data protection strategies including replication and disaster recovery.
Ideal Candidate Profile
Experienced in supporting storage solutions tailored for GPU clusters and AI/ML workflows at multi-petabyte scale.
Skilled in automation tools (e.g. Terraform, Ansible, Helm) and observability platforms (e.g. Prometheus, Grafana) for managing complex HPC storage infrastructures.
Demonstrated ability to troubleshoot and resolve issues spanning storage, compute, network, and I/O in multi-tenant, high-performance production environments.
