





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 employer, mid-level requirement, metro location, and visible ML-platform role increase candidate competition.
Highly specialized ML-platform and HPC skills limit cross-industry transferability.
Explicit 5+ years and deep ML-platform, distributed-systems, and tooling expertise create stringent filters.
Maintain and optimize CrowdStrike's ML infrastructure supporting billions of events daily, ensuring platform reliability and incident resolution.
Diagnose and debug issues in distributed ML systems including Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM with root cause analysis on production incidents.
Develop and implement observability solutions, performance optimizations, and incident response procedures to sustain system stability and capacity planning.
5+ years of experience in distributed systems engineering and debugging ML platforms in production.
Deep expertise in at least three technologies: Ray, Spark, JupyterHub, SLURM, Kubernetes, Airflow, MLflow, or similar ML platform tools.
Strong programming skills with expert Python debugging, multi-language proficiency, and Linux/Unix environment experience.
Work Experience Required: 5+ years in relevant distributed systems and ML platform engineering roles.
Experienced in high-throughput, scalable ML infrastructure in production environments with strong troubleshooting and debugging skills across multiple ML and cloud-native technologies.
Demonstrated capability in system performance profiling, optimization, and capacity planning in large distributed ML clusters.
Familiarity with observability tooling, incident management, and collaboration with ML engineers to improve reliability and operational workflows.