





Medium — Tier‑1 brand and metro location but senior, specialized ML-platform role reduces applicant density.
High — specialized ML-platform and distributed-systems skills tightly tied to AI/ML infrastructure.
High — explicit 12+ years plus mandatory ML-platform, distributed-systems expertise and specific toolset.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Maintain and optimize CrowdStrike's critical ML infrastructure supporting billions of processed events daily.
Diagnose and resolve complex distributed system issues across platforms like Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM ensuring reliability and performance.
Develop observability solutions, debugging tools, runbooks, and conduct incident response to maintain platform stability metrics such as SLAs and latency.
12+ years of experience in distributed systems engineering.
5+ years debugging ML platforms in production environments.
Deep expertise in at least 3 of the following: Ray, Spark, JupyterHub, SLURM, Kubernetes performance tuning.
Proficiency in Python debugging and familiarity with cloud platforms (AWS, GCP, Azure, OCI) and containerization (Kubernetes, Docker).
Experienced in operating and troubleshooting large-scale, ML platform infrastructures under high-throughput conditions.
Skilled in performance profiling, capacity planning, and root cause analysis of distributed ML systems.
Familiar with building debugging frameworks, incident management processes, and collaborating with ML engineering teams for platform stability.