





Strong employer brand and metro location increase competition, but niche ML-platform skills moderate applicant density.
Role requires specialized ML-infrastructure tools and HPC/cluster experience, so skills are less transferable across industries.
Explicit 12+ years, 5+ years ML-platform experience, and mandatory expertise in multiple specialized tools make filtering strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Maintain and optimize CrowdStrike's ML infrastructure supporting billions of events daily to ensure platform reliability.
Diagnose, debug, and resolve complex distributed system and ML platform issues including root cause analysis of production incidents.
Develop tools, monitoring solutions, and runbooks for observability, performance optimization, and incident response.
12+ years experience in distributed systems engineering.
5+ years experience debugging ML platforms in production.
Expertise in at least one technology among Ray, Spark, JupyterHub, SLURM, or Kubernetes performance optimization.
Programming expertise in Python and proficiency with ML platform and infrastructure tools including Kubernetes, Docker, and Cloud platforms (AWS/GCP/Azure/OCI).
Strong track record in managing and debugging production distributed ML infrastructure at scale, especially with Ray, Spark, or SLURM.
Experience building observability, incident response procedures, and mentoring teams in debugging and platform optimization.
Background that includes open-source ML infrastructure contributions, on-call incident management, or chaos engineering for high-throughput systems.