





Tier-1 brand, metro Bengaluru location, and broad specialized skillset increase applicant competition significantly.
Role requires deep ML-platform, GPU/SLURM and Ray/Spark expertise, making cross-industry transfers difficult.
Explicit 12+ years, 5+ years ML platform experience, and mandatory niche tech expertise make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Maintain and optimize CrowdStrike's mission-critical ML infrastructure supporting billions of events daily.
Diagnose and resolve complex distributed system issues across platforms including Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM to ensure platform reliability.
Develop and implement debugging tools, performance optimizations, monitoring solutions, runbooks, and incident response workflows.
12+ years of experience in distributed systems engineering.
5+ years of experience debugging ML platforms in production environments.
Deep expertise in at least 3 of the following technologies: Ray, Spark, JupyterHub, SLURM, Kubernetes.
Expertise in Python debugging, multi-language programming, Linux/Unix; experience with Kubernetes, Docker, and cloud platforms (AWS/GCP/Azure/OCI).
Proven capability in hands-on debugging and optimization of complex ML platform components and distributed systems at scale.
Experience collaborating with ML teams to resolve workflow and platform issues with a focus on reliability and performance.
Strong knowledge of observability, incident management, and production monitoring in large scale, AI-native infrastructure environments.