





Tier-1 brand and Bangalore metro increase competition despite the role's niche ML-platform specialization.
High because deep ML-platform and distributed systems expertise is required and not easily transferable across industries.
High due to explicit 12+ years requirement and mandated deep ML-platform expertise and tooling.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Maintain and optimize mission-critical ML infrastructure handling billions of events daily with a focus on platform reliability and performance.
Diagnose and resolve complex issues across distributed systems components including Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM.
Develop debugging tools, observability solutions, runbooks, and collaborate with ML teams on incident response and post-mortems.
12+ years experience in distributed systems engineering.
5+ years experience debugging ML platforms in production environments.
Deep expertise in at least three of the following: Ray, Spark, JupyterHub, SLURM, Kubernetes.
Expertise in Python debugging and proficiency with Kubernetes, Docker, and at least one cloud platform (AWS/GCP/Azure/OCI).
Experienced in performance profiling, optimization, and capacity planning for large scale ML infrastructure.
Skilled at root cause analysis, incident management, and developing diagnostic tools for distributed ML environments.
Demonstrates ability to collaborate closely with ML engineers and contribute to debugging guides or open-source ML infrastructure projects.