





Senior, niche ML-infrastructure role with GPU/serving specialization reduces applicant density.
Highly domain-specific ML infra and GPU/serving expertise limits cross-industry transferability.
Explicit 10+ years and mandatory architecture/platform experience make hiring filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own architecture and scaling of real-time data processing and distributed messaging systems with strict latency and throughput requirements.
Lead capacity planning, scheduling, and utilization of GPU infrastructure for model training and inference; optimize cost and latency for scalable ML serving systems.
Drive system reliability, observability, incident response, and collaborate cross-functionally to deliver technical roadmaps and scalable infrastructure improvements.
10+ years experience building backend and infrastructure systems with architecture ownership at scale.
Deep hands-on experience with large-scale databases, high-throughput messaging systems, and real-time job queues.
Strong written communication skills for technical and business audiences across time zones.
Degree: BTech/MTech/PhD in Computer Science or equivalent.
Experienced in scaling GPU infrastructure and optimizing ML inference performance, including capacity planning and autoscaling strategies.
Proven track record of mentoring senior engineers and influencing technical decisions in large complex codebases.
Familiarity with stack elements like Django, Celery, Redis, PostgreSQL, Google Cloud, or relevant ML inference tools (vLLM, TensorRT, Triton, Ray Serve) is preferred but not mandatory.