





Metro location and broad ML+infra requirements increase candidate density.
High because production ML infrastructure and inference optimization skills are highly specialized and industry-specific.
High due to explicit 8+ years and mandatory production ML, infra, and PyTorch expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and scale production machine learning inference backends and pipelines for diverse models including LLMs, vision, embeddings, and classical ML models at large dataset scale (terabytes to petabytes).
Automate full ML lifecycle workflows including training, deployment, evaluation, rollback, monitoring, and retraining for robust and efficient inference systems.
Optimize system latency, throughput, reliability, and costs through techniques such as batching, caching, parallelism, quantization, and scalable infrastructure improvements.
8+ years of experience in software engineering, machine learning engineering, or ML infrastructure.
Strong backend engineering skills in Python with expertise in APIs, distributed systems, testing, debugging, and operational ownership.
Proven experience building and operating production ML systems, especially on AWS.
Hands-on experience with PyTorch or equivalent modern ML frameworks and automating ML lifecycle from experimentation to deployment and monitoring.
Can independently own and deliver end-to-end technical solutions for large-scale ML production systems with high autonomy.
Experienced in balancing system engineering and ML skills to optimize production ML serving tradeoffs like latency, throughput, scaling, and cost efficiency.
Ability to mentor other engineers while remaining deeply hands-on and making pragmatic architectural decisions in fast-moving environments.