





Tier-1 brand, metro location, generalist 'Backend' title, and mid-level experience create high applicant competition.
Role mixes general backend skills with ML inference and document-AI specifics, so industry fit is moderately sensitive.
Explicit years, on-call production experience, Temporal/Kubernetes/GPU inference requirements make shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the end-to-end architecture of Sarvam's vision model serving harness, balancing accuracy, latency, and cost-per-page at national scale.
Design and implement durable, resumable document processing workflows with Temporal, ensuring reliability, idempotency, and graceful degradation for critical documents like KYC and loans.
Manage inference serving layer, including multi-model routing, batching, GPU utilization, autoscaling, and observability infrastructure to meet strict SLOs and cost targets.
5–6+ years in backend engineering with experience operating high-throughput production systems on-call.
Strong proficiency in Go and/or Python.
Production experience with distributed systems and durable execution engines like Temporal, including workflow orchestration, retry semantics, and idempotency.
Experience with Kubernetes in production, autoscaling, resource management, and ML/LLM inference serving (batching, caching, latency budgeting).
Engineer with strong distributed systems background and judgment over trade-offs in accuracy, latency, and cost in large-scale ML inference serving.
Experience setting and delivering on SLOs, managing reliability incidents, and iterating system cost-effectively without quality loss.
Familiarity with ML inference production environments and orchestration, comfortable owning the technical direction and mentoring junior engineers.