ML Engineer (Training Infra), Foundational Models
Sarvam AIMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and optimize the distributed training infrastructure for Sarvam's foundational AI models, including large GPU clusters and custom kernel development.
Design and implement advanced parallelism strategies (data, tensor, pipeline, sequence, expert) tailored to various model architectures and scales.
Ensure reliability and fault tolerance of long-running training jobs, addressing performance bottlenecks and operational stability to prevent costly failures.
Minimum Requirements
Bachelor's or Master's in Computer Science or closely related technical field, or equivalent experience.
3+ years experience in ML training infrastructure or large-scale distributed systems; exceptional early-career candidates with strong systems background considered.
Hands-on experience with distributed training frameworks such as Megatron-LM, DeepSpeed, FSDP, NeMo; experience managing real pretraining runs including on-call responsibilities.
Deep knowledge of GPU architecture and CUDA programming; proficiency with GPU profiling tools and strong PyTorch internals expertise.
Ideal Candidate Profile
Experienced with designing and tuning high-performance distributed training systems at scale, including custom CUDA/Triton kernel development.
Background in training very large models (10B+ parameters) on extensive GPU clusters (1000+ GPUs) with strong operational ownership.
Proven contributions to open-source training infrastructure projects and familiarity with cluster orchestration/job scheduling systems at scale.
