





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Mid-level Pune metro role with demand but niche GPU/MLOps skills reduce applicant density.
Specialized AI infra and GPU cluster experience limits transferability across non-ML industries.
Explicit 3-5 years plus mandatory cloud, Kubernetes, IaC, and GPU/MLOps skills required.
Design and manage AI infrastructure systems including GPU clusters and Kubernetes orchestration for scalable AI model training and deployment.
Build, optimize, and maintain distributed training frameworks, data pipelines, and automate infrastructure provisioning using tools like Terraform or Ansible.
Mentor junior team members, troubleshoot complex infrastructure issues, and collaborate with ML engineers to resolve training performance bottlenecks.
Bachelor's degree in Computer Science, Engineering, or related field.
3-5 years of experience in DevOps, cloud infrastructure, or SRE roles.
Proficiency in Linux administration, container orchestration (Kubernetes, Docker), Python, and Infrastructure as Code tools.
Experience with cloud platforms (AWS, GCP, or Azure) and understanding of machine learning workflows.
Experienced in hands-on management of GPU clusters or HPC environments and distributed training frameworks (PyTorch DDP, DeepSpeed).
Familiar with Kubernetes administration, cloud architecture, and MLOps practices, preferably with relevant certifications (CKA, AWS/GCP/Azure).
Effective collaborator with ability to mentor junior staff and optimize AI infrastructure for performance and cost (FinOps).