





Strong employer brand, metro location and platform title offset by seniority and niche GPU/ML platform specialization.
Specialized GPU and AI/ML platform skills reduce transferability despite general cloud/devops applicability.
Explicit 8+ years and mandatory GPU, cloud, Kubernetes, IaC, and certification requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end operational health, reliability, and performance of AI/ML platforms across Azure, AWS, and on-premises GPU environments.
Manage GPU infrastructure including maintenance, optimization, and troubleshooting of GPU nodes and clusters using Kubernetes and related tools.
Automate platform provisioning and maintain CI/CD pipelines while ensuring security, compliance, and observability of AI workloads and infrastructure.
Bachelor's/Master's degree in Computer Science, Engineering, or related field.
8+ years in Platform/Infrastructure/DevOps/SRE roles; at least 3+ years in AI/ML platform engineering.
Strong hands-on experience with Azure, AWS, Kubernetes, GPU infrastructure, Linux system administration, and IaC tools like Terraform and Ansible.
Work Location: Chennai, Office working, Permanent employment type.
Experienced in managing hybrid-cloud AI/ML infrastructure involving large-scale GPU clusters and multi-cloud environments.
Skilled in troubleshooting complex platform issues across application, GPU, networking, and storage layers under SLA-driven conditions.
Capable of collaborating with cross-functional teams including Data Scientists and ML Engineers to optimize production AI workloads and maintain enterprise-grade platform reliability.