





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Well-funded Tier-1 startup, metro Bangalore, mid-level role but niche GPU/MLOps skills limit applicant density.
High industry specificity: GPU cluster and MLOps expertise not easily transferable outside AI infrastructure roles.
Explicit years plus mandatory GPU, CUDA, Kubernetes, and inference-platform requirements make shortlisting highly strict.
Own and build the AI platform enabling AI/ML engineers to deploy, version, evaluate, and monitor workloads with self-service APIs and workflows.
Manage GPU infrastructure including cluster design, provisioning, scheduling, isolation, and utilisation optimization for multi-tenant enterprise environments.
Run and optimize large-scale LLM inference production workloads meeting latency, throughput, and cost targets while ensuring secure, multi-tenant, observable platform operations.
2-4 years experience building and operating production software and infrastructure with end-to-end ownership.
Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience.
Strong programming skills in Python and/or Go; deep experience with Kubernetes, Docker, Linux, networking, and distributed systems.
Hands-on experience with GPU cluster operations, NVIDIA CUDA ecosystem, GPU scheduling and troubleshooting, and production LLM inference serving using tools like vLLM, Triton, or KServe.
Background in platform engineering, infrastructure, SRE, distributed systems, or MLOps with a focus on AI workloads and GPU infrastructure.
Proven ownership of secure, scalable, multi-tenant SaaS services and backend APIs in cloud environments (AWS/GCP/Azure) using Terraform-driven infrastructure as code.
Experienced in building observability and automated CI/CD pipelines driving measurable improvements in GPU utilisation, inference cost, and model evaluation quality.