Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Implement, maintain, and optimize AI platform capabilities on Google Kubernetes Engine (GKE) in Google Cloud Platform (GCP) to enable product teams to build and deploy Generative AI applications and self-hosted models.
Collaborate with application, data science, and platform teams to troubleshoot integration issues, implement cloud-native AI tools and observability standards, and improve platform efficiency and developer experience.
Execute on technical roadmaps to improve platform reliability, reduce operational overhead, and accelerate delivery cycles for AI product features.
Minimum Requirements
5+ years of relevant work experience.
Hands-on experience deploying and scaling self-hosted Large Language Models (LLMs) on GKE, including GPU resource management.
Strong practical knowledge of Google Cloud Platform services including GKE, GCE, Istio service mesh, GitOps tools (e.g., ArgoCD, Flux), Infrastructure as Code (Terraform), and AI observability tools like OpenTelemetry.
Proficiency in Python or Go programming for automation, tooling, and custom resource development.
Ideal Candidate Profile
Experienced with cloud-native AI infrastructure and generative AI workloads, including integration with GenAI frameworks and production deployment in Kubernetes.
Skilled in platform engineering with a focus on software engineering best practices, clean code, testing, and automation within a large cloud environment.
Demonstrates capability to collaborate cross-functionally, provide clear technical documentation, and contribute to platform standards and reusable templates for developer tooling.
