





Mid-level metro role with offshore option and popular platform title yields high applicant competition.
Highly specialized AI/LLM infrastructure, GPU and Azure focus limits transferability across general industries.
Explicit 5–8 years requirement and extensive must-have infra and AI/LLM tech stack create high shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and manage the complete infrastructure stack below the application layer including compute, container orchestration (Kubernetes), RAG pipelines, CI/CD, observability, load testing, and security.
Design, build, and maintain scalable, reliable, and secure infrastructure for AI and LLM production platforms supporting multiple product teams.
Lead InfoSec compliance, security governance, incident response, and platform operational excellence including on-call support.
5–8 years experience in Platform Engineering, Site Reliability Engineering (SRE), or Infrastructure Engineering with strong ownership of infrastructure (full-stack experience alone is insufficient).
Hands-on experience deploying and managing production-grade AI or LLM infrastructure including RAG pipelines and GPU compute.
Strong skills in Kubernetes, Docker, Helm, Azure cloud platform (especially Azure DevOps, Azure Key Vault, Azure Monitor), Terraform for IaC, GitOps (ArgoCD or Flux), and observability tools (Datadog, Elasticsearch, OpenTelemetry).
Capability to manage multiple infrastructure priorities, lead InfoSec security reviews, and provide incident response and on-call support.
Experienced infrastructure engineer specializing in AI/LLM platform production environments with demonstrated ownership of scalable containerized infrastructure and RAG pipelines.
Comfortable independently managing security compliance processes and InfoSec reviews for customer-facing products.
Able to juggle competing demands across multiple product teams and operate autonomously without dedicated project management support.