Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Architect and scale cloud-native platform foundations supporting AI-driven products and internal platforms, ensuring reliable, secure, and cost-efficient infrastructure.
Define and enforce DevOps standards (CI/CD pipelines, deployment patterns, observability, security, cost governance) across the AI Factory, acting as final escalation point for complex infrastructure and reliability challenges.
Mentor senior engineers and influence platform strategy, reliability practices, and cloud cost governance across multiple teams without direct people management responsibility.
Minimum Requirements
8-12 years of experience in DevOps, Platform, or Infrastructure Engineering with proven Staff/Principal-level scope and multi-team impact.
Bachelor's or Master’s degree in Computer Science, Engineering, or related field.
Strong hands-on expertise with AWS, GCP, Azure and deep knowledge of Docker, Kubernetes, CI/CD pipelines, infrastructure automation (Python/Bash), monitoring and observability tools (Prometheus, Grafana, ELK/Loki).
Work Experience Required: 8-12 years in relevant domain; Notice Period: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced senior individual contributor comfortable operating at platform level with multi-team influence and handling escalations for complex infrastructure and reliability issues.
Proactively identifies systemic risks including capacity, cost, security and scaling; drives reliability engineering practices such as SLOs, incident response, and resilience planning.
Strategic thinker who prioritizes long-term maintainability and acts as a technical authority communicating complex trade-offs clearly to leadership and stakeholders.
