Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessStrong Tier-1 brand, mid-level generalist platform role, and metro location increase applicant competition.
Kubernetes, cloud, and CI/CD skills are broadly transferable across industries, reducing background sensitivity.
Explicit experience, Kubernetes, cloud, compliance, and tooling requirements create strict shortlisting filters.
Job Description
Structured overview of role & requirementsAbout This Role
Own platform support and operations for Walmart Cloud Native Platform (enterprise Kubernetes), including cluster health management, incident response, and on-call duties to meet service level objectives.
Manage cluster lifecycle activities like provisioning, upgrades, maintenance, and security compliance, ensuring platform reliability and efficiency.
Design and build automation pipelines and AI-assisted tooling to reduce manual effort and accelerate incident response and platform operations.
Minimum Requirements
Bachelor's degree in Computer Science, Software Engineering, Information Technology, or related field with minimum 5 years experience in infrastructure engineering, platform operations, or DevOps; master’s degree with 3 years also acceptable.
Strong hands-on expertise in Kubernetes administration and operations, including cluster lifecycle, RBAC, autoscaling, with Certified Kubernetes Administrator (CKA) preferred.
Proficiency with Kubernetes networking, Istio service mesh, container networking, load balancing, CI/CD tools (Jenkins, GitHub Actions), and scripting (Python, Bash).
Experience with cloud platforms (GCP or Azure), observability tools (Prometheus, Grafana, OpenObserve), security/compliance (PCI-DSS, SOC2), and incident management as L2/L3 engineer.
Ideal Candidate Profile
Experienced infrastructure or platform engineer skilled in enterprise Kubernetes platforms and capable of full cluster lifecycle management for large-scale production environments.
Comfortable developing advanced automation including AI/LLM based tools to reduce operational toil and improve incident response efficiency.
Demonstrated ability working cross-functionally with application, SRE, security teams, and experienced with GitOps workflows and service mesh technologies in cloud-native infrastructure.
