Production Support Engineering / Site Reliability Engineering (SRE)
NTT Ltd.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own end-to-end support and operational management of multiple applications/platforms, including incident triage, root cause analysis, change management, and deployment.
Provide on-call production support for mission-critical applications ensuring resolution within Recovery Time Objective (RTO).
Develop automated solutions to reduce repetitive tasks and improve stability and operational excellence with an SRE mindset, particularly for AI Cloud environments.
Minimum Requirements
Minimum 12 years of production support experience with demonstrated implementation of SRE processes.
Hands-on experience and certification in Google Cloud Platform (GCP).
Strong technical expertise in Python, shell scripting, Linux/Unix command line, AI/ML workload infrastructure (including GPU-based environments), and supporting Generative AI applications.
Experience with big data technologies (Hadoop, Spark, Elasticsearch) and monitoring tools like Splunk and AppDynamics.
Work Experience Required: 12+ years production support.
Ideal Candidate Profile
Experienced in supporting and resolving scalability, capacity, and performance issues in AI/ML and Generative AI workloads and platforms.
Proven ability to liaise effectively between infrastructure operations and business projects in large-scale environments.
Comfortable working in 24x7 shift models with strong problem-solving and communication skills in a complex, global technology services environment.
