Associate - AI Tooling Ops - Platform Reliability Engineer
JefferiesMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessMid-level SRE title, 3+ years, Pune metro, and broad Kubernetes/observability requirements increase competition.
Core SRE and cloud skills transfer across industries, though post-trade platform experience adds some domain specificity.
Explicit 3+ years plus mandatory observability, Kubernetes, cloud, and messaging skills create strict shortlisting filters.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure stability, reliability, scalability, and operational excellence of AI Tooling Infrastructure on AWS Kubernetes.
Monitor platform health, perform incident management, and drive systemic improvements in reliability and performance.
Develop automation, observability solutions, and dashboards to reduce manual intervention and improve service efficiency.
Minimum Requirements
Bachelor's degree in Computer Science, Engineering, IT, or related discipline.
3+ years experience in Site Reliability Engineering, Platform Reliability Engineering, DevOps, Production Support, or Application Support.
Strong programming/scripting skills in Python, Go, C#, Java, or C++.
Experience with Linux/Unix and Windows Server, distributed application support, monitoring tools (Grafana, Datadog, Prometheus, OpenTelemetry), and Kafka-based messaging platforms.
Ideal Candidate Profile
Experienced in operational support and automation within complex, distributed, cloud-native environments (AWS Kubernetes).
Skilled at cross-team collaboration including development, infrastructure, and business stakeholders across global regions.
Strong analytical capabilities to diagnose multilayered system issues and improve platform observability and reliability.
