Associate - AI Tooling Ops - Platform Reliability Engineer
JefferiesMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessMid-level SRE title, metro location, and broad skillset create high candidate competition.
Core SRE skills transfer across industries, though finance-specific platform knowledge modestly increases domain sensitivity.
Explicit 3+ years requirement plus mandatory observability, Kubernetes, cloud, and messaging skills raises strictness.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure stability, reliability, scalability, and operational excellence of AI tooling infrastructure on AWS Kubernetes.
Monitor platform health, perform incident triage, troubleshooting, and drive systems improvements to minimize business impact.
Develop automation, dashboards, and observability capabilities to reduce manual interventions and improve service efficiency.
Minimum Requirements
Bachelor's degree in Computer Science, Engineering, IT, or related discipline.
3+ years experience in Site Reliability Engineering, Platform Reliability Engineering, DevOps, Production or Application Support.
Strong programming/scripting skills in Python, Go, C#, Java, or C++.
Experience with Linux/Unix and Windows Server environments, observability platforms (Grafana, Datadog, Prometheus, OpenTelemetry), distributed applications support, and Kafka-based event-driven architectures.
Ideal Candidate Profile
Experienced in production support with strong troubleshooting skills across applications, middleware, databases, messaging, infrastructure, and cloud.
Proficient in building and maintaining monitoring and alerting systems using Grafana, Prometheus, Datadog, OpenTelemetry, and related tools.
Familiar with DevOps practices including CI/CD, source control, Infrastructure as Code, and automation to improve operational efficiency and reliability.
