Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Drive reliability, scalability, and operational excellence of key post-trade processing platforms focusing on monitoring, incident management, and performance improvements.
Design and implement scalable, resilient solutions in collaboration with engineering, infrastructure, and business teams, including automation to reduce manual intervention.
Build and enhance observability and deployment capabilities using tools like Grafana, Prometheus, and OpenTelemetry; support Kafka-based enterprise messaging and event-driven architectures.
Minimum Requirements
Bachelor’s degree in Computer Science, Engineering, IT, or related field.
3+ years of experience in SRE, Platform Reliability Engineering, DevOps, Production or Application Support.
Strong programming/scripting skills in Python, Go, or Java; working knowledge of Linux/Unix and Windows Server environments.
Experience with observability tooling, incident and problem management, and familiarity with Kafka or similar messaging platforms.
Ideal Candidate Profile
Experienced in production support and reliability engineering within distributed, event-driven platforms, especially involving Kafka.
Technically strong in automation, monitoring, and troubleshooting distributed systems in fast-paced environments.
Able to collaborate effectively with technical and business stakeholders to improve system resilience and operational visibility.
