Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design and manage lifecycle of large scale distributed systems including system design, build, deployment, and monitoring to maintain high availability and performance.
Integrate and leverage AI, Generative AI, and AIOps tools such as AWS Bedrock, Azure OpenAI, Moogsoft, and Datadog to automate incident detection, root cause analysis, and resolution workflows, including developing AI/ML driven monitoring and alerting systems.
Develop software and automation solutions using high-level programming languages (Python, Ruby, GoLang) to improve system efficiency, capacity, and operational tasks across public and private cloud environments.
Minimum Requirements
Experience in designing and managing large scale distributed systems and infrastructure on public and private clouds.
Proficiency in one or more programming languages: Python, Ruby, or GoLang with Object Oriented Programming familiarity.
Hands-on experience with integrating and deploying AI/Generative AI and AIOps platforms (e.g., AWS Bedrock, Azure OpenAI, Moogsoft, Dynatrace, Splunk, Datadog).
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Operates strongly at the intersection of site reliability engineering and AI-driven automation focusing on incident management and proactive anomaly detection.
Experienced in architecting solutions incorporating AI/ML models and prompt engineering techniques tailored for cloud environments.
Capable of managing high availability architecture and scaling infrastructure globally using DevOps best practices and advanced observability tools.
