Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and maintain observability infrastructure using Elastic Stack for Azure and AWS workloads.
Lead incident response for distributed, multi-tenant cloud services including runbook creation and continuous improvement.
Design and implement proactive support tooling to reduce reactive support efforts and improve operational maturity.
Minimum Requirements
8+ years experience in cloud platform engineering, SRE, or infrastructure roles supporting commercial SaaS products.
Hands-on expertise with Elastic Stack: Elasticsearch, Kibana, Elastic Fleet, KQL/Query DSL.
Strong knowledge of Azure cloud services (AKS, Entra ID, Key Vault, Service Bus, Cosmos DB, Private Endpoints) and multi-cloud experience including AWS.
Experience with Infrastructure as Code tools (Azure Bicep, Terraform) and CI/CD tools (Azure DevOps, GitHub Actions).
Ideal Candidate Profile
Experienced operating and troubleshooting distributed, multi-tenant cloud SaaS workloads with strong incident response capabilities.
Skilled in cross-functional collaboration with SRE, R&D, and proactive support teams to close observability gaps and improve tooling.
Proficient scripting skills (Bash, Python, PowerShell) and familiarity with monitoring and alerting processes in complex cloud environments.
