AI Site Reliability Engineer (AI SRE)
Elfonze Technologies Private LimitedMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessNiche AI SRE skillset but broad requirements and mid-senior level create moderate applicant competition.
Highly specialized AI/LLM SRE and MLOps experience reduces cross-industry transferability.
Explicit 7-10 years and many mandatory cloud, SRE, MLOps, and tooling requirements impose strict filters.
Job Description
Structured overview of role & requirementsAbout This Role
Own reliability, scalability, security, and operational excellence of AI/ML production platforms and AI-enabled applications.
Lead incident response, define SLOs/SLIs, automate operational workflows, and maintain AI service health monitoring including AI-specific metrics.
Collaborate cross-functionally to ensure production readiness, mentor junior engineers, and influence engineering standards for AI/ML platforms.
Minimum Requirements
7-10 years of experience in Site Reliability Engineering, DevOps, cloud operations, MLOps, or related fields.
Strong proficiency in Python and/or Go; experience with Kubernetes, Docker, Linux, networking, API gateways, and cloud platforms (Azure, AWS, or Google Cloud).
Experience with observability tools (e.g., OpenTelemetry, Prometheus), CI/CD and infrastructure automation tools (e.g., Terraform, Jenkins).
Work Experience Required: 7-10 years; Location: India; Other filters: Not explicitly mentioned.
Ideal Candidate Profile
Experienced with production AI/ML lifecycle management including Generative AI, LLM APIs, RAG, and AI agent frameworks.
Able to translate reliability signals into prioritized engineering actions and communicate clearly with technical and non-technical stakeholders.
Demonstrated leadership in incident management, capacity planning, automation, and operational excellence at enterprise scale AI platforms.
