





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Remote, mid-level SRE role with popular title and broad applicant pool drives high competition.
Requires core SRE skills plus GPU and FinOps experience, so transferability across industries is moderately constrained.
Explicit 4–5 years requirement and mandatory Kubernetes, GPU, FinOps, GCP, and IaC skills enforce high strictness.
Own and optimize infrastructure cost efficiency, focusing on Kubernetes overprovisioning reduction and cost telemetry for backend teams.
Lead GPU throughput optimization experiments on on-premise GPU clusters, partnering with AI service owners.
Build tools, dashboards, and processes enabling backend teams to manage their own cost and reliability budgets, and contribute to reliability instrumentation and select platform security workstreams.
4-5 years of hands-on systems engineering experience.
Strong backend engineering skills with production experience in Python, Go, or Rust and ability to own services end to end.
Expertise with Kubernetes at scale including scheduling, resource management, autoscaling, and cost-aware autoscaling tools.
Must be able to work during EST time zone hours (8 PM to 4 AM IST).
Experienced in hybrid cloud and on-prem infrastructure, including GCP, Terraform IaC, CI/CD, and on-prem GPU clusters.
Skilled in GPU workload profiling, batching, server tuning, and utilization metrics for AI/ML workloads.
Demonstrates a FinOps mindset with measurable outcomes on infrastructure costs and can independently drive platform-security related tasks without heavy DevOps reliance.