





Tier-1 brand, mid-level generalist role in Pune with broad sought-after skills attracts many qualified applicants.
Strong specialization in AI inference, GPU/system programming, and observability makes cross-industry moves difficult.
Explicit 5+ years requirement plus mandatory systems programming, CI/CD, and GPU/inference expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, implement, and maintain system software and infrastructure to run AI workloads across multiple inference backends such as Llama.cpp, Ollama, PyTorch, vLLM, WinML, and TRT-RTX.
Develop and maintain infrastructure for deploying, benchmarking, qualifying AI applications and models, including automated build, integration, and deployment systems.
Analyze system workload metrics, implement data processing pipelines, build visualizations, and collaborate to add system-level diagnostics and debugging capabilities.
5+ years of software development experience with proficiency in C/C++, C#, Java, or another systems programming language and at least one scripting language such as Python.
Degree: B.Tech or higher in Computer Science, IT, Software Engineering, or related field.
Experience with system APIs, multithreaded software, process and resource management, debugging, performance analysis.
Familiarity with databases and SQL, source control systems (Git, Perforce), and CI/CD tools like Jenkins.
Experienced in building scalable system software, runtime infrastructure, or backend components supporting AI workloads.
Skilled in system profiling, performance optimization, debugging, observability tooling (e.g., Grafana, Kibana), and resource utilization analysis for CPU/GPU.
Hands-on experience with Linux/Windows system programming, containerization (Kubernetes, Docker), integration or debugging of AI inference frameworks (Llama.cpp, PyTorch, etc).