





Tier‑1 brand, popular AI/ML title, and metro location drive strong competition.
Highly domain-specific inference and quantization expertise limits cross-industry transferability.
Explicit 4–10 years requirement and mandatory inference, quantization, and runtime skills create stringent filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own design, development, and validation of accuracy evaluation pipelines for deep learning inference on large data-centre hardware.
Implement and maintain accuracy KPIs, automated accuracy evaluation tools, and optimizations across multiple inference runtimes (TensorRT, ONNX Runtime, Triton, etc).
Perform deep accuracy analysis including debugging, quantization validation, architecture-driven degradation identification, and execution of accuracy recovery experiments.
Bachelor's degree or higher in Engineering, Information Systems, Computer Science, or related field.
4+ years of Software Engineering or related experience; Software Test Engineering experience specifically: minimum 2+ years.
Strong Python programming skills and hands-on experience with inference runtimes (TensorRT, ONNX Runtime, Triton).
Experience with deep learning model architectures (transformers, CNNs, RNNs), accuracy analysis, and quantization techniques (INT8, FP16, calibration, QAT).
Proven ability to develop scalable and automated accuracy evaluation pipelines across multiple ML frameworks and hardware platforms.
Experience working with large language models (LLMs), generative AI, and complex multi-modal models to identify accuracy bottlenecks and optimize inference.
Demonstrated expertise in debugging accuracy failures with strong model architectural understanding and statistical accuracy analysis methods.