Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer)
Rafay SystemsMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Optimize and serve trained ML/LLM models to real-world users focusing on speed (TTFT, tokens/s), accuracy, and cost efficiency.
Apply model compression techniques (quantization, pruning, distillation) to reduce model size without accuracy loss.
Tune runtime and serving frameworks, manage memory/caching for long contexts and multiple users, and profile GPU/TPU bottlenecks for hardware co-optimization.
Minimum Requirements
4+ years of experience in systems programming, HPC, or ML infrastructure.
Strong proficiency in Python or Go; familiarity with CUDA/Triton kernel development.
Production experience with inference engines like vLLM, TensorRT-LLM, or SGLang and profiling tools like Nsight.
Deep knowledge of transformer architectures and related hardware performance concepts.
Ideal Candidate Profile
Experienced in balancing inference latency, throughput, and accuracy trade-offs in ML deployment.
Skilled in low-level optimization including CUDA/Triton kernel tuning and distributed multi-GPU setups.
Hands-on with high-performance inference systems and tooling to identify and mitigate hardware bottlenecks.
