





Mid-level ML inference role at a reputable startup in metro India, making competition medium.
Specialized inference, pruning, and LLM experience reduces cross-industry transferability.
Explicit 1–3 years plus mandatory ML frameworks and inference skills increases shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Adapt and optimize advanced language and vision transformer models to run efficiently on Cerebras AI hardware.
Implement and validate models focusing on speculative decoding, large-model pruning, compression, sparse attention, and sparsity-driven inference techniques.
Collaborate with ML researchers and cross-functional teams to improve low-latency, high-throughput AI inference performance at scale.
Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
1–3 years of experience in software engineering or machine learning, including internships.
Proficiency in Python and at least one ML framework (e.g., PyTorch, Transformers).
Understanding of deep learning concepts and experience with Generative AI and machine learning systems.
Experience with model optimization techniques such as speculative decoding, pruning, compression, sparse attention, and quantization focused on inference.
Familiarity with large language or computer vision models and running experiments for model tuning.
Comfortable working in Linux environments and collaborating closely with ML, software, and hardware teams on cutting-edge AI accelerator platforms.