





Senior niche ML-systems role at a notable AI-hardware startup reduces general applicant density.
Highly specialized ML, compiler, runtime, and hardware integration skills limit cross-industry transferability.
Explicit 10+ years requirement plus deep learning, C++, compiler, and low-level optimization skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead end-to-end bring up of state-of-the-art open-source ML models, frameworks, and data engineering on Cerebras CSX AI chip systems.
Own cross-stack activities including model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
Identify, debug, and resolve performance and correctness issues across model code, compiler IRs, runtime behavior, and hardware utilization, and prototype tool/API improvements.
Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field.
10+ years of relevant experience in AI software and systems engineering.
Proficiency in Python, C/C++, deep learning frameworks (PyTorch, TensorFlow), and low-level optimization techniques.
Strong debugging skills in performance, numerical accuracy, and runtime integration across full AI toolchain.
System-minded generalist comfortable working rapidly across the entire AI software stack in fast-paced environments.
Deep understanding of deep learning model internals such as attention mechanisms, mixture of experts (MoE), and diffusion models.
Experienced in optimization techniques including tackling NP-hard problems and improving compiler and runtime performance on cutting-edge AI hardware.