






Mid-level data role, metro location, and broad skillset drive high applicant competition.
Core data engineering skills are transferable but speech/NLU research specialization raises fit sensitivity to medium.
Explicit 5+ years plus mandatory Python, Spark and SQL requirements increase shortlisting strictness to high.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and maintain data backend pipelines for speech recognition, NLU, and AI models ensuring accurate named entity aggregation from multiple sources.
Develop, test, and monitor research tools and infrastructure related to text processing, data extraction, and CI/CD for data pipelines.
Implement multilingual text processing including data scraping, cleaning, and relevance sorting to support voice-powered automotive AI.
MSc/MTech in Computer Science, Engineering, or equivalent; excellent BSc/BTech may be considered.
Minimum 5 years professional work experience, with at least 2 years utilizing Python extensively.
Skills in Python scripting (Pandas), JavaScript, Spark/PySpark, SQL/RDF querying, and Unix/Linux user level.
Proficiency in English, experience working in international and distributed teams.
Experienced in developing and maintaining data pipelines for large scale AI or speech recognition systems, including named entity management.
Comfortable with multi-language text data parsing and research tool development, with prior exposure to cloud infrastructure and orchestration frameworks.
Strong technical proficiency in data engineering tools (Git, DVC, Airflow) and familiarity with Elasticsearch, data visualization, and natural language understanding domains.