Vacancy
Senior ML Data Engineer
Location:
Other, Eastern Europe
Seniority:
Senior
Technologies:
BigData, Data, Machine Learning, Python

Zoolatech is looking for a Senior ML Data Engineer to join the team of our client and contribute to its AI and frontier-model initiatives.

In this role, you’ll design and build the data systems that power frontier-scale machine learning research and applied AI products, with a particular focus on spatial intelligence and multimodal data. Your primary goal will be to ensure researchers and engineers can reliably discover, curate, transform, and operate large-scale datasets as projects move from experimentation into production.

You’ll collaborate closely with ML researchers, applied ML engineers, and system architects to transform ambiguous research needs into scalable, production-ready data pipelines. You’ll remain deeply hands-on while providing technical leadership in data architecture, quality, and operational excellence.

This is an opportunity to influence how frontier models are built, evaluated, and deployed by ensuring their underlying data systems are robust, observable, reproducible, and designed for rapid iteration.

  • Act as the technical lead for data engineering efforts supporting frontier-model research and applied ML systems.

  • Design, build, and maintain scalable batch and streaming pipelines for multimodal data, including documents, images, and spatial metadata.

  • Partner with researchers and architects to transform experimental workflows into reliable and repeatable data systems.

  • Lead the development of dataset curation, versioning, and lineage workflows that enable rapid experimentation and reproducibility.

  • Establish and maintain standards for data quality, validation, observability, and cost efficiency across AI data pipelines.

  • Contribute to data architecture decisions spanning research and production environments.

  • Identify gaps and inefficiencies in existing data workflows and develop proofs of concept to evaluate potential improvements.

  • Mentor other engineers through code reviews, design discussions, and hands-on collaboration.

  • Bachelor’s or master’s degree in a relevant field, or equivalent practical experience.

  • 6+ years of experience designing and operating large-scale data systems.

  • Strong Python and SQL skills, with experience in distributed data processing using Spark, Databricks, or equivalent technologies.

  • Proven experience building scalable batch and streaming pipelines for ML training, evaluation, or inference.

  • Experience working with large-scale multimodal or unstructured data, such as documents, images, video, or spatial data.

  • Experience designing lakehouse or equivalent large-scale data architectures.

  • Experience orchestrating reliable production data pipelines using Airflow, Dagster, or equivalent technologies.

  • Strong knowledge of dataset curation, versioning, lineage, data quality, and reproducibility.

  • Production experience with AWS or GCP, including orchestration, observability, and performance and cost optimization.

  • Staff-level technical leadership, including architecture ownership, hands-on contribution, mentorship, and effective collaboration in ambiguous environments.

  • Experience with Kafka, Pub/Sub, or equivalent streaming technologies.

  • Experience with annotation workflows or experiment-tracking tools.

  • Experience optimizing data pipelines for GPU-backed training or large-scale inference.

Benefits:
  • Paid Vacation
  • Sick Days
  • Floating Holidays
  • Sport/Insurance Compensation
  • English Classes
  • Charity
  • Training Compensation
Apply for this job
Benefits:
  • Paid Vacation
  • Sick Days
  • Floating Holidays
  • Sport/Insurance Compensation
  • English Classes
  • Charity
  • Training Compensation
Similar Vacancies
Looking for More Opportunities?
Explore similar open positions that match your experience.