Articul8

Machine Learning Engineer - Data Pipeline

Dublin, California, IrelandFull timePosted about 1 month ago
Apply on Articul8 →

Sign in to see who you know at Articul8.

About us

At Articul8 AI, we relentlessly pursue excellence and create exceptional AI products that exceed customer expectations. We are a team of dedicated individuals who take pride in our work and strive for greatness in every aspect of our business. We believe in using our advantages to make a positive impact on the world and inspiring others to do the same.

Job Description

We are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis. You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end. You will work closely with other researchers and engineers to empower our next generation of domain-specific models. We value rapid prototyping, iterating, and shipping new systems quickly.

Required Qualifications

  • BS/MS/PhD in Computer Science or a related field.
  • Proficiency in at least one deep learning framework, such as PyTorch.
  • Experience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem.
  • Strong expertise in large stateful distributed systems and data processing.
  • Strong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes).
  • Proficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code.
  • Excellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality.

Key Competencies

  • Active Github contributions are a big plus.
  • Experience in building large-scale datasets.
  • Familiar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, S...