Machine Learning Engineer - Data Pipeline
About us
At Articul8 AI, we relentlessly pursue excellence and create exceptional AI products that exceed customer expectations. We are a team of dedicated individuals who take pride in our work and strive for greatness in every aspect of our business. We believe in using our advantages to make a positive impact on the world and inspiring others to do the same.
Job Description
We are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis. You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end. You will work closely with other researchers and engineers to empower our next generation of domain-specific models. We value rapid prototyping, iterating, and shipping new systems quickly.
Required Qualifications
- BS/MS/PhD in Computer Science or a related field.
- Proficiency in at least one deep learning framework, such as PyTorch.
- Experience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem.
- Strong expertise in large stateful distributed systems and data processing.
- Strong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes).
- Proficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code.
- Excellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality.
Key Competencies
- Active Github contributions are a big plus.
- Experience in building large-scale datasets.
- Familiar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, S...