Research Engineer, QC Automation
ABOUT HUD
HUD https://www.hud.ai/ is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
ABOUT THE ROLE
We're looking for Research Engineers to automate QC for training data created by companies using HUD’s infrastructure. You’ll build the systems that scale quality to help us meet our continued strong demand.
RESPONSIBILITIES
- Create QC systems based on true understanding and human judgement, without relying heavily on LLMs
- Define and enforce quality standards for training data
- Design experiments and metrics to grade agent outputs
- Partner with data vendors to debug quality issues and diagnose agent failure modes, provide actionable feedback, and improve their data generation processes
- Translate QC learnings into systems for auditing supplier-generated datasets, including sampling strategies, validation pipelines (rule-based and model-assisted), and feedback loops
- Continuously integrate QC learnings into infrastructure tools and data vendor portal to reduce anomalies, inconsistencies, and edge cases
EXPERIENCE
You may be a good fit if you have:
- Proficiency in Python, Docker, and Linux environments
- Strong understanding of what “good data” means and how to measure it
- Built scalable data validation pipelines and automated QA/QC systems end-to-end without a fully prescribed roadmap
- Experience working on benchmarks and evals - you can reason about what makes a task realistic, a rubric reliable, an environment usable, and a trajectory useful for RL training
- Early-stage startup experience with ability to work independently in fast-paced environments
Strong candidates may also:
- Be detail-oriented and able to spot subtle inconsistencies or edge cases in data
- Be comfortable d...