Altruist

Staff Back End Engineer, Evals - Hazel AI

San Francisco, United StatesFull timeStaff$275,000 - $325,000 / yearPosted 20 days ago
Apply on Altruist →

Sign into see who you know at Altruist.

About Altruist Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time and resources, and deliver superior outcomes for their clients. We're looking for exceptional talent to help us achieve our mission of making financial advice better, more affordable, and accessible to all. If you're passionate about challenging the status quo and want to do the most important work of your life, we'd love to meet you! But first, our values Kindness - Kindness doesn’t just equal niceness. We listen to understand. We embrace, and encourage healthy debate and diverse perspectives. We approach conflict openly, honestly, and respectfully.Brilliance - Humility is the skill we’re most proud of and possessing a growth mindset is always top of mind. We take ownership in everything we touch; regularly using our unique superpowers to reach a common goal as a team. We succeed and fail as one.Grit - When challenges arise, we stay laser focused on achieving our mission and finding a way forward, even when it’s hard. We are nimble and maintain a sense of urgency, swiftly adapting to change and overcoming obstacles.About Hazel: Hazel.ai is building the AI engine for wealth management that unlocks 10x growth, efficiency and value for financial advisors and their clients in a regulated industry. Since its launch last September, Hazel has organically and rapidly grown its user base. Hazel is a part of Altruist’s broader mission to make financial advice better, more affordable, and accessible to all. This role is hybrid, with four in-office days per week at our San Francisco FiDi location. The opportunity: Architect our evaluation platform from first principles – the observability, scoring, golden datasets, verification agents, and CI/CD integration that define standards of quality. You'll work shoulder-to-shoulder with backend engineers, product managers, and a growing bench of subject matter experts, including practicing CFPs, CPAs, and tax planners, to translate fiduciary-grade requirements into automated quality signals. Your impact: Design and build Hazel's evals platform end-to-end – online scoring, offline benchmarks, regression suites, LLM-as-judge pipelines, and human-in-the-loop review workflows across every Hazel surface. Build production observability and monitoring for AI quality: hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows. Architect data curation pipelines that turn real advisor interactions into evaluation datasets – with rigorous sampling strategies, labeling protocols, dataset versioning, and the privacy and consent controls required for regulated finance. Build and steward Hazel's golden datasets in close partnership with SMEs and a network of prac...