Meltplan

AI Systems QA Engineer

Bengaluru, IndiaFull timePosted 20 days ago
Apply on Meltplan →

Sign into see who you know at Meltplan.

MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI systems. Perform end-to-end system testing, regression testing, and performance testing for AI-driven applications. Validate model outputs for accuracy, consistency, safety, hallucination detection, and edge cases. Build automated test pipelines and quality benchmarks for AI systems. Collaborate closely with AI/ML engineers, product teams, and platform engineers to improve system reliability. Analyze failures, identify root causes, and provide actionable feedback to improve model behavior. Develop datasets, prompts, and testing scenarios to measure model performance across multiple use cases. Monitor production performance and continuously improve evaluation metrics and testing standards. Ensure compliance with responsible AI and quality assurance best practices. What We're looking for:  Bachelor’s degree in Computer Science, Engineering, or related field 5–8 years of experience in QA/testing, preferably in AI/ML or data-driven systems Strong experience in AI/LLM evaluation framewo...