AI Research Engineer
Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications. Since public launch, Emergent has crossed $100M in ARR and grown to over 10M users across 190+ countries. We're backed by Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator. We're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox. We're hiring builders who want ownership, speed, and impact at global scale. The Role: We're looking for a Research Engineer to characterize, measure, and advance the capabilities of our coding agents. You will turn ambiguous notions of "agent quality" into clear, defensible metrics that the team, leadership, and the field can rely on, and you will use those metrics to drive both incremental wins and moonshots in agent performance. This is a deep-work role at the intersection of agent behavior, evaluation research, and applied training. You will define what good looks like for long-horizon coding agents, build the evaluation dataset and methodology that produces those signals, mine production data for failure modes most teams never see, and run targeted training, fine-tuning, RL, memory, and prompt-optimization experiments that translate research advances into shipped improvements. You will operate with strong independence, make hard calls in inherently subjective and probabilistic systems, and own outcomes end-to-end. If you treat models as objects of study rather than black boxes, take pride in moving benchmark numbers with rigor, and want to apply the frontier of agent research at the scale of millions of real applications, this is your role.What You'll Do: Architect the next version of the Emergent agent. Shape the core architecture and make the foundational design choices that define how the agent thinks, learns, and improves over time. Characterize agent behavior at depth. Develop a deep, evidence-grounded understanding of how the agent succeeds and fails across the full range of real-world usage, and convert that understanding into rigorous, quantitative measurement. Design and ship evaluations across reasoning, planning, tool use, code correctness, long-horizon execution, security, and agent reliability. Define the metric, build the dataset, validate against known signals, and ship dashboards that make regressions impossible to miss. Drive step-function gains. Take on the ambitious bets that meaningfully advance the state of the art, the 10-point leaps on hard capabilities, not incremental polish. Pick the problems where the upside is large and the path is uncertain. Climb public benchmarks. Mo...