Staff Research Engineer/Scientist
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Our Core AI Research team develops novel methods for enterprise agents that reason over multimodal information, use tools, take reliable action across stateful workflows, and improve through feedback. We work across LLM model post-training, agent harnesses, training environments, evaluations, ML, search and reasoning systems, partnering closely with product, engineering, infrastructure, security, and domain experts.
About the role
As a Staff Research Scientist, you will independently lead a major workstream in agent learning and recursive self-improvement. You will turn systematic failures and successful trajectories into hypotheses, experiments, training signals, and deployable improvements to model weights and/or the executable harness around the model.
This is a research role for someone who can move between scientific reasoning, training code, agent systems, and production constraints.
What you get to do in this role
- Design and execute end-to-end research projects that improve long-horizon enterprise agents across planning, reasoning, memory, tool use, retrieval, computer use, multi-agent coordination, and verification.
- Research model post-training methods such as continued pretraining, supervised fine-tuning (SFT), RL, DPO/GRPO, reward modeling, and distillation.
- Research harness-level optimization across prompts and task framing, tool and schema design, skills, MCP-backed providers, subagents, context and memory management, agent-loop policy, and reliable verifiers.
- Build improvement flywheels that mine trajectories and production-safe signals, identify recurring failure modes, generate or curate data, propose interventions, and measure generalization before promotion.
- Create realistic, stateful training environments and benchmarks for enterprise workflows, with programmatic verifiers and calibrated human or model-based graders where deterministic grading is not possible.
- Run rigorous ablations and scaling experiments; reason explicitly about variance, contamination, reward hacking, distribution shift, cross-model transfer, cost, and latency.
- Develop capabilities across one or more modalities - language, documents, images/video, and speech/audio - and across multilingual or cross-lingual settings.
- Build reproducible distributed pipelines for training, rollout generation,...