ServiceNow · Posted 10 days ago
Senior Research Scientist, Agent Evaluation
The posting
key requirement, as the employer wrote it
We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces.
You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
What you'll do Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation Take problems from research question to prototype to shipped feature, owning them end to end Build and harden the pipelines and scoring logic behind customer-facing evaluation Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels
What we're looking for 5+ years in ML, applied AI, prompt engineering, agentic AI, including shipping something real to users Strong Python skills Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context Experience designing evaluation methodologies, not just running evaluations Hands-on production work with LLM APIs — prompt engineering, structured output, cost and latency tradeoffs Clear communication with technical and non-technical audiences Good to have Experience with AI-assisted development tools (Claude code, Windsurf, or similar)
ServiceNow
- Open roles in India
- 58
- Hiring in
- Hyderabad, Bengaluru, Mumbai, Gurugram
- Applications through
- SmartRecruiters
Counted from the roles we read off ServiceNow's own hiring page today.
ServiceNow
American technology company
- Founded
- 2004
- Headquarters
- Santa Clara
- Employees
- 10,371
- Industry
- enterprise software
- Revenue
- $3.5B (2019)
- Stock market
- Listed on New York Stock Exchange
Facts from Wikidata, the open, community-edited database behind Wikipedia — check the link if something looks out of date. Funding rounds, investors and employee ratings are not shown: no free source carries them reliably.