Skip to main content
TryApplyNow logo
ServiceNow logo

Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations

ServiceNow
Full Timemanager
Santa Clara, California, United StatesPosted 9 days ago

Role Overview

ServiceNow is hiring a Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations. This is a full-time role in Santa Clara, California. Part of ServiceNow's Risk hiring, posted last week. Full responsibilities, required qualifications, and the apply link are listed in the description below.

Salary Context

Salary is not disclosed in this posting. Market median for Manager-level Risk roles is $140k-$187k (based on 96 comparable listings). Many employers share specifics during the interview process or after an initial screen.

Resume Keywords to Include

Make sure these keywords appear in your resume to improve ATS scoring

PythonSQLPandasNumPySaaSDriftROIOR

Job description

The Advanced Technology Group (ATG) at ServiceNow is a customer-focused innovation group building intelligent software and smart user experiences using existing and latest advanced technologies to enable end-to-end, industry-leading work experiences for customers. We are a group of researchers, applied scientists, engineers, and product managers with a dual mission. We build and evolve the AI platform, and partner with teams to build products and end-to-end AI-powered work experiences. In equal measure, we lay the foundations, research, experiment, and de-risk AI technologies that unlock new work experiences in the future.

Job Description

We are seeking an exceptional, data-driven Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations to establish and lead AI evaluation practices for both ServiceNow and our customers. As ServiceNow shifts enterprise workflows from simple generation to complex, autonomous agents, ensuring system reliability, safety, and accuracy is paramount.

In this role, you will lead a specialized team of AI evaluation engineers and data scientists. Your team will build the infrastructure, rigorous validation frameworks, and benchmarks that quantify the performance of Now Assist agentic workflows across multi-step orchestration, tool-calling, and enterprise-grounded reasoning. You will bridge the gap between frontier AI research and hard production metrics, directly impacting the trust and adoption of autonomous workflows for millions of enterprise users.

 

What You Get To Do In This Role

  • Build the Evaluation Infrastructure: Design, own, and scale automated testing and evaluation harnesses (unit evals, integration evals, and production drift monitors) to measure agent quality and eliminate regressions.
  • Define Enterprise AI Benchmarks: Create standard, repeatable evaluation frameworks tailored to complex business workflows—assessing multi-agent orchestration, intent routing, multi-step planning loops, and long-term memory accuracy.
  • Validate Grounding & RAG Pipelines: Partner with search and data fabric teams to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines, hybrid search, and semantic re-ranking systems.
  • Model Selection Optimization: Rigorously benchmark frontier LLMs (e.g., OpenAI, Anthropic, Google, and proprietary ServiceNow models) to evaluate trade-offs across execution capabilities, latency, context-window efficiency, and inference costs.
  • Lead a High-Performing Team: Recruit, mentor, and foster an AI-native engineering team, driving engineering best practices, prompt-infrastructure stability, and production-grade rigor.
  • Cross-Functional Leadership: Collaborate with Core Product, Machine Learning Platforms, and Engineering leads to translate baseline performance statistics into actionable product improvements and model fine-tuning targets.

 

To be successful in this role you have:

  •  8+ years of professional software engineering or machine learning experience, including 3+ years managing or technically leading high-performing AI/ML teams.
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • Strong foundational knowledge of frontier AI SDKs and deep experience deploying or testing agentic/probabilistic software architectures (multi-agent orchestration, tool execution, and probabilistic feature deployment).
  • Demonstrated experience implementing rigorous AI metrics (e.g., ROUGE, BLEU, G-Eval, LLM-as-a-judge patterns, and custom deterministic evaluation code) at an enterprise scale.
  • Proficiency in Python and familiarity with data analytics infrastructures (SQL, Pandas, NumPy) alongside standard MLOps tracking platforms.
  • Experience with complex knowledge infrastructure, SaaS platform architectures, or relational datasets (e.g., Knowledge Graphs, CMDBs).
  • Ability to translate deeply technical evaluation data into executive-level risk assessments, ROI summaries, and strategic roadmap recommendations.
  • Bachelor’s or higher degree in Computer Science, Data Science, Machine Learning, or a highly quantitative field (Master's or Ph.D. is a plus).

For positions in this location, we offer a base pay of $201,300 - $352,300, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.  

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

About ServiceNow

ServiceNow logo

ServiceNow

servicenow.com

RiskOn-site

241 other open roles at ServiceNow on TryApplyNow.

Frequently Asked Questions

How do I apply for the Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations position at ServiceNow?

Use the Apply button above to submit your application directly to ServiceNow. Most applications take less than 5 minutes if your resume and contact details are ready, and you'll be routed to the employer's official application system to finish.

Where is the Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations position at ServiceNow located?

This position is based in Santa Clara, California. ServiceNow has not indicated remote or hybrid options for this role, so candidates should plan for on-site work.

What does a Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations at ServiceNow earn?

ServiceNow has not disclosed a salary range in this posting. Many employers share specifics later in the interview process; you can also ask during a recruiter screen if compensation transparency is important to you.

When was the Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations role at ServiceNow posted?

This role was posted on July 13, 2026 (9 days ago). It's still listed as actively hiring; we re-confirm openings against the source system multiple times per day and remove closed roles.

AI-powered job search

Get every job scored to your resume

Upload your resume and get jobs ranked, your resume tailored, and employee contacts found automatically.

Get started free

No credit card to start