Role Overview
Scale AI is hiring a Senior Software Engineer, Agent Oversight. This is a full-time role in San Francisco, CA; New York. Part of Scale AI's Fullstack hiring, posted last week. Full responsibilities, required qualifications, and the apply link are listed in the description below.
Salary Context
Salary is not disclosed in this posting. Market median for Senior-level Fullstack roles is $149k-$196k (based on 94 comparable listings). Many employers share specifics during the interview process or after an initial screen.
Resume Keywords to Include
Make sure these keywords appear in your resume to improve ATS scoring
Job description
About Scale
Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust.
About the Team
Applied Intelligence Systems team is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power Agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like.
About the Role
As a Software Engineer on Agent Oversight, you will build the platform infrastructure that lets our production agents be observed, evaluated, and improved at scale. This includes building observability tooling, evaluation harnesses, and the pipelines that connect them to improvement loops. Whether building foundational infrastructure or partnering closely with ML engineers on production workflows, you will own your systems end-to-end while maintaining rigorous technical standards.
You will:
- Design and build core platform capabilities for deploying, monitoring, and evaluating agentic applications in production
- Build reliable APIs and data pipelines that capture agent telemetry, evaluation signals, and performance metrics at scale
- Work alongside ML engineers where platform work intersects with evaluation or improvement systems — bringing enough ML fluency to reason about model behavior, evaluation quality, and improvement loops while owning the software systems that make those workflows reliable
- Own the reliability, scalability, and observability of platform components serving multiple concurrent enterprise and government customers
- Work cross-functionally with product, forward deployed engineering, and customers to translate real-world deployment requirements into platform features
- Build features end-to-end: system design, implementation, debugging, and testing
- Participate in high-velocity experimentation to validate platform capabilities against real customer usage
Requirements:
- 4+ years of professional software engineering experience, with strong fundamentals in backend/distributed systems, APIs, and data pipeline design
- Hands-on experience building production software for ML/LLM-powered products or platforms, such as evaluation systems, observability/monitoring, experimentation infrastructure, agent runtimes, model-serving-adjacent services, or telemetry/data pipelines
- Working knowledge of how LLM or ML systems behave in production: evaluation signals, failure modes, prompt/tool-calling workflows, experiment results, data quality issues, and the tradeoffs between offline evals and live customer behavior
- Experience partnering closely with ML engineers or applied researchers to turn prototypes, eval loops, or model-improvement workflows into reliable platform capabilities, without needing to own model training, modeling strategy, or research direction
- Experience building infrastructure or platforms that other engineering teams build on top of (internal platform, developer tools, or similar)
- Track record of taking ownership of features or components end-to-end — from design through production — within a larger platform or system
- Comfortable operating in an ambiguous, fast-changing domain where tooling and best practices are still being defined
- Strong problem-solving skills and the ability to work independently or as part of a tight-knit, cross-functional team
- Excited to work directly with ML engineers and customer-facing teams, including challenging assumptions in designs and metrics when platform behavior, model behavior, and customer needs intersect
- Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow
Nice to have:
- Deep experience building or maintaining observability, monitoring, or evaluation systems for ML/LLM-powered products in production
- Familiarity with agent architectures — tool use, planning, multi-agent orchestration
- Exposure to MLOps, feature stores, model serving, or experiment infrastructure
- Experience working in regulated or enterprise contexts
- Experience reviewing others’ technical designs or mentoring engineers at a senior/staff level
Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend.
PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.
About Us:
At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications.
We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at accommodations@scale.com. Please see the United States Department of Labor's Know Your Rights poster for additional information.
We comply with the United States Department of Labor's Pay Transparency provision.
PLEASE NOTE: We collect, retain and use personal data for our professional business purposes, including notifying you of job opportunities that may be of interest and sharing with our affiliates. We limit the personal data we collect to that which we believe is appropriate and necessary to manage applicants’ needs, provide our services, and comply with applicable laws. Any information we collect in connection with your application will be treated in accordance with our internal policies and programs designed to protect personal data. Please see our privacy policy for additional information.
About Scale AI

Scale AI
scaleai.org
141 other open roles at Scale AI on TryApplyNow.
Frequently Asked Questions
How do I apply for the Senior Software Engineer, Agent Oversight position at Scale AI?
Use the Apply button above to submit your application directly to Scale AI. Most applications take less than 5 minutes if your resume and contact details are ready, and you'll be routed to the employer's official application system to finish.
Where is the Senior Software Engineer, Agent Oversight position at Scale AI located?
This position is based in San Francisco, CA; New York. Scale AI has not indicated remote or hybrid options for this role, so candidates should plan for on-site work.
What does a Senior Software Engineer, Agent Oversight at Scale AI earn?
Scale AI has not disclosed a salary range in this posting. Many employers share specifics later in the interview process; you can also ask during a recruiter screen if compensation transparency is important to you.
When was the Senior Software Engineer, Agent Oversight role at Scale AI posted?
This role was posted on July 14, 2026 (8 days ago). It's still listed as actively hiring; we re-confirm openings against the source system multiple times per day and remove closed roles.
How much experience does the Senior Software Engineer, Agent Oversight role at Scale AI require?
This is a senior-level position. Most senior roles call for 5+ years of directly relevant experience. Scale AI lists their specific requirements in the description below, so review the must-have qualifications closely before applying.
Similar Jobs
Lead Full Stack Developer - Python (Global Security)
Royal Bank of Canada
Senior Full Stack Developer, Vice President
BlackRock
Software Engineer
FanDuel
Deployments Software Engineer
Physicalintelligence
Senior Fullstack Software Engineer, Growth
Algolia
More Jobs at Scale AI
View all →Field Engineer, Public Sector
Scale AI
SWE Fellow - Human Frontier Collective (Canada)
Scale AI
Machine Learning Fellow - Human Frontier Collective (Canada)
Scale AI
Staff Machine Learning Research Engineer, Agent Post-training - Enterprise GenAI
Scale AI
Subject Matter Expert
Scale AI
AI-powered job search
Get every job scored to your resume
Upload your resume and get jobs ranked, your resume tailored, and employee contacts found automatically.
Get started freeNo credit card to start