ServiceNow AutoSynthData Targets Training for Enterprise AI Agents
ServiceNow introduced AutoSynthData, a pipeline for creating and validating synthetic training tasks based on failures of the target model.

ServiceNow AutoSynthData is a method for creating synthetic training tasks for enterprise AI agents, described by the company in a ServiceNow-AI post on Hugging Face dated October 2, 2026. The method is intended to target specific model failures rather than collecting large amounts of manually prepared data. In the EnterpriseOps Gym experimental benchmark, ServiceNow reports improved success rates after synthetic fine-tuning.
Enterprise agents do not work only with text. In many scenarios, they use tools and change system states, for example in IT service workflows. This increases the importance of tasks where it is possible to clearly verify whether an agent followed the required procedure and achieved the correct result.
ServiceNow AutoSynthData Creates Tasks from Model Failures
The pipeline uses two inputs: failures of the target model and successful solutions from a stronger teacher model. Based on these, it creates new tasks intended for training, specifically supervised fine-tuning (SFT).
Each created task contains a system specification, a user request, and a verifier. It is therefore not merely generating prompts and expected textual responses. The pipeline checks candidate tasks by executing a reference solution and running negative verifier tests. The goal is to establish that the solution works and that the verifier does not accept incorrect results.
This approach focuses on executable tasks with deterministic outcome checks. For agents in stateful environments, this is particularly important because an answer that sounds correct does not necessarily mean that the operation was correctly performed in the system.
Results in EnterpriseOps Gym
ServiceNow evaluated the method in the EnterpriseOps Gym benchmark. In the Hybrid environment, the authors report that Pass@1 increased from 63.01% to 68.55% after synthetic SFT. In the ITSM environment, it rose from 18.77% to 27.18%.
Pass@1 expresses success on a single attempt. The figures therefore indicate the model’s result after fine-tuning on that evaluation, not a general guarantee of reliability in enterprise operations.
EnterpriseOps Gym is a publicly described benchmark of stateful enterprise environments. Its authors list 1,150 tasks and 512 tools across eight areas. The benchmark is designed to evaluate agents that must work with tools and progressively modify the environment’s state.
The Benefit Is So Far Limited to Controlled Tests
AutoSynthData addresses a practical problem in agent fine-tuning: preparing high-quality tasks, reference solutions, and reliable evaluation rules is difficult. If training data is generated according to specific model failures, fine-tuning may be more targeted than when using generally collected data.
However, the published results come from the method’s authors and from a controlled benchmark. They are not independently replicated evidence that the improvement will transfer in the same way to production enterprise systems. The announcement also does not confirm that AutoSynthData is available as publicly usable code or a commercial product.
It also remains open how the method will perform with other models, sensitive corporate data, and workflows involving real operational, security, and permission-related constraints.
What to Watch Next
- whether ServiceNow will release an implementation, training data, or reproducible AutoSynthData configurations,
- whether independent replications outside the EnterpriseOps Gym benchmark will confirm the results,
- how the method performs in environments closer to production, including data-sensitive workflows and permission-controlled workflows.
Sources
- Hugging Face Blog — ServiceNow-AI – Confirms the AutoSynthData announcement, pipeline description, and the experiment results reported by the authors in Hybrid and ITSM.
- arXiv — EnterpriseOps-Gym – Describes the EnterpriseOps Gym benchmark, its scope, and the limitations of current agents in stateful enterprise tasks.
- Hugging Face Datasets — ServiceNow-AI/EnterpriseOps-Gym – Confirms the publicly available benchmark dataset card and its connection to the arXiv paper.
Verified and updated: 10/02/2026 06:24



