AutoSynthData: Generating Training Data for Enterprise Agents
ServiceNow CoreAI developed AutoSynthData to generate training data for enterprise agents, addressing specific capability gaps. This system identifies a target model's failures and a stronger teacher's successes to create new tasks, ensuring the model learns what it struggles with. AutoSynthData validates individual samples and reviews generation at the batch level to prevent repetitive or unbalanced datasets, as demonstrated with EnterpriseOps Gym.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 2, 2026, 04:01 UTC
IngestedOffset at this time: UTC+0Oct 2, 2026, 05:00 UTC
- Published
- Oct 2, 2026, 04:01
- Ingested
- Oct 2, 2026, 05:00
- Source type
- Official
- Tier
- First-party
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.
The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded.
At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym ( Malay et al., 2026 ), using the released dataset . We begin by describing the environment an agent operates in and what makes a task useful for training.
What makes a useful agentic task?
An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions.
A task is instantiated within this environment. We use the following abstraction:
task = (system specification, user prompt, verifier)
System specification
The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles.
The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty.
Agent-facing task
The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties.
Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.
Realism. The user prompt should resemble something a user would plausibly ask in the target environment. The space of executable behaviors is usually much larger than the space of realistic workflows.
Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.
Verifier
The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties.
Consistency. It should agree with the user prompt, the system specification, and the task-specific environment state.
Soundness. It should reject trajectories that fail to satisfy the task or violate relevant constraints.
Completeness. It should accept valid solutions rather than encode one particular reference trajectory.
These properties matter directly during training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions.
Overview
Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model.
AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals which gaps remain and can guide the next round of generation.