ServiceNow CoreAI details AutoSynthData, synthetic training tasks for enterprise agents

ServiceNow CoreAI has published a Hugging Face blog post on AutoSynthData, a pipeline for generating training data for enterprise agents. The starting point is a familiar problem: a model can be broadly capable and still stumble in one company's environment, with a workflow it handles poorly, a tool combination it misuses or a constraint it ignores. A single failure is informative, but training needs many new tasks that exercise the same capability in different situations. Those tasks must also be completable in the environment, resemble real requests, and come with a reliable way to check success.
AutoSynthData uses a target model's failures and a stronger teacher's successes to decide what the model should learn next, then generates and validates new tasks for those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. The team illustrates the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset.
The post defines a task as a triple: system specification, user prompt and verifier. The specification holds the constraints the agent operates under: system instructions, environment policies and, where relevant, initial state such as a seeded database or a set of knowledge articles. The user prompt should be feasible (at least one trajectory exists that satisfies it within the environment and policy), realistic (something a user would plausibly ask) and difficult (it exposes a weakness of the current agent, since reliably solved tasks add little signal). The verifier should be consistent with the prompt, specification and state; sound, meaning it rejects trajectories that fail the task or break constraints; and complete, meaning it accepts valid solutions rather than encoding one reference trajectory. A lax verifier can reward wrong behavior, and an overly strict one can penalize valid solutions.
The process begins by evaluating the target model on diagnostic tasks. In the EnterpriseOps Gym experiment, both the target model and a stronger teacher are run on evaluation tasks, and the runs are examined for the capability being tested, the tools and workflow involved, where the target fails and how the teacher succeeds, the properties a correct final state must satisfy, and the dimensions that can vary while preserving the capability. These findings are distilled into sanitized capability specification cards. The generator receives only the cards, not the original prompts, entities, trajectories or verifier details, and uses them to create new tasks with different prompts, states and solution paths. It varies entities, initial state, workflow composition, tool combinations, wording and difficulty, and the teacher then demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), those demonstrations teach the target model to apply the capability in new situations.
Dataset construction has two phases. In the target phase, workers generate independent tasks in parallel from the capability specifications; each candidate goes through validation, execution, solver evaluation and repair before acceptance. In the multiply phase, accepted target samples are expanded into novel variants, each with its own user request, environment state, entity configuration, reference trajectory and verifier, and each must pass the same checks. A multiplied sample cannot seed another multiplied sample, which anchors expansion to the vetted set and limits drift. Architecturally, a shared controller coordinates generation, quality control, coverage and dataset construction, while an environment-specific adapter handles execution, task and state management, reference replay, deterministic verification, solver execution and task profiling.
Quality control works at two levels. For each sample, solver evaluation measures difficulty. In the configuration used, the team favors tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three. A positive gate runs the reference trajectory in the target environment and checks the resulting state against the candidate's verifier, exposing mismatches among prompt, initial state, solution and success criteria. A negative gate can mutate parts of the expected outcome and confirm those states no longer pass verification, which catches weak verifiers that award success without the intended behavior. Candidates that fail go to a critic before being discarded. The critic looks for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic or a mismatch with the intended capability, and its findings guide a targeted repair under a fixed retry limit. The repaired task must pass the gates again.
At batch level, a meta-review examines accepted samples, rejected samples and generation behavior. It asks which task families are overrepresented, which capability dimensions are missing, whether the same kinds of examples keep appearing, whether particular targets keep failing generation, whether critiques show systematic problems, and what guidance should change for the next batch. The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions and directs more work toward gaps.
The post frames the whole loop as moving the training frontier: AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary, hard enough to expose weaknesses but solvable enough for the teacher to give reliable demonstrations. After post-training, the updated model is evaluated in the same environment; tasks it now solves reliably become less useful, and persistent failures guide the next round of generation.
Key facts
- AutoSynthData, built at ServiceNow CoreAI, uses a target model's failures and a stronger teacher's successes to generate and validate new training tasks for enterprise environments.
- A task is a triple of system specification, user prompt and verifier; the generator sees only sanitized capability specification cards, never the original evaluation prompts, trajectories or verifier details.
- Data is built in two phases, target and multiply, and a multiplied sample cannot seed another multiplied sample, which limits drift.
- In the configuration used, the team favors tasks the target model solves on at most one of three trials and the stronger solver solves on at least two of three.
- Each candidate passes positive verification (reference trajectory replay) and negative verification (mutated outcomes must fail), plus a critic-guided repair loop with a fixed retry limit; batch-level meta-review balances coverage and diversity.
Why it matters
Enterprises need agents that work in their own systems, under their own rules and with their own data, and a generally capable model can still fail there. The hard part, as the post frames it, is converting observed weaknesses into many new, executable, verifiable tasks. AutoSynthData's answer is to aim generation at the edge of what the target model can do: tasks the model fails but a stronger teacher can solve, with the curriculum shifting as the model improves. Its emphasis on checking every candidate (does the reference solution run, does the verifier reject wrong outcomes) addresses the risk that a lax verifier rewards incorrect behavior during training.
Who it affects
The audience is teams that want to adapt a model to a specific enterprise environment through post-training, particularly with SFT, where the stronger teacher's demonstrations become the training examples. It also speaks to anyone building verifiers and agentic environments, since the post lays out explicit properties for tasks (feasibility, realism, difficulty) and for verifiers (consistency, soundness, completeness). The illustration uses EnterpriseOps Gym (Malay et al., 2026) and its released dataset.
How to use it
The post presents a method and its design rather than a how-to. Its ingredients are an environment with executable state and a deterministic verifier, a target model, a stronger teacher or solver, and diagnostic evaluation tasks. From those, the evaluation runs are distilled into capability specification cards, the generator produces tasks from the cards, and accepted samples are used for post-training. The updated model is then re-evaluated to guide the next round. The architecture splits a shared controller from an environment-specific adapter, so adapting it to a new environment means supplying the adapter functions the post lists: execution, task and state management, reference replay, deterministic verification, solver execution and task profiling.
How solid is it
The design is described in concrete, checkable terms: the task abstraction, the two construction phases, the positive and negative gates, the critic and repair loop, and the batch-level meta-review. The one numeric detail is the difficulty filter of at most one of three target trials solved and at least two of three for the stronger solver. The portion of the post available here contains no experimental results, benchmark scores, accuracy gains or before and after comparisons, so the effect of the pipeline on model performance cannot be assessed from it. The names of the target and teacher models, the number of generated tasks and the dataset size are also not given in the visible text.
Risks and caveats
The visible text does not say how many repair retries are allowed, only that the limit is fixed. Difficulty is defined relative to a particular target model and teacher, so the filter is tuned to those models. The method depends on a stronger teacher being able to solve the tasks and on the environment offering deterministic verification and reference replay. The post notes that individually valid samples can still form a repetitive or unbalanced dataset, which is why batch-level review exists. The text is cut off at the sentence 'Our experiments focus on SFT, bu', so what the post says about reinforcement learning or other training methods is not visible.
“AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary”
— ServiceNow CoreAI, AutoSynthData blog post on Hugging Face