DeskForge builds a 1.2M-observation desktop dataset to train computer-use agents

Computer-use agents have to pick the right target in cluttered desktop scenes, where several applications, overlapping windows and look-alike controls compete for attention. The authors of this paper say existing training data rarely pairs such scenes with dense annotations or varies them in a controlled way.
Their answer is DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision. It varies application states, content, window layout, appearance and resolution. It fuses screenshots, accessibility trees and window geometry into dense element annotations, and it records the outcome of each action it executes.
From this environment the authors built DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. They then fine-tuned four vision-language models on 200K grounding examples drawn from the corpus.
All four models improved across held-out desktop conditions and on all five external GUI grounding benchmarks. For Qwen3.5-4B, accuracy rose by 11.51 percentage points on ScreenSpot-Pro and by 10.11 points on OSWorld-G.
The gains also carried over to long-horizon tasks. Under a fixed planner, the fine-tuned action models solved more tasks on WebArena-Infinity and OpenApps. Qwen3.5-4B went from 31 to 50 of 119 tasks on WebArena-Infinity and from 3 to 15 of 100 tasks on OpenApps.
The authors conclude that controllable composition of real desktop environments is a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset and the fine-tuned model are available from the project page.
Key facts
- DeskForge composes and explores real applications, varying application states, content, window layout, appearance and resolution, and fuses screenshots, accessibility trees and window geometry into dense element annotations.
- The resulting DeskForge-1M corpus holds 1.2M annotated desktop observations with 159.7M element instances.
- Four vision-language models were fine-tuned on 200K grounding examples from the corpus; all four improved on all five external GUI grounding benchmarks.
- Qwen3.5-4B gained 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G.
- Under a fixed planner, Qwen3.5-4B solved 50 of 119 WebArena-Infinity tasks (up from 31) and 15 of 100 OpenApps tasks (up from 3).
Why it matters
Grounding, meaning locating the correct control on screen, is a basic skill for any agent that operates a desktop. The authors argue that training data for it is the bottleneck, because existing data rarely pairs complex scenes with dense annotations or varies them in a controlled way. DeskForge tackles that by generating the data from real applications with controlled variation, and the reported results suggest the approach helps both grounding and longer multi-step tasks.
Who it affects
Researchers and teams training or fine-tuning vision-language models for computer use are the direct audience, since the work offers a data source and a way to produce more of it. Anyone benchmarking agents on ScreenSpot-Pro, OSWorld-G, WebArena-Infinity or OpenApps will find these as reference numbers.
How to use it
The authors say the framework code, the dataset and the fine-tuned model are available from the project page at https://saidgurbuz.github.io/deskforge/. The paper's recipe is to draw grounding examples from DeskForge-1M (200K were used) and fine-tune a vision-language model on them. No compute cost, training time or licence is stated in the source.
How solid is it
The claims come from the authors' own abstract. The evidence is broad in one respect: all four fine-tuned models improved on all five external GUI grounding benchmarks and across held-out desktop conditions. The detailed numbers, however, are given only for Qwen3.5-4B. The other three models are not identified and their gains are not quantified. Baseline and absolute accuracy on ScreenSpot-Pro and OSWorld-G are not given, only the gains in points. The task-completion results are small-sample counts (119 and 100 tasks) under a fixed planner that is not named.
Risks and caveats
Results rest on one abstract, so details such as which applications DeskForge composes, which other models were tuned and how the planner works are not available here. The 11.51 and 10.11 figures are absolute gains in points, not relative improvements, and they cannot be turned into final accuracy without the baselines. The task-completion gains, such as 3 to 15 of 100 on OpenApps, start from a low base and come from a single model. Whether the released fine-tuned model is Qwen3.5-4B or another one is not stated.
“controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use”
— DeskForge paper abstract