GUI-HARVEST optimizes GUI agent harnesses without retraining the model

GUI-HARVEST optimizes GUI agent harnesses without retraining the model

A paper introduces GUI-HARVEST, an automatic optimizer for the executable harness that surrounds a GUI model. That harness decides how observations are assembled, how actions are executed, and how verification, recovery and termination are controlled. The idea is to let a GUI agent improve itself while the backbone model stays frozen: the harness code changes, the model does not.

The authors say that optimizing this harness automatically is harder for GUI agents than for non-GUI agents, because of three coupled challenges. First, model intent has to be reconciled with the visual effects actually observed on screen. Second, failures have to be diagnosed even though execution outcomes vary from run to run. Third, recurrent failure patterns have to be found across tasks and turned into reusable runtime changes.

GUI-HARVEST answers each challenge with one component. To ground diagnosis in what actions really did, it aligns model outputs and executed actions with before-and-after screenshots, so each finding is tied to a specific interface transition. To handle execution variability, it treats repeated runs of the same task as a joint evidence unit and uses within-task comparisons to locate behavioral differences that matter for the outcome. Finally, it consolidates verified findings across tasks into recurring failure patterns and maps them to bounded source-code edits, with predictions recorded before evaluation. It then checks the predicted behavioral effects alongside task performance through repeated execution.

On OSWorld-Verified, the authors report consistent held-out gains across six backbone models: general-purpose open, GUI-specialized open, and proprietary. One example given is Qwen3-VL-32B-Instruct, which gains 12.33 points on the full suite. In a transfer test, the frozen harness improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps, without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms two baselines, Self-Harness and Meta-Harness. The authors read this as suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is on GitHub.

Key facts

  • GUI-HARVEST automatically optimizes the executable harness around a GUI model while the backbone model stays frozen.
  • It has three parts: before-and-after screenshot alignment, repeated runs of one task as joint evidence, and cross-task failure patterns mapped to bounded source-code edits with predictions recorded before evaluation.
  • On OSWorld-Verified it reports consistent held-out gains across six backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite.
  • The frozen harness transfers to WindowsAgentArena, improving GPT-5 by 13.87 percentage points at 50 steps without further optimization.
  • With the same backbone and initial harness, it outperforms Self-Harness and Meta-Harness; code is available on GitHub.

Why it matters

This paper targets the layer around the model: the harness code that assembles observations, executes actions, and controls verification, recovery and termination. GUI-HARVEST improves that layer automatically while the backbone stays frozen, which means the same approach is tested on open, GUI-specialized and proprietary models. Its specific twist is grounding diagnosis in before-and-after screenshots, so that a proposed fix traces back to an observed interface transition rather than to a guess about what the model meant.

Who it affects

Teams building GUI or computer-use agents on top of existing models are the direct audience, especially those who cannot or do not want to fine-tune the backbone. Researchers working on harness optimization for agents are also affected, since the paper positions itself against Self-Harness and Meta-Harness.

How to use it

The code is available at https://github.com/GaryYang12345/GUI-HARVEST. The workflow it describes starts from a backbone and an initial harness, runs tasks repeatedly, and produces bounded source-code edits that are validated by repeated execution. Evaluation in the paper uses OSWorld-Verified and WindowsAgentArena.

How solid is it

The results are reported by the paper's authors and rest on two benchmarks: OSWorld-Verified for held-out gains and WindowsAgentArena for transfer. The gains span six backbone models, which is broader than a single-model demonstration. The comparison with Self-Harness and Meta-Harness uses the same backbone and initial harness. The authors frame the generalization explanation as a suggestion, not a proof.

Risks and caveats

The 12.33 gain is stated in 'points' without saying whether it is percentage points or relative, and no baseline scores are given for Qwen3-VL-32B-Instruct or GPT-5. The margins by which GUI-HARVEST beats Self-Harness and Meta-Harness are not given. The names of the six backbone models are not listed except Qwen3-VL-32B-Instruct and GPT-5, and GPT-5's membership in the six is not stated. No compute cost, number of optimization iterations or number of repeated runs is given, so the practical expense of the method is unclear.

“We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models.”

— GUI-HARVEST paper abstract