H Company releases Holo4 computer-use models in two sizes

H Company releases Holo4 computer-use models in two sizes

H Company has released Holo4, a new series of agentic models that interact with software through any available interface: GUIs, code, MCP and APIs. It comes in two sizes, 27B dense and 35B-A3B Mixture of Experts, and both are available on the H Models API. Alongside it, the company is releasing Holotron4 Nano, an updated version of its Holotron 3 small model.

The pitch is one model for every surface. Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, choosing whichever fits the task. H argues that most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while tool-calling models are stuck in front of an application with no API. Real business tasks mix these approaches. Holo4 runs on desktops, the web, Android, in a code sandbox and against business APIs; it is the same model in each case and is called the same way. The authors say it scores well on academic benchmarks but was built for real business workflows.

On results, H says Holo4 models improve significantly over their Qwen base and trail only the strongest closed models on long workflows. On OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. H adds that it gets there with orders of magnitude fewer parameters and at a much lower cost. On OSWorld 2.0 (desktop control) and AutomationBench (API use), the company says Holo4 competes with frontier models at a much lower cost per task. Reference scores of 70.2% for Opus 5 and 66.2% for GPT-5.6 Sol are given as max-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart. H open-sources every trajectory behind its public-benchmark scores; they can be replayed at trajectories.hcompany.ai or downloaded from Hugging Face.

The chart notes flag that the comparisons are loose. On OSWorld 2.0, costs are estimated from the input and output tokens of each run; Holo4 is priced at H Models API rates from a single run. Qwen3.8 27B uses its model-card score with cost from the tokens of H's own run at Alibaba Cloud list prices. Qwen3.6 35B-A3B is a single run in H's harness, priced at Alibaba Cloud list prices with cache hits at 20% of the input price. Other closed and open-weight points come from the official OSWorld 2.0 leaderboard, and releases, harnesses and task subsets differ. On AutomationBench v1.0.6, Holo4 and the two Qwen models were measured in H's internal harness, while other models use public-set scores with cost per task from the official leaderboard, which runs on the private set. H says it will report Holo4 on the private set once it is evaluated.

The post shows three example tasks, run with the same prompt and harness for Holo4 27B and its base model, Qwen3.8 27B. In FreeCAD, building a detailed Eiffel Tower model took Holo4 27B 84 calls and 1.3M tokens versus 60 calls and 1.0M tokens for Qwen3.8 27B. Building the H company logo took Holo4 94 calls and 1.5M tokens versus 118 calls and 1.9M for Qwen. In Godot, a self-playing Pac-Man-style game took Holo4 68 calls, 2.4M tokens and 268 lines, versus 197 calls, 11.4M tokens and 327 lines for Qwen.

On training, Holo4 was trained with supervised and reinforcement learning on a large set of environments and tasks, including those from H's Agentic Task Factory. That internal set of pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP. H also rebuilt its harness, the loop that executes the model's actions and manages context over hundreds of steps, using feedback from OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were a reliable memory that can keep track of hundreds of steps and a shell on the desktop machine itself.

Holotron4 Nano applies the same post-training recipe to NVIDIA's Nemotron 3 Nano Omni, as a follow-up to Holotron 3 and as a member of the NVIDIA Nemotron Coalition. H says it significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes, with gains measured as absolute percentage-point improvements. It argues the recipe transfers well and that nothing in it is size-specific.

Both Holo4 sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to Holotron4 Nano. H says it will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.

Key facts

  • Holo4 comes in 27B dense and 35B-A3B Mixture of Experts sizes, both on the H Models API, with Holotron4 Nano (built on NVIDIA's Nemotron 3 Nano Omni) released alongside.
  • One model works across desktops, web, Android, a code sandbox and business APIs, using GUIs, code, MCP and APIs.
  • On OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5; Holo4 35B-A3B reaches 30.9%.
  • H says Holo4 uses orders of magnitude fewer parameters at a much lower cost per task, and open-sources every trajectory behind its public-benchmark scores.
  • Training used tasks from the Agentic Task Factory, which has produced about 10,000 tasks; weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF.

Why it matters

Most agent models are built for a single interface: screen clicking or tool calling. Holo4 is presented as one model that mixes GUIs, code, MCP and APIs within a single business task, called the same way on every platform. The other angle is cost. H says a 27B model trails only the strongest closed models on long workflows while costing far less per task. On OSWorld 2.0 the gap to Opus 5.5 is still large, 61.7% against 81.8%, so the argument rests on price and size rather than on beating the frontier.

Who it affects

Teams automating business workflows that cross desktop apps, web pages and APIs, and that would otherwise stitch together separate GUI and tool-calling models. Developers who want to run agent models themselves are also covered, since weights are published in several formats including 4-bit GGUF. Anyone comparing computer-use agents gets the open trajectories to inspect.

How to use it

Both sizes are available today on the H Models API, with a quickstart linked from the post. Weights for Holo4-27B, Holo4-35B-A3B and Holotron4 Nano are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF. Trajectories can be replayed at trajectories.hcompany.ai or downloaded as a dataset. Optimized DSpark drafter checkpoints are promised in the coming days to speed up inference. No prices for the H Models API are given, and the license of the weights is not stated.

How solid is it

This is the developer's own report, and it is detailed. Every trajectory behind the public-benchmark scores is open-sourced, which lets outsiders check the runs. The chart notes are candid that comparisons are uneven: releases, harnesses and task subsets differ between points, and Holo4 is excluded from the line connecting the closed models. On AutomationBench, Holo4 was measured in H's internal harness while other models use public-set scores and leaderboard costs from the private set; H says it will report Holo4 on the private set once evaluated. The text names both Opus 5.5 (81.8% on OSWorld 2.0) and Opus 5 (70.2%, on a different set from OpenAI's launch chart) and does not say whether they are the same model. The three example tasks give calls, tokens and lines but the outputs are not described.

Risks and caveats

The headline OSWorld 2.0 result trails the strongest closed model by a wide margin, and the 35B-A3B model reaches only 30.9%. The claims of far fewer parameters and much lower cost are the authors' own and unquantified: no numeric per-task cost figures are given, and no parameter counts are given for the closed models. In the example tasks Holo4 used fewer calls and tokens on Pac-Man but more on the Eiffel Tower model than Qwen3.8 27B (84 calls against 60). The text does not say what hardware the models need.

“It scores well on academic benchmarks, but we built it for real business workflows.”

— H Company, Holo4 announcement