Omni-IO Skills harness lifts GPT-5.6 Sol and Claude Sonnet 5 to 100% input support

The paper starts from a familiar gap. General-purpose agents can plan, reason and act over long horizons, but their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets and code. The authors name two ways to close the gap and find both lacking. Extending a foundation model to more modalities ties capability growth to costly model updates. Assembling specialist models and tools leaves open how procedures, dependencies, intermediate assets and cross-turn revisions should be coordinated.
Their answer is Omni-IO Skills, described as a plug-and-play Agent Harness that makes existing agents omni-native. It rests on four parts: hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs. These graphs schedule independent operations concurrently and register successful outputs so they can be reused downstream and across turns, over replaceable execution backends.
The harness ships 27 Skills covering 38 representative tasks. The tasks span seven artifact modalities and four capability families: understanding, generation, reasoning and retrieval.
The evaluation uses UniM-90 with two host agents, GPT-5.6 Sol and Claude Sonnet 5. With the harness, their input-support rates rise from 40.00% and 38.89% respectively to 100%. Their relative Semantic--Quality Coupled Score goes from 26.99 to 74.94 for GPT-5.6 Sol and from 27.82 to 77.78 for Claude Sonnet 5. Strict Structure Score reaches 100.00 and 99.78, listed in the same order as the hosts.
The authors conclude that these results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
Key facts
- Omni-IO Skills is a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration and a persistent Asset Registry.
- It includes 27 Skills covering 38 representative tasks across seven artifact modalities and four capability families: understanding, generation, reasoning and retrieval.
- On UniM-90, input-support rates rise from 40.00% (GPT-5.6 Sol) and 38.89% (Claude Sonnet 5) to 100%.
- Relative Semantic--Quality Coupled Score goes from 26.99 to 74.94 for GPT-5.6 Sol and from 27.82 to 77.78 for Claude Sonnet 5.
- Strict Structure Score reaches 100.00 and 99.78, and the host agent's reasoning core is left unchanged.
Why it matters
The paper targets a practical bottleneck: agents that reason well but produce work unevenly across text, images, audio, video, documents, 3D assets and code. The authors argue that retraining a foundation model for each new modality ties progress to costly model updates, while stitching together specialist tools leaves coordination unsolved. Moving that coordination into a harness, with the host agent's reasoning core untouched, is the route they propose.
Who it affects
The work is aimed at people building on existing general-purpose agents who want multimodal production abilities without touching the underlying model. The two hosts tested are GPT-5.6 Sol and Claude Sonnet 5. Teams that need multi-step workflows over several asset types, with intermediate results reused across turns, are the audience the Asset Registry and execution graphs are designed for.
How to use it
The paper describes Omni-IO Skills as plug-and-play: the harness attaches to an existing agent through its Skills and a standardized multimodal execution interface, and execution backends are replaceable. Workflows are expressed as Declare Execution Graphs, so independent operations can run concurrently. No code, repository, license or release availability is mentioned, so how to obtain or install it is not stated.
How solid is it
The claims come from the paper's abstract alone. The headline numbers are specific: input support up to 100% on UniM-90 for both hosts, relative Semantic--Quality Coupled Score of 74.94 and 77.78, and Strict Structure Score of 100.00 and 99.78. But the abstract does not say what UniM-90 is beyond its name, nor how many tasks or items it contains. It does not define the Semantic--Quality Coupled Score or Strict Structure Score, nor what 'relative' is relative to. No baseline comparisons against other agent harnesses or specialist-model pipelines are given.
Risks and caveats
No latency, cost or compute figures are given, so the overhead of the harness is unknown. No authors or institutions are named in the abstract. The Strict Structure Score values are not tied to a specific host beyond the order of mention. The conclusion that harness-level composition is a practical route is the authors' own, drawn from two host agents on one benchmark.
“These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.”
— Omni-IO Skills paper abstract