Skip to content
AI Primer
release

H Company releases 27B and 35B-A3B Holo4 computer-use models

Holo4 ships as a dense 27B model and a 35B-A3B mixture-of-experts model, with API access and downloadable weights. H Company reports a 61.7% score on OSWorld 2.0 and publishes replayable benchmark trajectories.

5 min read
H Company releases 27B and 35B-A3B Holo4 computer-use models
H Company releases 27B and 35B-A3B Holo4 computer-use models

TL;DR

The 27B model card gives its weights a noncommercial license despite the commercial API launch. The benchmark table reports 41.5% full task success behind the 61.7% OSWorld 2.0 score, and its AutomationBench notes disclose training-data collection from a split containing 480 of the 600 public tasks.

Models and licenses

Holo4 combines screen clicks and typing with code execution and MCP or API calls in the same agent loop. The two model cards identify different Qwen foundations and a 262,144-token maximum context in their configurations:

  • Holo4-27B: Dense, based on Qwen3.8-27B. Weights are CC BY-NC 4.0.
  • Holo4-35B-A3B: Mixture of experts with 3B active parameters, based on Qwen3.6-35B-A3B. Its weights are Apache 2.0.

Both are on the H Models API and in the Hugging Face collection as BF16, FP8, NVFP4 and 4-bit GGUF weights. The public API models documentation still lists Holo3 IDs, even though H Company's launch says Holo4 is available there.

OSWorld 2.0

Holo4-27B reached 61.7% average partial score and 41.5% full task success on long desktop workflows in H Company's results table. The MoE scored 30.9% partial and 12.3% full success; the Qwen3.8-27B base scored 48.0% partial, a 13.7-point gap to Holo4-27B.

The tweet compares Holo4 with GPT-6 Sol at 60.5%, its xhigh setting. The full cost-performance chart puts GPT-6 Sol max at 64.4% and Opus 5.5 at 81.8%. H Company estimates $1.22 per Holo4-27B task from its API rates; those outside scores come from different harnesses, effort settings and task subsets.

Other benchmark scores

H Company reports these Holo4-27B figures in its benchmark table:

  • OSWorld: 85.2%, versus 84.3% for the Qwen3.8-27B base.
  • ALE-CLI: 44.1% average score on the 105-task Linux split of Agents' Last Exam.
  • AndroidWorld: 85.1%.
  • AutomationBench: 45.4% at an estimated $0.05 per task. The overview in H Company's benchmark tweet instead says 45.5%; the same tweet's cost line says 45.4%.

The dense model leads the MoE on every benchmark in that table, including OSWorld at 85.2% versus 80.8% and AndroidWorld at 85.1% versus 77.6%.

AutomationBench's public split

The 45.4% AutomationBench figure comes from 600 public v1.0.6 tasks, according to H Company's evaluation notes. Of those, 480 belong to a split from which H collected training data; on the other 120 held-out tasks, the 27B scored 49.3% and the MoE 31.7%. H says it will report private-set results after evaluation. Its cost chart mixes these internal public-set scores with other models' costs from the official leaderboard's private set.

Agentic Task Factory

The factory builds interactive environments and verifiable tasks from product documentation, website screenshots and open-source software. H reports about 4,000 web-app tasks, 3,000 MCP-server tasks and 3,000 desktop or OS tasks, including environments that expose the same state through GUI and MCP.

A task passes the factory's gates only when its verifier fails on the untouched seed, passes on the intended final state, rejects near misses, and an agent solves it through the real interface. Failed attempts feed an audit-and-harden loop.

Two RL experts

H Company trained on 127 billion tokens, roughly three quarters from successful agentic trajectories, per its training summary. Its published recipe divides that trajectory mix into desktop (45%), web (14%), MCP and API (12%), and mobile (3%).

After supervised fine-tuning, asynchronous online RL trains two LoRA experts: one for desktop and web, another for terminal, MCP and API. H says it merges them with equal weight into the fine-tuned model, with no additional training after the merge.

Harness memory and desktop shell

The agent harness now preserves context across hundreds of steps and can run a shell on the desktop machine itself. H Company charts its OSWorld 2.0 partial score rising from 6.1% to 59.8% through successive August and September harness changes, before the 61.7% release result.

The change log embedded in the launch post includes context passthrough during memory compaction, a fix to vision serving, desktop control through the shell, and raising the run limit from 200 steps and two hours to 500 steps and six hours. Those milestones combine harness and serving changes, rather than isolating a model-weight improvement.

Replayable trajectories

H Company released 7,366 agent traces from both Holo4 models, spanning OSWorld, OSWorld 2, AndroidWorld, AutomationBench, PinchBench and Agents' Last Exam. The dataset card describes step-level actions, tool results, screenshots and verifier results; data/index.json carries task, score, duration and step counts.

The bundle is downloadable and browsable as replays. Credentials and personal data are masked, sensitive screenshots are replaced with placeholders, and a few tasks are omitted, according to the dataset card.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 4 threads
TL;DR2 posts
Models and licenses1 post
Agentic Task Factory1 post
Two RL experts1 post
Share on X