On September 28, 2026, H Company announced Holo4, a family of computer-use models with two variants: dense model Holo4-27B and mixture-of-experts model Holo4-35B-A3B. H Company says both are available through its H Models API and that their weights are published on Hugging Face.
How Holo4 is designed to interact with software
H Company describes Holo4 as able to work through a graphical user interface (GUI), write and run code, and call tools through the Model Context Protocol (MCP) or APIs. The company lists desktop, web, Android, code-sandbox and business-API environments.
That combination gives the models more than one way to act on a task: they can interact with screen controls, execute code or call connected tools, depending on the environment. These are capabilities H Company describes for the model family.
Holo4-27B and Holo4-35B-A3B compared
| Variant | Model design | Parameters | Weight license |
| Holo4-27B | Dense vision-language model, based on Qwen3.8 | 27 billion | CC BY-NC 4.0 |
| Holo4-35B-A3B | Mixture-of-experts model, based on Qwen3.6 | 35 billion total; 3 billion active | Apache 2.0 |
The licenses apply to the downloadable Holo4 weights. Holo4-27B’s model card also identifies an Apache 2.0 license for its upstream Qwen3.8-27B model; that is separate from the CC BY-NC 4.0 license listed for Holo4-27B’s weights. H Company also offers hosted access through the H Models API.
What H Company’s benchmark results mean
H Company reports an 85.2% OSWorld score for Holo4-27B and 80.8% for Holo4-35B-A3B. The company says its comparisons across models use different evaluation harnesses, effort settings and task subsets, so results across systems do not all come from the same conditions. Holo4’s OSWorld 2.0 and ALE-CLI results use one run; most of its other benchmark scores average two to four runs.
AutomationBench’s public v1.0.6 set contains 600 tasks. H Company says 480 were in the split used to collect training data, and reports separate results on the remaining 120 held-out tasks: 49.3% for Holo4-27B and 31.7% for Holo4-35B-A3B. That distinction matters when assessing what the public-set figures say about performance beyond tasks in the training-data split.
Training and the planned next release
H Company says supervised fine-tuning for Holo4 used 127 billion tokens, with about three quarters consisting of successful agent trajectories. It reports that the trajectories covered desktop tasks (45%), web (14%), MCP and API use (12%), and mobile tasks (3%). The company then trained two specialized LoRA experts and merged them into the fine-tuned model.
H Company also says its Agentic Task Factory generated about 10,000 tasks: approximately 4,000 for web apps, 3,000 for MCP servers and 3,000 for desktop and operating-system tasks. The company planned to release optimized DSpark drafter checkpoints in the days following its September 28 announcement.