H Company's Holo4 Handles Screens, Code, and APIs With 79% Fewer Tokens

H Company's new Holo4 family runs across GUIs, code, MCP, and APIs with a single model, competing with frontier systems at a fraction of the cost.

·
·
H Company's Holo4 Handles Screens, Code, and APIs With 79% Fewer Tokens
Read5 min
TypeNews
TopicAgents · Api
  • Holo4 released in 27B dense and 35B-A3B MoE variants, both open weights.
  • Single model handles desktop, web, Android, code sandboxes, MCP, and APIs.
  • Holo4 27B scores 61.7% on OSWorld 2.0 vs 81.8% for Opus 5.5, at far lower cost.
  • Trained via SFT plus RL on ~10,000 synthetic tasks from H's Agentic Task Factory.
  • Weights available on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF.
  • All benchmark trajectories open-sourced for step-by-step replay and auditing.

H Company releases Holo4 for GUI, code, and API automation

Paris-based H Company has released Holo4 models, a family of computer-use models that can click through graphical interfaces, write and run code in sandboxes, and invoke tools through REST APIs or the Model Context Protocol, known as MCP. Both checkpoints are available through the H Models API, with downloadable weights for self-hosting.

The release includes a 27-billion-parameter dense model and a 35-billion-parameter Mixture of Experts model that activates about 3 billion parameters per token. H publishes the weights in FP16, FP8, and GGUF formats.

One agent across screens, shells, and APIs

Software agents often depend on a single control method: visual interaction for desktop applications, browser automation for websites, or structured tool calls for services with APIs. Holo4 uses one architecture across desktop operating systems, browsers, Android devices, code sandboxes, MCP servers, and REST endpoints.

A single checkpoint can navigate menus in FreeCAD, call an enterprise API, and run a shell script within the same workflow. Hybrid tasks can also expose one application state through both a graphical interface and MCP, allowing the model to select a shorter or more reliable route.

Checkpoint Architecture Practical profile
Holo4-27B 27B dense Higher published OSWorld 2.0 score; all parameters participate in inference
Holo4-35B-A3B 35B Mixture of Experts, about 3B active Lower active parameter count per token; suited to efficiency testing at higher volume

Desktop scores reveal the trade-offs

OSWorld 2.0 evaluates whether an agent can complete tasks in real desktop applications. H reports a 61.7% score for Holo4-27B and 30.9% for Holo4-35B-A3B. The company’s cited Opus 5.5 result is 81.8%, leaving the stronger Holo4 checkpoint 20.1 percentage points behind that closed-model reference.

Model OSWorld 2.0
Opus 5.5 81.8%
Holo4-27B 61.7%
Holo4-35B-A3B 30.9%

The results place the dense 27B checkpoint well ahead of the larger MoE checkpoint on this desktop benchmark. The 35B-A3B label describes total and active parameters rather than a capability ranking. H also evaluates API-oriented work with AutomationBench, though the release summary cited here provides no corresponding score.

H’s release post includes a Pac-Man project built in Godot. Holo4-27B completed the task in 68 tool calls and consumed 2.4 million tokens, while its Qwen3.8 27B base used 197 calls and 11.4 million tokens for a comparable attempt under the same prompt and harness. Those figures represent about 65% fewer tool calls and 79% fewer tokens. Actual serving cost also depends on hardware, quantization, caching, and API rates.

Synthetic tasks teach route selection

H trained the models with supervised learning and reinforcement learning on a large set of environments and tasks. Its Agentic Task Factory creates those environments from software documentation and application screenshots, then defines outcomes that can be checked automatically.

Some generated tasks expose the same underlying state through a visual interface and MCP. That setup trains the model to compare interaction paths, such as clicking through several dialogs or making one structured tool call, while preserving the same target result.

H also revised its execution harness with multi-step memory management and a native desktop shell. Memory management helps preserve task state across long sequences of observations, actions, tool responses, and errors. The shell gives the model a direct route for file operations, scripts, and command-line tools when the environment permits them.

Weights, API access, and deployment

Both checkpoints are available through the Models API, with commercial use allowed. H describes the service as pay as you go. For comparison, the previous Holo3-1-35B-A3B model costs $0.25 per million input tokens and $1.80 per million output tokens, with a free tier limited to 10 requests per minute. Developers should confirm Holo4-specific rates in the API console before estimating production costs.

Deployment Format or runtime Use case
Managed API H Models API Hosted inference and a model-ID change from existing Holo3 pipelines
Datacenter self-hosting BF16, FP8, or NVFP4 through vLLM GPU deployments requiring control over data, latency, or throughput
Local inference 4-bit GGUF through llama.cpp Consumer hardware, prototyping, and offline evaluation
Edge deployment Holotron4 Nano A related smaller model based on NVIDIA Nemotron 3 Nano Omni

Workloads that fit the model

Holo4’s mixed interaction modes align with workflows that cross several applications or encounter incomplete API coverage:

  1. Back-office automation across CRMs, ERPs, ticketing systems, and legacy interfaces.
  2. QA and regression testing across web, desktop, and Android applications.
  3. Research and data-entry workflows combining browsing, spreadsheet edits, scripts, and API writes.
  4. Internal tools that require self-hosted weights or auditable action logs.
  5. High-volume automation where active parameter count and token consumption affect operating cost.

Production evaluations should account for more than aggregate benchmark scores. GUI agents can fail when layouts, permissions, dialogs, or application versions change. API and shell access introduce separate risks around credentials, destructive actions, and untrusted output. Sandboxed execution, least-privilege credentials, action limits, confirmation gates, and complete logs remain appropriate controls.

A tiered deployment can route routine tasks to Holo4-35B-A3B, send harder visual work to Holo4-27B, and escalate low-confidence or high-impact cases to a stronger model or human reviewer. Workload-specific tests should determine that routing because the published scores cover only selected environments.

Open trajectories expose each decision

H publishes the trajectories behind its reported benchmark results through a trajectory viewer and Hugging Face. A trajectory records the observations, model decisions, tool calls, and resulting state changes for an evaluation run, allowing researchers to inspect where an agent succeeded, looped, or failed.

Open weights, multiple quantizations, and replayable trajectories give developers several ways to reproduce the results and compare alternative prompts, harnesses, and runtimes. Independent replication remains necessary because computer-use benchmarks are sensitive to environment versions, step limits, and execution settings.

Holo4 gives teams one model family for visual control, code execution, and structured tool use, with managed and self-hosted deployment paths. Its strongest published desktop result still trails the cited closed-model baseline, while its open artifacts make the efficiency and failure claims easier to inspect.

Trending
  • No trending articles

Comments

avatar

Next Reads