H Company Releases Holo4 Agentic Models, Scoring 61.7% on OSWorld 2.0 at a Fraction of Frontier Cost
H Company has released Holo4, a series of agentic models in 27B dense and 35B-A3B Mixture of Experts sizes that operate across GUIs, code, MCP, and APIs using a single interface. The models score 61.7% (27B) and 30.9% (35B-A3B) on OSWorld 2.0, trailing closed frontier models like Opus 5.5 (81.8%) but at far lower cost.
H Company has released Holo4, a new series of agentic models built to operate computer interfaces the way humans do — through GUIs, code execution, MCP, and APIs — using a single model and calling convention across all of them.
The release comes in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts model. Both are available now through the H Models API in FP16, FP8, and GGUF formats. H Company also released Holotron4 Nano, an updated version of its Holotron 3 model.
Benchmark performance
On OSWorld 2.0, a benchmark for long-horizon desktop control tasks, Holo4 27B scores 61.7%, and Holo4 35B-A3B scores 30.9%. For comparison, according to H Company, Anthropic's Opus 5.5 scores 81.8% on the same benchmark. H Company says Opus 5 scores 70.2% and GPT-5.6 Sol scores 66.2%, citing OpenAI's launch chart data and max-effort partial rewards on the v2026.08.08 offline set.
H Company positions Holo4's value proposition around cost-efficiency rather than raw score: the company claims Holo4 competes with frontier closed models on OSWorld 2.0 and AutomationBench (a benchmark for API-based tool use) at substantially lower per-task cost, given its smaller parameter count. The company states Holo4 is built on a Qwen base (Qwen3.8 27B and Qwen3.6 35B-A3B) and shows significant improvement over that base model.
H Company has open-sourced the trajectories behind its benchmark scores, allowing outside parties to replay each step of a benchmark run at trajectories.hcompany.ai or download them from Hugging Face.
Training approach
According to H Company, Holo4 was trained using supervised learning and reinforcement learning across a large set of environments and tasks, including tasks generated by the company's internal "Agentic Task Factory" — a pipeline that the company says has produced roughly 10,000 tasks across web apps, MCP servers, and desktop environments by building interactive environments and verifiable tasks directly from documentation and screenshots of real software.
H Company also says it rebuilt its training harness — the execution loop that manages an agent's actions and context over hundreds of steps — based on failure analysis from OSWorld 2.0 runs. The company states the two largest changes were adding persistent memory that tracks state across hundreds of steps, and giving the agent direct shell access on the desktop machine.
In examples shared by the company comparing Holo4 27B against its Qwen3.8 27B base model on tasks like building a 3D model in FreeCAD and constructing a Pac-Man-style game in Godot, Holo4 27B reportedly completed tasks using fewer API calls and less code than the base model under identical prompts and harness conditions.
Pricing and availability
Holo4 27B, Holo4 35B-A3B, and Holotron4 Nano are available now through the H Models API. Pricing has not yet been disclosed. Model weights are available in FP16, FP8, and GGUF formats via Hugging Face.
What this means
Holo4's core claim — a single model that operates GUIs, code, MCP, and APIs interchangeably without task-specific selection — addresses a real fragmentation problem in computer-use agents, where GUI-trained models fail against API-only tasks and vice versa. Whether that generalization holds up outside curated benchmarks and cherry-picked examples like the FreeCAD and Godot demos remains to be independently verified.
The gap to closed frontier models is still substantial: 61.7% versus 81.8% on OSWorld 2.0 is a 20-point deficit against Opus 5.5. H Company's argument rests entirely on cost-per-task, a metric it controls the framing of via its own cost-performance charts. Buyers evaluating Holo4 for production workflows should weigh the completion-rate gap directly against their tolerance for failed or incomplete agentic runs, since a cheaper model that fails more often on long workflows can still cost more in aggregate. The open-sourcing of trajectories is a genuinely useful transparency move that lets third parties scrutinize scores rather than take benchmark charts at face value.
Related Articles
Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware
Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.
Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context
Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.
NVIDIA Releases Nemotron 3 Diarization, a 100M-Parameter Open-Weight Model Ranked #1 on Voice Arena's Diarization-Bench
NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who is speaking and when in audio conversations. It ranked #1 among 17 system configurations on Voice Arena's Diarization-Bench with a 14.72% diarization error rate, supporting up to eight speakers in both live and recorded audio.
Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time
Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.
Comments
Loading...