Holo3.1: Fast & Local Computer Use Agents

Holo3.1 is H company's open computer-use model family with first quantized checkpoints for fast local agents, lifting AndroidWorld scores to 79.3% and cutting step times.

Tuesday June 2, 2026 Source: huggingface.co
TL;DR — Quick Answer

Holo3.1 is a family of open-weight computer-use models from H company on Hugging Face, released June 2, 2026. It is the first Holo release to ship quantized checkpoints — FP8, Q4 GGUF, and NVFP4 — for fast, fully private local agents. On AndroidWorld the 35B-A3B model improves from 67% to 79.3%, and on DGX Spark NVFP4 plus harness optimizations cut average agent step time from 6.8 seconds to 3.3 seconds.

Key Takeaways

Holo3.1: Fast & Local Computer Use Agents — AI news article illustration

H company has released Holo3.1, a family of computer-use models that for the first time ships quantized checkpoints — FP8, Q4 GGUF, and NVFP4 — built for fast, fully local agents. It is the follow-up to Holo3, the state-of-the-art computer-use model the team released last March on Hugging Face.

Computer Use Across Every Environment

Following Holo3, H company says adoption was immediate across browser automation, business software, internal tools, and desktop applications. But production teams kept hitting the same wall: strong performance in one environment does not transfer to another. Holo3.1 is built to fix exactly that, spanning web, desktop, and mobile, and integrating with different agent frameworks.

Mobile Automation Jump

Mobile is the biggest single gain. On AndroidWorld, the Holo3.1 35B-A3B improves from 67% to 79.3%, while the smaller 4B and 9B variants jump from 58% to 72%. Native function-calling support joins the structured JSON outputs from Holo3, so teams can drop the model into third-party harnesses.

Fast, Local, Private

The quantized checkpoints make local execution practical for the first time:

Four Sizes, Four Jobs

The family ships in four sizes: Holo3.1-0.8B for ultra-lightweight local agents, 4B for cost-efficient deployment, 9B for balanced performance, and 35B-A3B for state-of-the-art results — each with optimized FP8, NVFP4, and Q4 GGUF weights.

Availability

All checkpoints are on Hugging Face under the Hcompany organization, alongside the Holo Models API for hosted deployment. H company frames the release as a major step toward universal computer-use agents: systems that run wherever the workflow lives.

What This Means

Holo3.1 makes computer-use agents a local, private workload rather than a cloud dependency. For enterprises that cannot send screens or data off-premises, that is the difference between an interesting model and a deployable system.

Frequently Asked Questions

What is Holo3.1?

A family of open-weight computer-use models from H company, released June 2, 2026, that operate across web, desktop, and mobile environments.

Can I run Holo3.1 locally?

Yes. For the first time Holo3 ships quantized FP8, Q4 GGUF, and NVFP4 checkpoints, so agents can run privately on a Windows or Mac machine or a DGX Spark on the local network.

How fast is Holo3.1 on DGX Spark?

NVFP4 W4A16 delivers 1.41x the token throughput of FP8 and 1.74x that of BF16, and combined harness optimizations cut average step time from 6.8 seconds to 3.3 seconds.

Is Holo3.1 open source?

The model family and quantized checkpoints are released open-weight on Hugging Face under the Hcompany organization, with an optional hosted Holo Models API.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news