Nous Research’s Hermes Agent has crossed 140,000 GitHub stars in under three months and, according to OpenRouter, is now the most used agent in the world. It is also the latest open-source framework to find a natural home on NVIDIA hardware — running always-on, self-improving agents locally on RTX PCs, RTX PRO workstations, and DGX Spark.
The Breakout Framework
Following the success of OpenClaw, the community has embraced Hermes for two qualities that have historically been hard to achieve in agents: reliability and self-improvement. Nous Research curates and stress-tests every skill, tool, and plugin that ships with the framework, so it works out of the box even with 30-billion-parameter-class local models — without the constant debugging other agent stacks require.
Self-Improving by Design
- Self-evolving skills — Hermes writes and refines its own skills, saving learnings after every complex task or piece of feedback
- Contained sub-agents — short-lived, isolated workers with focused context and tools, ideal for small local context windows
- Active orchestration — Hermes is an orchestration layer, not a thin wrapper, enabling persistent on-device agents instead of task-by-task execution
Hardware That Runs 24/7
Hermes is built to run continuously — responding, planning multistep tasks, executing, and self-improving. Quality of hardware directly determines quality of experience, and NVIDIA GPUs are purpose-built for that workload. DGX Spark, with 128GB of unified memory and 1 petaflop of AI performance, can run 120-billion-parameter mixture-of-experts models all day.
Qwen 3.6 Under the Hood
Hermes is provider- and model-agnostic, and Alibaba’s new Qwen 3.6 open-weight models pair especially well. The Qwen 3.6 35B runs on roughly 20GB of memory while surpassing the previous-generation 120B models, and the 27B matches 400-billion-parameter accuracy at one-sixteenth the size.
Getting Started
Developers can grab the framework from the Hermes GitHub repository, pair it with a preferred local model and runtime — LM Studio and Ollama both ship Hermes support out of the box — and run it locally for a private, always-on agent that improves with every task.
What This Means
The combination of self-improving, always-on agents and local NVIDIA compute points to a future where agents are not a cloud request but a persistent part of the machine itself. Hermes has the adoption, and the hardware now has the memory to back it.