06 / Agent infrastructure Open source · v0.2.0
Hermes Warm CompactionFaster context enginefor Hermes Agent.
Hermes Agent is an open-source AI agent from Nous Research. When a chat fills half of what its model can read at once, Hermes compacts the chat into a summary. My plugin writes that summary on the server's cache, so the server reads only what's new.
- Role
- Sole developer
- Released
- October 2026
- Software
- Python, Hermes plugin API
- Runs on
- Unpatched Hermes Agent
One DGX server and 10 synthetic chats of more than 100,000 tokens. Tokens are the word pieces a model reads. hermes-lcm is another compaction plugin for Hermes.
The problem
The built-in compressor makes the server read the whole chat again.
Last request
In the server's cache
Built-in summary request
Start differs: all read again
In the cacheRead againBuilt-in prompt
The server keeps the last request in its cache.
The built-in compressor puts a new prompt in front.
The start no longer matches, so the server reads everything.
A server's cache helps only when a request starts the same way. The built-in summary request starts differently, so the server reads the whole chat again.
Why the cache misses, in technical terms
A model server can keep the work it did on a request's prompt. This is a prefix cache. When the next request starts with the same tokens, the server reuses that work and reads only the tokens after the match.
At each compaction, the Hermes built-in compressor sends a separate summary request with its own prompt in front of the chat. That request no longer starts like the cached one, so the server reads the whole chat again. The built-in compressor also cuts long messages before it summarizes them.
The approach
My plugin sends the last request again, with one instruction at the end.
Last request
In the server's cache
Warm request
Only the new part read
New history
Handoff, then the newest rows
In the cacheReadMy instruction and handoff
My plugin resends the last request with one instruction added.
The start matches. The server reads only the new rows.
The model writes a handoff that replaces the older rows.
The handoff has five headings, and the newest rows stay word for word. On DGX, at least 99.9% of each request came from the cache.
How the warm request works, in technical terms
At each compaction, the plugin sends the last main-model request of the session again. It adds the rows that came after that request and one handoff instruction at the end. The server can reuse the cached prefix of the earlier request, so it reads only the new rows.
The reply is a Markdown handoff with five headings: Goal, User instructions, Current state, Key facts, and Next step. The plugin replaces the older rows with it. A tail of recent rows stays word for word: by default 2.5% of the context window, from 10,000 to 25,000 tokens.
The speedup needs the same model and a server with a prefix cache. The plugin works for manual /compress and for automatic compaction, which starts at half of the context window by default.
Decisions
Five design choices keep Hermes unpatched, the chat intact, and the log private.
Changes to Hermes
–
Nonedocumented plugin APIs only
- No patch to Hermes code
- No wrapper around it
A missing plugin API
–
Built-incompressor stays in use
- Each API checked at load
- A warning in the log
A failed warm request
–
Fallbacka second way to summarize
- Fallback summary
- Then a fixed summary, no model
Newest messages
–
Exactkept word for word
- Older rows go into the handoff
- 10,000 to 25,000 tokens kept
Message text in the log
–
Noneone line per compaction
- No request bodies or keys
- Which path ran: warm, fallback, or fixed
The plugin uses documented plugin APIs only. Hermes stays unpatched.
If an API is missing, Hermes keeps its built-in compressor.
If the warm request fails, a fallback summary runs.
The newest messages stay word for word.
The log never holds message text, request bodies, or keys.
Hermes stays unpatched, a failed request falls back, the newest messages stay exact, and the log holds no message text.
How each choice works, in technical terms
Use documented APIs only. The plugin registers through the Hermes plugin API. It doesn't patch Hermes code, subclass the built-in compressor, or wrap Hermes code at runtime. It needs Hermes 45871e10, from October 2, 2026, or a later version.
Load safely. At load, the plugin checks each API it uses. If one is missing, it registers nothing, writes a warning to the log, and Hermes keeps its built-in compressor.
Fall back in steps. If the warm request can't run or its reply fails a check, a fallback summary runs through the Hermes auxiliary model route. If that also fails, the plugin writes a fixed-format summary without a model. When Hermes cancels a compaction, the history stays unchanged.
Keep the newest messages. A tail of recent rows stays word for word. By default it's 2.5% of the context window, from 10,000 to 25,000 tokens.
Keep the log private. Each compaction writes one log line that names its path: warm, fallback, or fixed. The log never holds message text, request bodies, or keys. To notice a key change, the plugin keeps a salted PBKDF2-HMAC-SHA-256 digest of the key, never the key itself.
To go back, set the context engine to compressor and start a new session. An integration check runs real Hermes code against a local test server.
Evidence
In one benchmark, my plugin was faster in every chat and kept more facts.
Compaction and the next reply, median on DGX
My plugin finished first in all 10 chats.
Facts kept after compaction
- 59/60facts correct, against 50/60 for the built-in compressor and 43/60 for hermes-lcm
- 10/10facts from the middle of a long message, against 0/10 for both
The other two cut long messages before they summarize them.
Compaction on a second server, median
LM Studio with a 4B model. The warm request ran in 9 of 10 chats, and the fallback in the other.
My plugin finished first in all 10 chats.
It kept 59 of 60 facts, including mid-message ones.
On a second server, it compacted about twice as fast.
In this benchmark my plugin compacted faster than both other engines in every chat, and it kept more facts.
The benchmark, in technical terms
All runs used synthetic chats of 100,000 or more prompt tokens on unpatched Hermes 45871e10, one session at a time. DGX is an OpenAI-compatible server on an NVIDIA DGX that reports cached tokens. LM Studio ran a Qwen3.5 4B model and reports no cached-token count.
| Measure | My plugin | hermes-lcm | Built-in |
|---|---|---|---|
| Compaction, median | 16.8 s | 43.5 s | 78.5 s |
| Compaction and next reply, median | 27.9 s | 66.1 s | 97.1 s |
| Compacted facts correct | 59/60 | 43/60 | 50/60 |
| Fact in the middle of a long message | 10/10 | 0/10 | 0/10 |
| Standing rule kept in the next reply | 10/10 | 5/10 | 6/10 |
| Fact in the recent rows | 10/10 | 10/10 | 10/10 |
| Next action and its target | 9/10 | 9/10 | 10/10 |
| Server | Mode | Warm request ran | Compaction, plugin / built-in | With next reply, plugin / built-in |
|---|---|---|---|---|
| DGX | Manual | 10/10 | 10.2 / 32.7 | 12.5 / 36.9 |
| DGX | Automatic | 10/10 | 10.1 / 39.9 | 11.5 / 42.9 |
| LM Studio 4B | Manual | 9/10 | 20.9 / 34.4 | 25.2 / 64.5 |
| LM Studio 4B | Automatic | 9/10 | 21.1 / 41.8 | 25.7 / 54.9 |
The two benchmarks were separate runs, so their DGX times differ. The results come from one synthetic task family, one session shape, and one model per engine. Cache reuse depends on the server. The three-way benchmark was recorded on October 5, 2026.