06 / Agent infrastructure Open source · v0.2.0

Hermes Warm CompactionFaster context enginefor Hermes Agent.

Hermes Agent is an open-source AI agent from Nous Research. When a chat fills half of what its model can read at once, Hermes compacts the chat into a summary. My plugin writes that summary on the server's cache, so the server reads only what's new.

Role
Sole developer
Released
October 2026
Software
Python, Hermes plugin API
Runs on
Unpatched Hermes Agent
Median time to compact a long chat

One DGX server and 10 synthetic chats of more than 100,000 tokens. Tokens are the word pieces a model reads. hermes-lcm is another compaction plugin for Hermes.

The built-in compressor makes the server read the whole chat again.

What the server readsIllustrative

A server's cache helps only when a request starts the same way. The built-in summary request starts differently, so the server reads the whole chat again.

Why the cache misses, in technical terms

A model server can keep the work it did on a request's prompt. This is a prefix cache. When the next request starts with the same tokens, the server reuses that work and reads only the tokens after the match.

At each compaction, the Hermes built-in compressor sends a separate summary request with its own prompt in front of the chat. That request no longer starts like the cached one, so the server reads the whole chat again. The built-in compressor also cuts long messages before it summarizes them.

My plugin sends the last request again, with one instruction at the end.

What the server readsIllustrative

The handoff has five headings, and the newest rows stay word for word. On DGX, at least 99.9% of each request came from the cache.

How the warm request works, in technical terms

At each compaction, the plugin sends the last main-model request of the session again. It adds the rows that came after that request and one handoff instruction at the end. The server can reuse the cached prefix of the earlier request, so it reads only the new rows.

The reply is a Markdown handoff with five headings: Goal, User instructions, Current state, Key facts, and Next step. The plugin replaces the older rows with it. A tail of recent rows stays word for word: by default 2.5% of the context window, from 10,000 to 25,000 tokens.

The speedup needs the same model and a server with a prefix cache. The plugin works for manual /compress and for automatic compaction, which starts at half of the context window by default.

Five design choices keep Hermes unpatched, the chat intact, and the log private.

What each design choice doesFive choices

Hermes stays unpatched, a failed request falls back, the newest messages stay exact, and the log holds no message text.

How each choice works, in technical terms

Use documented APIs only. The plugin registers through the Hermes plugin API. It doesn't patch Hermes code, subclass the built-in compressor, or wrap Hermes code at runtime. It needs Hermes 45871e10, from October 2, 2026, or a later version.

Load safely. At load, the plugin checks each API it uses. If one is missing, it registers nothing, writes a warning to the log, and Hermes keeps its built-in compressor.

Fall back in steps. If the warm request can't run or its reply fails a check, a fallback summary runs through the Hermes auxiliary model route. If that also fails, the plugin writes a fixed-format summary without a model. When Hermes cancels a compaction, the history stays unchanged.

Keep the newest messages. A tail of recent rows stays word for word. By default it's 2.5% of the context window, from 10,000 to 25,000 tokens.

Keep the log private. Each compaction writes one log line that names its path: warm, fallback, or fixed. The log never holds message text, request bodies, or keys. To notice a key change, the plugin keeps a salted PBKDF2-HMAC-SHA-256 digest of the key, never the key itself.

To go back, set the context engine to compressor and start a new session. An integration check runs real Hermes code against a local test server.

In one benchmark, my plugin was faster in every chat and kept more facts.

Measured results10 synthetic chats each

Compaction and the next reply, median on DGX

Built-in compressor97.1 shermes-lcm66.1 sMy plugin27.9 s

My plugin finished first in all 10 chats.

Facts kept after compaction

  • 59/60facts correct, against 50/60 for the built-in compressor and 43/60 for hermes-lcm
  • 10/10facts from the middle of a long message, against 0/10 for both

The other two cut long messages before they summarize them.

Compaction on a second server, median

Built-in compressor41.8 sMy plugin21.1 s

LM Studio with a 4B model. The warm request ran in 9 of 10 chats, and the fallback in the other.

In this benchmark my plugin compacted faster than both other engines in every chat, and it kept more facts.

The benchmark, in technical terms

All runs used synthetic chats of 100,000 or more prompt tokens on unpatched Hermes 45871e10, one session at a time. DGX is an OpenAI-compatible server on an NVIDIA DGX that reports cached tokens. LM Studio ran a Qwen3.5 4B model and reports no cached-token count.

My plugin, hermes-lcm 0.21.0-rc2, and the built-in compressor on DGX, automatic compaction, 10 chats
MeasureMy pluginhermes-lcmBuilt-in
Compaction, median16.8 s43.5 s78.5 s
Compaction and next reply, median27.9 s66.1 s97.1 s
Compacted facts correct59/6043/6050/60
Fact in the middle of a long message10/100/100/10
Standing rule kept in the next reply10/105/106/10
Fact in the recent rows10/1010/1010/10
Next action and its target9/109/1010/10
My plugin and the built-in compressor, median seconds, 10 chats for each engine and mode
ServerModeWarm request ranCompaction, plugin / built-inWith next reply, plugin / built-in
DGXManual10/1010.2 / 32.712.5 / 36.9
DGXAutomatic10/1010.1 / 39.911.5 / 42.9
LM Studio 4BManual9/1020.9 / 34.425.2 / 64.5
LM Studio 4BAutomatic9/1021.1 / 41.825.7 / 54.9

The two benchmarks were separate runs, so their DGX times differ. The results come from one synthetic task family, one session shape, and one model per engine. Cache reuse depends on the server. The three-way benchmark was recorded on October 5, 2026.

Results and records

Contact

Get in touch.

I'm open to full-time AI product engineering roles.

Email hello@shayanbianconi.com or see more of my work on GitHub.