05 / Developer tools Open source · v1.2.0

Hermes Classic GoldA theme and telemetry packfor Hermes Agent.

Hermes Agent is an open-source AI agent from Nous Research. Classic Gold gives its desktop app a gold theme and a telemetry tape, a strip of numbers about the running session. It shows the model, the cost, and the computer's memory at a glance.

Role
Sole developer
Built
July–September 2026
Software
JavaScript, Python, Hermes plug-in SDK
Runs on
Hermes Desktop on Windows
Hermes Desktop with the Classic Gold theme: a dark window with a dotted gold caduceus behind a pixel HERMES-AGENT wordmark, a small pixel pet in the corner, and a gold telemetry tape above the task box with the model, provider, tokens, speed, cost, VRAM, and RAM.
Classic Gold before the first session. Each value fills in as Hermes records it.

The tape had to survive Hermes updates.

Where the first version kept the tapeIllustrative

Version 1.2 installs as a plug-in in the Hermes profile, which updates leave in place.

What went wrong, in technical terms

During an agent session I want a few facts at a glance: which model and provider run, how full the context is, what the session costs, whether the prompt cache hits, and how much memory is left on the machine that runs the model.

The first version added its status bar as a patch to the Hermes source. A Hermes update then reverted the patch, or stopped on a Git conflict.

A plug-in draws the tape, and a small backend reads the numbers.

The tape and its sourcesSchematic
The Classic Gold tape and where its values come fromA backend host holds the Hermes session record, a graphics card, main memory, and the pack's Python backend. The tape sits above the task box in Hermes Desktop. The model menu opens from the tape. The backend reads the session record for context, cache, and cost, and reads the host's memory for the RAM and VRAM gauges. The drawing is a schematic.SESSIONRECORDBACKENDGPURAMBACKEND HOSTHERMESMODELCONTEXTCACHE--COST--RAMVRAMMODELSESSIONHARDWARE+ Give Hermes a task

The tape and the backend install inside the Hermes profile as a plug-in, so a Hermes update leaves them in place.

How the plug-in works, in technical terms

Version 1.2 installs through the public Hermes plug-in SDK (software development kit). It adds a renderer plug-in for the theme and the tape, and a small Python backend for telemetry, both inside the Hermes profile. The context is the text the model holds at once, and the tape shows how full it is. The Hermes checkout stays untouched, so a normal update can't conflict with it.

The backend runs on the host that runs Hermes, so RAM and VRAM describe the machine that does the work, even when it's a remote server.

Five design choices keep the tape honest, the install safe, and messages private.

What each design choice doesFive choices

The tape shows only numbers it knows, the installer never guesses or removes your edits, and the backend never reads messages.

How each choice works, in technical terms

Show unknown as unknown. The tape never invents a value. Speed is the average over a finished turn, and a cost that is only estimated shows as --.

Refuse to guess the target. With more than one Hermes profile, the installer refuses to choose. Every command that writes has a dry run, and --yes skips a prompt but never picks a profile.

Remove only what it can prove it installed. The installer records what it writes. Uninstall removes or restores only recorded files whose hashes still match, and leaves later user edits in place. Migration from the old source patch restores a file only when its backup matches the current Hermes version exactly.

Hide fields before they crowd. The tape drops RAM and VRAM below 1180 px, then speed, cost, and time below 1000 px. Below 880 px it hides completely. These rules win over user settings.

Read no message text. The backend reads the session record for cost and display fields only. It never reads prompts or messages, and the install scripts make no network requests.

What the tape shows, and when it shows --
FieldShowsShows -- when
SpeedOutput tokens ÷ the full turn timeThe turn is still running
CostThe provider's actual session cost, or 0.00 for an included planThe cost is unknown or only estimated
Cache hitsCache reads ÷ (input + cache reads + cache writes)No cache read is recorded yet
RAMUsed and total memory of the backend hostpsutil isn't available
VRAMUsed and total memory of all NVIDIA GPUs on the hostnvidia-smi isn't available

I traced a context‑count error and sent a fix to each project involved.

The error, the fixes, and the testsThree in review

What the context gauge showed

Acting model's request8,000Advisor's request230,000Gauge showed238,000

Tokens are the word pieces a model reads. No single model held 238,000 of them.

Fixes sent to each project

Release 1.2.0

  • 206tests passed
  • 2systems in automated testing: Windows and Ubuntu

Automated checks also include code style and security scans.

Three fixes are in review, each with passing tests, and release 1.2.0 passed its 206 tests.

The error and the fixes, in technical terms

With mixture-of-agents, Hermes could count an 8,000-token acting request and a 230,000-token advisor request as 238,000 tokens in one context. The gauge then showed a context that no single model had. I traced the fault into Hermes and its compaction plug-in, and sent a fix to each layer.

Hermes Agent, "Separate context usage and preserve acting output budgets". Context pressure uses the acting request, and billing keeps the combined total. 361 tests passed across 24 related files.

hermes-lcm, "Rearm the compaction budget after LCM progress". After three successful compactions in one long turn, automatic compaction stopped. A new test runs five compactions in one turn and passes only with the fix.

Classic Gold, "Fix the active context gauge for MoA and compaction". A 120,500-token active context stayed at 120,500 when the compressor reported 513,000. 207 renderer and 28 backend tests passed.

Release 1.2.0 passed 206 Node tests, the telemetry backend tests, and bundle tests on Node 18. CI runs on Windows and Ubuntu, with lint and CodeQL.

Contact

Get in touch.

I'm open to full-time AI product engineering roles.

Email hello@shayanbianconi.com or see more of my work on GitHub.

Email me