AI, ML, and networking — applied and examined.
Hands-On with Nous’s New Agent: AI is Finally Writing Its Own Manuals
Hands-On with Nous’s New Agent: AI is Finally Writing Its Own Manuals

Hands-On with Nous’s New Agent: AI is Finally Writing Its Own Manuals

Looking at this installation log terminal interface, it still has that familiar geeky vibe
(This terminal interface is so soothing to look at, without those flashy pop-ups)

Today is Sunday. I woke up early, casually scrolled through GitHub notifications, and saw that Nous Research’s hermes-agent just released version v0.8.0.

Taking a glance at the commit history, they fixed IPv6 timeouts overnight and conveniently purged 115 pieces of dead code from the production environment. Talk about dedication.

Stop Letting AI Be a Goldfish

Do you know the biggest flaw of most so-called AI assistants nowadays?

It’s their goldfish memory. Once you close the web page, it formats everything. The next time you ask it to set up a dev environment, it still stumbles from step one and falls into the same traps. It feels like hiring a rookie on their first day of work, every single day.

But this time, Hermes introduced “closed-loop learning.”

Its logic works like this. After spending five or six turns solving a complex task, it doesn’t just wash its hands and walk away. Instead, it writes a Markdown-formatted Skill file itself. It faithfully records the steps, the pitfalls it encountered, and how to verify the solution.

Now, this is interesting.

It’s like hiring a temp worker to cook. Previously, you had to hold their hand and teach them exactly how much oil to use every time. Now, after finishing the job, they pull out a little notebook, write an SOP (Standard Operating Procedure), and stuff it into your kitchen drawer. Next time a similar situation arises, they just follow the manual.

Looking at this architecture diagram, you can tell they are trying to solve the problem of cross-platform persistence
(Looking at this feature panorama, it covers almost every interaction scenario you can think of)

Moreover, looking at the tinker-atropos reinforcement learning submodule in their repository, these guys really want to lower the barrier to entry. The parameter redundancy of large language models is often like buying an entire supermarket just to cook one dish. Hermes uses Trajectory Compression to weed out useless garbage actions.

Getting the most accurate things done with the least amount of compute. It’s truly fierce.

Tearing Down the Wall Called “Frontend”

I’ve always been hostile towards tools that force you to log into a specific webpage to use them. Frankly, what real geek stares at a Web UI all day?

Hermes’ approach really suits my taste. It essentially turns itself into an incorporeal ghost.

You can throw it onto a $5/month VPS, or just run it in a serverless container like Modal. It goes to sleep when idle, costing almost nothing.

And then?

Then you talk to it via Telegram, Discord, or even WhatsApp.

This image is quite interesting, a winged helmet bringing the cyberpunk vibe to the max
(This minimalist yet mechanical design language really appeals to old-school hackers)

Think about it. You’re walking down the street, suddenly get an inspiration, and directly send a voice message to the bot on WeChat or Telegram. The Hermes in the cloud wakes up, runs RPC calls to tools, and pushes the results back to your phone when finished.

In the latest version, they even added a Termux testing path. This means you can stuff it into an old, dust-collecting Android phone as a local hub. However, I didn’t test the voice module. Rumor has it that installing the full version directly on Termux will blow up Android’s incompatible dependencies, so they made a special .[termux] version. Guess that’s a compromise.

But either way, I really like this unorthodox approach that doesn’t cater to big tech ecosystems.

A Blatant “User-Poaching” War

To be blunt, OpenClaw used to be completely unrivaled in this local network assistant domain.

Everyone was using it. But Hermes is playing “ruthlessly” this time.

It directly built in a hermes claw migrate command line tool. One-click migration. You just type this one line, and it seamlessly rips everything over—from OpenClaw’s SOUL.md persona file to your conversation memory, and even API keys.

The open-source repo dashboard on Github, with over 380 contributors. This community activity is a bit scary
(Just looking at this screen full of PRs and Commits, you can feel how terrifyingly fast the underlying architecture is iterating)

If we look at the two together, the philosophical differences are actually huge.

OpenClaw is more like an obedient marionette; it does whatever you tell it to do, heavily relying on manually configured skill trees from humans.

But Hermes is playing the automation game. It has built-in cron scheduled tasks. You can tell it in plain English, “Backup my database every night and write an audit report while you’re at it.” It runs fully autonomously in the background and can push the results straight to your Slack.

But there’s a catch you need to know. Running these long-lived resident tasks is a bottomless pit for token consumption.

If you use top-tier models like Claude-3.5 or GPT-4o, your end-of-month bill will probably make you wince. Fortunately, it can seamlessly switch to cheaper models like Kimi, GLM, or those on OpenRouter. This is exactly why it emphasizes “toolchain independence” — the framework is ironclad, but the models come and go.

A Garbage Dump in the Brain

I sometimes think about a terrifyingly profound question.

Since it can autonomously generate Skills and continuously write experiences into an SQLite database during use… what if it learns the wrong things?

My shallow understanding is: what if it gets stuck in a rabbit hole with a specific environment error, uses an extremely bizarre and inefficient method to bypass the bug, and then smugly solidifies this “bad experience” into a Markdown skill?

Every time it encounters this issue in the future, it does the exact same thing. Wouldn’t it just become an “intelligent garbage dump”?

This is also a blind spot in knowledge. Current LLMs still rely too heavily on human intervention for self-correction. This is probably why I noticed its release notes specifically mentioning a summary mechanism for “FTS5 session search and cross-session recall.” Perhaps they are trying to use another layer of large models to periodically clean up these junk experiences.

If this logic holds up, the future MCP (Model Context Protocol) ecosystem is going to be even wilder. You plug in an external GitHub tool, and not only can it call it, but it can also figure out your coding habits on its own.

Or maybe I’m overthinking it. A lot of times, the spaghetti code written by human programmers isn’t much better than AI’s anyway (o_O)

Food for Thought

By the way, today’s coffee beans are really over-roasted; the burnt bitterness is rushing straight to my head.

Looking at the line The agent that grows with you on my screen, I can’t help but feel that we might truly be witnessing the end of a certain era of classical software. In the future, when we buy software, we might no longer be buying hard-coded logic, but a “colleague” who writes its own diary.

So, if one day your “colleague” quietly puts a configuration you broke into its blacklist, how are you going to reconcile with it?


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *