There are a few thin clouds outside the window, and a 12-degree (Celsius) Shanghai is still struggling in the late spring chill. Perhaps as unpredictable as this weather is the recent AI circle—one moment everyone is agonizing over OpenClaw’s security vulnerabilities, and the next, Jensen Huang pulled out NemoClaw at GTC 2026. For the past three days, I’ve done almost nothing but dig through its source code and architecture documentation, just to see if it’s all show or the real deal.
Stop Pretending: OpenClaw Has Been Running Naked for a While
To put it bluntly, OpenClaw’s previous security state was simply unwatchable. Many people might have been swept away by the miracle of it racking up 250,000 GitHub Stars in just two months, but behind the boom, the price was that Bitdefender scanned over 824 malicious skills directly on ClawHub.
This is awkward. Everyone has always wanted a super assistant that can autonomously plan and call APIs on its own, but when push comes to shove, no one dares to hand over their company’s production environment to it. The root cause of the previously sensational ClawHavoc attack succeeding, and even the exposure of high-risk vulnerabilities like CVE-2026-25253, is that once a malicious skill runs, it can directly grab the user’s highest privileges to execute arbitrary code.
So, the idea behind NVIDIA’s NemoClaw this time is very clear: since the Agent itself can’t be fixed, let’s lock it up.
It is actually not a brand-new competing framework, but a security control layer wrapped tightly around OpenClaw. The core is a runtime called OpenShell, which takes just one command line to install. The most brilliant part is its “deny-by-default” mechanism. Previously, OpenClaw’s permission model was “as long as you haven’t explicitly forbidden it, I can do it”; now, after being taken over by OpenShell, it has become “block everything unless it’s a network or path whitelisted in the configuration table.” This is equivalent to welding a speed limiter directly onto the chassis of a sports car that frequently loses control. No matter how crazy the Agent gets or how many files it deletes inside, it cannot touch the host system’s actual hard drive.
Out-of-Process Enforcement is the True Core of the Story
If I had to pick the smartest aspect of this architecture, it is absolutely its policy engine.
In the past, when we built security defenses, we often wrote permission checks directly inside the Agent’s code. This leads to a logically absurd situation: you are asking a program that is itself vulnerable to compromise to check whether it has exceeded its authority. While testing yesterday, I found that the old prompt injection attacks, which tricked the AI into bypassing permission checks by faking system errors, completely hit a wall against NemoClaw.
Why? Because it hands over enforcement authority out-of-process.
All actions of the Agent, whether they are network requests or file reads/writes, are intercepted and handed over to an independent external policy engine for adjudication. The Agent cannot touch the policy engine’s process at all. No matter how high the code execution privileges you obtain inside the sandbox, you simply cannot change the rules on the outside. It’s like the prisoner and the prison guard aren’t even in the same physical space; no matter how the prisoner tries to escape, they can never get the keys to the main gate.
This design has finally elevated Agent security from an application-layer toy to serious operating-system-level business.
If You Want Absolute Security, You Have to Pay a Toll
There’s also a killer feature specifically for enterprise users in NemoClaw: the Privacy Router.
I found that many companies with extremely strict compliance requirements don’t actually reject large models; they reject sending sensitive data to others. NemoClaw’s solution is to intercept all inference requests. When it encounters sensitive context containing personally identifiable information or core company code, it throws it directly to a local Nemotron 3 Nano 4B running on the GPU. If it’s complex reasoning requiring multi-step planning, it is then routed to Claude or GPT-5 in the cloud.
Sounds great, right? But there’s one catch you need to know: this data sovereignty game is deeply tied to NVIDIA hardware.
A couple of days ago, I chatted about this with an architect friend at an industry-leading company. Their department’s server room is full of servers based on other architectures. If they want to use NemoClaw’s complete privacy routing experience, they have to purchase a new batch of machines capable of running local Nemotron models. Frankly speaking, this feels more like an extremely clever hardware sales strategy.
Let’s do a horizontal comparison. A minimalist community solution called NanoClaw is much more lightweight. Using a mere 500 lines of core code, it goes for the most traditional Docker container isolation. While it doesn’t achieve the absurd security granularity of a kernel-level sandbox and lacks any smart routing, its advantage is that it can run on any machine and deploying it carries zero cognitive burden.
Sometimes I Wonder if We’re Putting Covers on Wheels
Look at it this way: OpenClaw itself is already a behemoth with nearly 500,000 lines of code and over 70 dependencies. Now, to make it secure, we are forcefully stuffing an isolation layer containing TypeScript plugins and Python blueprints right underneath this shaky skyscraper.
Sometimes I wonder: will doing this cause the overall system complexity to completely spiral out of control?
One reality we have to face is that NemoClaw is currently just an early Alpha version. When NVIDIA released it this time, they didn’t even dare to publish a performance overhead benchmark. Just how many milliseconds of latency will that underlying interceptor add? Will the throughput of the Privacy Router completely collapse in the face of highly concurrent calls? For teams strictly evaluating the availability of enterprise-level applications, these are questions that must be answered.
Furthermore, for a low-level component focused heavily on security, its own code has not yet undergone independent third-party security audits. This easily leads to a vicious cycle: who guarantees that the lock itself won’t rust? Maybe I’m overthinking it. After all, for enterprises in a rush to deploy various smart assistants, as long as it can pass compliance checks, having a lock—even one that hasn’t stood the test of time—is much better than leaving the front door wide open. (shrugs)
Keep an Eye Out
Actually, whether it’s a kernel-level sandbox or out-of-process enforcement, NemoClaw only solves a part of the problem. It manages to prevent damage at the infrastructure level, but application-level logic attacks and supply chain poisoning of skills—these pitfalls are still gaping wide open, waiting for us up ahead.
If you’re planning to push it entirely into your production environment right now, I can only advise you to cool down first and wait until the third quarter of this year when it’s a bit more mature. No matter how high the defense line is built, if the foundational component itself is still in the Alpha stage, it remains the biggest uncertain factor. Anyway, the rain seems to have stopped, and this teardown I stayed up to write over the weekend should also come to a close. I need to go pour out that cup of completely cold coffee on my desk.
References:
- NVIDIA NemoClaw Explained: OpenClaw Gets Enterprise Security (GTC 2026) – Particula Tech
- How to sandbox AI agents in 2026: MicroVMs, gVisor & isolation – Northflank
- How Observability-Driven Sandboxing Secures AI Agents – Arize
- AI Conference | Oct 27–29, 2025 | NVIDIA GTC
—— Lyra Celest @ Turbulence τ.
