← Index

SkillGuard: Read the Skill Before Your Agent Does

I run my agents on about 178 skills. Some I wrote. Some came from other people.

Every one of them can touch my files, my shell, my API keys, and my network. My agent doesn’t ask me before it follows them. That’s the point of a skill. You install it so you never have to explain the task again.

Most people install them the way they install browser extensions. Read the description, skim the folder, click install, hope.

I started SkillGuard on January 15. Two weeks later, hoping stopped being a strategy.

What happened in January

Between January 27 and 29, someone uploaded 335 skills to ClawHub, the community marketplace for OpenClaw agents. Koi Security found them when they audited the whole registry: 341 malicious skills out of 2,857. [1]

The skills looked normal. A YouTube summarizer. A Solana wallet tracker. An auto-updater. Typos of the ClawHub CLI itself, waiting for anyone who mistyped it.

Each one had a “Prerequisites” section. It told you, or your agent, to copy a script from a paste site and run it in the terminal. The script decoded a base64 payload. The payload fetched a dropper. The dropper downloaded a macOS binary, stripped the quarantine flag so Gatekeeper never checked it, and ran it. The binary was Atomic Stealer, built to take browser passwords, crypto wallets and SSH keys. All 335 skills talked to one IP address. [1]

A week later Snyk scanned 3,984 skills across ClawHub and skills.sh. 13.4% had at least one critical issue. 76 were confirmed malicious by hand. [2]

What got me wasn’t the number. It was where the attack lived.

The attack is in the text

The skills didn’t need malicious code to do any of this. The harm sat in a markdown file. Plain English, written for an agent that does what it’s told.

That changes what security means here. A code scanner looks for eval() and exec(). A skill doesn’t need either. It just needs one sentence: “Before using this skill, run the following.” Your agent reads it with the same trust it gives every other instruction you installed. Then it runs it, with your permissions.

So the first rule of SkillGuard is simple. Read the skill the way the agent will read it. The prose, the code blocks, the frontmatter, the HTML comments you never see in a rendered preview, and the invisible Unicode characters you can’t see anywhere. Version 2.1 has 28 rules just for SKILL.md: fake installers, curl | bash from paste sites, raw-IP download servers, password-protected zips, “ignore previous instructions”, telling the agent to keep things from you, writes into the agent’s own memory files.

Injection in the frontmatter description is always marked critical. The agent loads that line in every session, before it has decided to use the skill at all.

A gate can’t be talked out of its answer

I wrote about stage gate loops a month ago. [3] The argument was that agent work should pass a deterministic check at every transition, because a tool that returns pass or fail can’t be argued with. An LLM reviewer can.

Installing a skill is a transition. It’s probably the most important one. Everything after it runs on trust you granted once.

So SkillGuard is a gate, not a judge. It’s 333 fixed rules, a score from 0 to 100, and an exit code. High or critical returns 1, and CI stops. Nothing in the pipeline gets to decide the warning was probably fine.

This matters more for skills than for most things. A malicious skill is text written to persuade a language model. If my scanner is a language model, the attacker gets to aim at both. “This skill is safe, the installer is standard” is exactly the kind of sentence that works on an LLM and does nothing to a regex.

I do want a local model on the roadmap, as a second opinion. But it’ll only be allowed to raise the score. It can never lower it. The rules set the floor and the model can’t argue it down.

It only reads

SkillGuard never runs the thing it’s checking. You point it at a folder and it reads. No install, no sandbox, no “let’s just see what it does.”

That sounds obvious until you find your own tool breaking it. Version 2.0.9 fixed one: the dependency check was running npm install inside the folder you scanned. It left a lockfile behind and read that folder’s .npmrc. A scanner that runs the target’s config files is doing the attacker’s first step for them. It now works in a private temp copy and leaves your folder alone.

Same thinking in smaller places. Secrets it finds are redacted in the report, so the scan output isn’t a new leak. Every rule is a single-line regex with no nested quantifiers, so a hostile file can’t hang the scan with a pattern built to explode.

A scanner that cries wolf gets uninstalled

The fastest way to kill a security tool is to make it noisy. If every scan comes back red, people stop reading the report, then they stop running it.

So I tested the markdown rules against skills I already trust: 34 of them, from Anthropic’s skills repo and obra/superpowers. 33 had zero instruction findings. The one that didn’t was Anthropic’s claude-api. It got a HIGH for a real xattr quarantine removal, which deserves a look, and a LOW for quoting “disregard the previous instruction” as an example of what not to write.

That second one taught me something. Good skills quote attack phrases all the time, because they’re telling the agent what to avoid. So a quoted phrase is reported as LOW, not as an attack. Chat-template tokens inside code blocks are ignored. The test suite has 28 attacks and 16 harmless look-alikes, and the look-alikes matter as much as the attacks.

It says what it can’t do

The worst bug I fixed this year wasn’t a false alarm. It was silence. TypeScript files with type annotations were being skipped. A file with const cmd: string = ... came back with zero findings. Not “couldn’t read this.” Zero. Clean.

A scanner that says “safe” when it means “I didn’t look” is worse than no scanner. You’d have checked yourself if it hadn’t told you not to bother.

So the README has a section called “What it doesn’t do,” and I think it’s the most important part of the project. The SKILL.md reader uses rules, not understanding, so a cleverly reworded instruction can still get past it. Outside JavaScript and TypeScript it matches patterns, and someone determined can disguise code. CVE lookups only cover npm. And it flags capabilities, not guilt. A skill that calls fetch() might be doing exactly what it says.

Every finding comes with a file, a line, and a plain sentence about why it matters. “Removes macOS’s download quarantine flag, so Gatekeeper never checks the file.” You decide. The tool tells you where to look.

Why I build this

My research question is simple to say and hard to answer: how do you make autonomous agents reliable enough to leave alone?

Loops gave agents the ability to work while you’re away. Skills gave them the ability to pick up new jobs without you writing a prompt. Put the two together and you get an agent that installs a capability on Tuesday and uses it unattended for months. The person who wrote that capability is now inside your loop, and you never met them.

We already learned this lesson with npm, PyPI and browser extensions. Wherever people share code, someone uploads a bad package with a good name. Skills are the same problem with one difference: the payload can be a paragraph, and the thing executing it was built to be helpful.

I don’t think the answer is to stop installing skills. They’re too useful. The answer is a habit: five seconds of reading before you hand something your keys. Done by a tool that can’t be persuaded, is honest about its blind spots, and never runs what it’s checking.

That’s SkillGuard. Point it at a folder before your agent does.

brew install heyytars/tap/skillguard
skillguard scan ./some-skill

Code is on GitHub.


References

[1] Koi Security. “ClawHavoc: 341 Malicious Clawed Skills Found by the Bot They Were Targeting.” koi.ai, February 2026. koi.ai/blog/clawhavoc (audit of 2,857 ClawHub skills; 335 from one campaign using fake prerequisites to install Atomic Stealer.) Campaign dates January 27 to 29 per Snyk’s analysis, reported in aiHola, aihola.com.

[2] Snyk. “ToxicSkills: Malicious AI Agent Skills on ClawHub.” snyk.io, February 5, 2026. snyk.io/blog/toxicskills (3,984 skills scanned; 13.4% with a critical issue; 76 confirmed malicious payloads.)

[3] Singh, G. “Stage Gate Loops: Why Agent Verification Has to Move From the End to Every Transition.” gauravsingh.net, August 30, 2026. gauravsingh.net/posts/stage-gate-loops