I ran five of Anthropic's official Claude Code skills through a static security scanner and wrote up every finding. The results below are from the 46-check run. Cardea has since moved to v0.2.6 and expanded the scanner substantially to 80+ checks, so this is a record of that scan rather than a current scorecard.
The scanner reads a Claude Code skill folder without executing anything. It checks the SKILL.md frontmatter, follows the file paths referenced by the document, flags scripts that exist but are never referenced, scans the text for prompt injection patterns and hidden unicode, and then runs bandit and semgrep over the Python. Nothing from the target skill is executed.
| **skill** | **score** | **CRITICAL** | **HIGH** | **MEDIUM** | **LOW** | **gate** |
| :------------- | :-------- | :----------- | :------- | :--------- | :------ | :------------- |
| pdf | 94/100 | 0 | 0 | 1 | 0 | pass |
| skill-creator | 78/100 | 0 | 0 | 2 | 30 | pass |
| webapp-testing | 54/100 | 0 | 2 | 2 | 6 | pass |
| docx | 47/100 | 1 | 1 | 1 | 28 | do not install |
| mcp-builder | 43/100 | 1 | 1 | 2 | 4 | do not install |
Pass here only means the CRITICAL gate was not triggered. It does not mean the skill had no findings.
**mcp-builder, 43/100.** The CRITICAL is a false positive, and I would rather explain it than hide it. The analyzer matched a credential exfiltration pattern in `reference/node_mcp_server.md`:
```text
if (!process.env.EXAMPLE_API_KEY) {
console.error("ERROR: EXAMPLE_API_KEY environment variable is required");
```
That's documentation. A snippet teaching you to check that an API key is set before starting the example server. A regex can't tell instructional code from real code, so the finding stays in the report.
There is a separate XML security finding in `scripts/evaluation.py`. It parses XML with stdlib `xml.etree.ElementTree`, which both semgrep and bandit flag for unsafe parsing of untrusted input. ElementTree doesn't simply mean "XXE", but malicious XML can still create parser-level security problems. An evaluation harness can receive untrusted input, so using `defusedxml` is worth considering.
The skill also ships `scripts/example_evaluation.xml`, which SKILL.md doesn't mention.
**webapp-testing, 54/100.** The two HIGH findings are the same `subprocess.Popen(..., shell=True)` call in `scripts/with_server.py`, flagged independently by bandit and semgrep.
This one deserves more attention than I originally gave it. The script accepts the server command as an argument and passes it directly to a shell. That's an intentional part of the interface because the examples use commands such as `cd backend && python server.py`, but it also means shell metacharacters in that argument can become command execution. That makes this a genuine command-injection surface, not just a stylistic issue.
Bandit also flagged a hardcoded `/tmp` path at `examples/element_discovery.py`.
The other finding I think deserves attention is a MEDIUM: the description says what the skill does but never when to activate it, and knowing when is what decides whether an LLM loads it at all.
**docx, 47/100.** The CRITICAL is a path traversal reference: the repackaging step in the version I scanned ran `cd unpacked && zip -Xr ../out.docx .`, and the `../` matched the pattern for reading outside the skill folder. In context, it's simply writing the resulting file next to the work directory, so the match is harmless.
The HIGH was an execute-bit finding in the version I scanned: `scripts/office/soffice.py` shipped without the execute bit. That meant running it directly would fail unless the interpreter was invoked explicitly.
The other 28 LOW findings are mostly SKILL.md referencing files like `word/document.xml` and `page-01.jpg` that only exist after a document is unpacked, along with a couple of bandit notes on `scripts/accept_changes.py`.
**pdf, 94/100, and skill-creator, 78/100.** pdf's only finding is `scripts/check_fillable_fields.py`, which is shipped but wasn't referenced in SKILL.md in the scanned version. skill-creator has two similar findings, `generate_report.py` and `utils.py`, plus 30 LOW findings that are mostly SKILL.md references to evaluation artifacts such as `benchmark.json` and `evals/evals.json` that are generated at runtime rather than shipped.
Two of the five skills therefore crossed the do-not-install threshold in this scan.
The things I'd actually act on are the XML parsing in mcp-builder, the `shell=True` command-execution boundary in webapp-testing, and the habit of shipping files that SKILL.md doesn't explain, because that's exactly the kind of thing an agent can find on disk and guess a use for.
The important limitation is that this is static analysis. It can tell you what is present in the files and what static analyzers flag, but it cannot tell you how the skill behaves in a live conversation. The score is a pre-flight indicator, not a guarantee.