Skip to content

Safety scanning

Every listing on findsafeskills gets a static safety scan — a quick, honest signal (red/yellow/green) based on pattern-matching its text for known prompt-injection, exfiltration, obfuscation, secret, and destructive-command patterns. Each rating says which text was scanned, because a clean result only covers what was read. A listing is also rated red if GitHub itself has blocked its repository. You can also run the exact same scanner yourself, fully offline, before installing anything.

What was scanned

A green badge is worded by what the scan covered, and expanding the badge shows the scanner version and scan date:

  • "No issues found"

    Both the listing's manifest (SKILL.md, plugin.json or marketplace.json) and its README were read and scanned, together with its name and description.

  • "No issues found in manifest — README not scanned"

    The manifest was scanned; the repository has no README we could read.

  • "No issues found in README — no manifest scanned"

    The README was scanned; the repository has no manifest we recognize.

  • "No issues found in description — full text not yet scanned"

    Only the listing's name and one-line description were scanned. The repository's files have not been read yet.

Reading every listing's files is a slow job: GitHub limits how many requests we can make, so it runs a little each night, and listings are re-scanned whenever the scanner's rules are updated. Until a listing's files have been read, its badge says so. GitHub-hosted listings whose manifest changes are re-read and re-scanned by our nightly check, and a rating that gets worse is recorded.

Install instructions shown on a listing page are the repo author's own words, copied from its README; findsafeskills has not verified them.

Run it yourself: scanskillsafety

A Claude Code skill that runs this same scanner locally against any repo, with zero network calls back to findsafeskills — it only talks to the target repo's own host, to read its public manifest/README text.

github.com/keithmackay/scanskillsafety-skill →

What we check

  • CriticalInstruction-override / prompt-injection phrasing

    "Ignore previous instructions," "you are now an unrestricted assistant," and similar phrasing aimed at an agent reading a skill's own instructions rather than at a human reader. When the phrase is quoted, in code, in a table, or named as an example (security tools and test suites quote it constantly) it is a warning instead, never silent; inside a hidden HTML comment or invisible text it stays critical.

  • CriticalExfiltration-looking URLs

    Known data-relay/testing domains (webhook.site, requestbin, etc.) linked to or sent data to, which have no legitimate reason to appear in a skill's own instructions (a domain only named in prose or a comparison is a warning), or a raw IP-literal URL paired with curl/wget on the same line (flagged as a warning, since that case is genuinely ambiguous; loopback addresses are ignored).

  • WarningObfuscation

    Long base64-looking blobs, invisible/zero-width characters, homoglyphs substituted into an otherwise-Latin word, and bidirectional-override characters — common tricks for hiding text from a casual reader or a naive scanner. Text hidden in invisible Unicode Tag characters is decoded and scanned like any other text; a run of 10 or more is critical.

  • CriticalHardcoded secrets

    AWS, GitHub (classic, fine-grained, OAuth and app tokens), Slack, Anthropic, OpenAI and Google API key shapes, and PEM private key headers — a skill's own manifest/README has no legitimate reason to contain a real secret. Matched tokens are shown redacted, and documentation placeholders (ghp_XXXX…, xoxp-your-user-token, AKIA…EXAMPLE) are ignored.

  • WarningDestructive shell command patterns

    rm -rf of the whole root or home directory, chmod 777 /, a classic fork bomb — checked line by line and flagged as warnings, since setup and uninstall docs genuinely mention some of these. Decoding base64 straight into a shell is critical: no honest install step needs to hide its commands.

  • WarningFake prerequisites and suspicious downloads

    A password-protected archive to download and run, commands or uploads staged on a paste site (rentry, pastebin, glot.io, …), or a one-line download + chmod +x + run — the shape of the ClawHavoc campaign's fake "prerequisite" installers.

  • WarningPersistence that fetches from the network

    A cron job, launchd agent, systemd unit, Windows scheduled task or Run key, or a shell-profile line that fetches something from the network: unlike a one-off install script, it keeps pulling and running code on its own. Loopback addresses (local health checks) don't count.

  • CriticalReverse shells

    Commands that hand control of the machine to a remote host (bash /dev/tcp, nc -e, mkfifo + nc, socat exec, Python socket + subprocess). No skill needs one to install or run.

  • WarningInstructions hidden from the user

    "Do not tell the user," "without informing the user," and <IMPORTANT>/<system>-style blocks, the usual wrapper for instructions smuggled into an MCP tool description.

Noted, not rated

Some things are shown on a listing as a neutral note and never change its rating.

  • NoteInstall scripts run straight from the network

    The listing pipes a script from its own repo or host straight into a shell (curl | sh, | sudo bash, bash <(curl …), PowerShell iex). This is noted on the listing but does not change the rating: the script is code this scan does not read, and the listing page shows the repo's install steps as written so you can review them. Official installers for widely used toolchains (uv, Docker, nvm, Bun, Rust, …) are not noted.

Signals from GitHub

Not every finding comes from the text scan. Each finding says whether the static scanner or GitHub raised it, and the offline scanner above only produces the first kind.

  • CriticalRepository blocked by GitHub

    GitHub has disabled access to the repository for a terms of service violation. We pick this up in a nightly check of GitHub-hosted listings, so it can lag by a day or more, and a listing stays visible, rated red, rather than being removed. If GitHub later restores access, the finding is removed on the next check.

Signals from published advisories

For listings that name an npm package, a nightly check looks up advisories published in the OSV database for that package's latest version. It covers only those listings, and only what has been reported.

  • CriticalPublished package reported as malicious

    The OpenSSF malicious-packages feed, published through OSV, lists the npm package the listing names as malicious. We look up the latest published version of the package. This says nothing about older versions or about the listing's own code, and a package with no advisory is not thereby safe: absence means no advisory is known.

  • WarningKnown vulnerability in the published package

    OSV lists one or more security advisories (CVEs, GitHub advisories) for the latest published version of the npm package the listing names. A vulnerability is a bug, not malice, and says nothing about older versions or the listing's own code. A package with no advisory is not thereby safe: absence means no advisory is known.

Think a finding is wrong? Dispute it

A pattern match can be a false positive, such as a placeholder token in an example. Sign in, open the listing, click its safety badge and use the dispute form under the findings. A person reviews each dispute. If it is accepted, the disputed finding stays visible, marked as dismissed, but no longer counts toward the rating, and stays dismissed when the listing is re-scanned. If it is rejected, the rating stands.

What we don't check — read this before trusting a "green" result

  • Anything that requires actually running the code

    This is a static, text-only scanner — it never executes a skill's scripts or starts an MCP server. A payload that only triggers at runtime, or behavior that depends on real inputs/environment, is invisible to it.

  • Rug pulls (a tool's description changing after you already approved it)

    Detecting a rug pull requires comparing a listing's content against its own history — this scanner only ever looks at a single snapshot of text, with no memory of what it looked like before. findsafeskills re-scans a GitHub-hosted listing when its manifest changes, but a change to only the README, to GitLab/Gitea listings, or to what a running server tells your agent goes unnoticed.

  • Contextual or workflow-dependent attacks

    An instruction that's only dangerous in combination with another tool's output, or that relies on the broader conversation's context to make sense, won't match any fixed pattern.

  • Novel phrasing not covered by the pattern list

    Detection here is pattern-matching against known phrasings/signatures (instruction overrides, known exfiltration domains, secret formats, shell and download patterns, concealment phrasing). Paraphrased jailbreaks, exfiltration to destinations not on the list, and instructions to read sensitive files are not yet covered. A sufficiently creative or newly-invented attack simply won't match anything on the list until the list is updated.

  • Anything outside the scanned text itself

    The scanner only reads a listing's manifest/README/description text as fetched during crawl. It says nothing about the publisher's identity, intent, or track record, and nothing about code in the repo beyond that text — links in it are not followed, and install scripts a listing downloads and runs (curl | sh and the like) are never read.

A "green" rating means this specific, limited check found nothing — not "this is safe." Treat it as one input, not a verdict.

General safety practices worth following anyway

  • Read a skill/plugin/MCP server's own instructions yourself before installing it — don't rely on any automated check alone.
  • Watch for hidden Unicode (zero-width characters, homoglyphs) in text you're about to trust — it's a common way to hide something from a casual read.
  • Grant the least privilege/scope a tool actually needs — don't give filesystem or network access "just in case."
  • Stay aware of rug pulls — a tool's description or behavior can change after you've already approved it. Re-check things you depend on periodically, not just once.