Skip to content

[Bug] Binary/PDF files trigger mass false positives in static pattern analysis #144

Description

@mimran-khan

Summary

Imagine hiring a proofreader to check your novel for plagiarism, but instead of giving them the manuscript, you hand them a printed photograph — and they start circling "repeated patterns" in the printer dots. That's absurd, but that's exactly what SkillSpector does with binary files.

When a skill directory contains a PDF, image, or any non-text file, SkillSpector reads it as UTF-8 with errors="replace", producing garbled decoded text. It then feeds this garbage to every static pattern analyzer. The MP2 "Context Window Stuffing" regex matches decoded JPEG stream artifacts (QE QE QE QE...) as repeated-text attacks. A single PDF file can produce 90+ false positive findings and push the risk score from 0 to CRITICAL (100), making the entire scan result meaningless.

Why This Matters — Real-World Scenario

Scenario: A compliance team publishes a security-training skill

A security team creates an AI agent skill that helps developers understand compliance requirements. They include a PDF of their organization's security policy (assets/security-policy.pdf) as reference material so the agent can cite specific sections. The skill itself is perfectly safe — it just reads questions and answers from the policy document.

They run SkillSpector as part of their publishing process:

Score: 100/100 — CRITICAL — DO_NOT_INSTALL
Findings: 89 (all MP2 "Context Window Stuffing")

The team is baffled. Their skill does nothing dangerous — it doesn't execute commands, access networks, or modify files. But the PDF's binary content (decoded JPEG stream headers) triggered 89 false alarms. They now have two choices: (a) remove the PDF reference material, making the skill less useful, or (b) ignore SkillSpector entirely because "it always cries wolf."

Either outcome defeats the purpose of having a security scanner.

Reproduction

binary-skill/
├── SKILL.md
└── assets/
    └── reference-doc.pdf   (any PDF, even a 1-page document)

SKILL.md:

---
name: binary-skill
description: A skill with PDF reference material
---
# Binary Skill
See assets/reference-doc.pdf for the full specification.
skillspector scan ./binary-skill/ --no-llm --format json -o report.json
python -c "
import json
data = json.load(open('report.json'))
print(f'Findings: {len(data[\"issues\"])}')
print(f'Score: {data[\"risk_assessment\"][\"score\"]}')
# Findings: 89 (all MP2 from the PDF)
# Score: 100 (CRITICAL)
"

All findings reference the .pdf file with matched_text containing binary garbage like QE QE QE QE... or \x00\x00\x00.

Root Cause

In src/skillspector/nodes/build_context.py:

def _walk_skill_files(skill_dir: Path) -> list[str]:
    """Walk skill directory and return sorted relative path strings."""
    paths: list[str] = []
    for item in skill_dir.rglob("*"):
        if not item.is_file():
            continue
        if any(skip in item.parts for skip in _SKIP_DIRS):
            continue
        # No binary file exclusion -- ALL files are included
        ...

And:

def _read_file_cache(skill_dir: Path, components: list[str]) -> dict[str, str]:
    for path in components:
        content = full.read_text(encoding="utf-8", errors="replace")
        file_cache[path] = content  # Binary content decoded as garbled text

There is no extension-based or content-based binary exclusion. Every file, regardless of type, is read as text and passed to all static pattern analyzers.

The MP2 regex ((\S)(?!\2).{1,19}?)\1{20,} then matches repeated byte sequences in the decoded binary as "context window stuffing."

Impact

  • Mass false positives: A single PDF produces 80-90+ MP2 findings
  • Score rendered useless: Any skill with a reference PDF is automatically CRITICAL
  • Blocks adoption in CI: Skills commonly include PDF specifications, design docs, or compliance certificates as reference material
  • Erodes user trust: Users learn to ignore scan results because "it always says CRITICAL"
  • Common binary file types affected: .pdf, .png, .jpg, .gif, .zip, .woff, .ttf, .pyc

Affected Version

SkillSpector v2.2.3

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions