Problem
No analyzer looks at image content bundled or cited in skills:
references.py treats markdown images as passive grammar,
mcp_tool_poisoning flags base64 blobs in metadata only, and the
LLM path is text-only (no image_url in any provider). SVG is
already visible to static patterns as text; PNG/JPG cited or bundled
are unscanned hiding places for payload text in metadata layers.
Proposed scope
- Per-skill inventory of LOCAL image files (markdown targets plus
loose files) in the inspection ledger, no verdicts. Remote-cited
URLs are inventoried but never fetched: no network access
mid-scan, by design.
- Stdlib-only extraction of embedded text layers (PNG text chunks
and EXIF, JPEG comments and EXIF, GIF comments) routed through
the existing static prompt-injection patterns, with findings
attributed to the image path. Bounded, never raises, no pixels,
no new dependencies.
- Omitted deliberately: pixel decoding (no OCR), remote
fetching, GIF plain-text extensions, WebP EXIF, BMP/WebP text
layers. Rendered QR payloads and steganography need a vision-LLM
design, proposed separately — not claimed here.
Acceptance
- Inventory present in the ledger for local files; remote URLs
cited-but-unfetched.
- Extracted image text flows through static patterns with tests;
textless images keep existing skip accounting.
Problem
No analyzer looks at image content bundled or cited in skills:
references.pytreats markdown images as passive grammar,mcp_tool_poisoningflags base64 blobs in metadata only, and theLLM path is text-only (no
image_urlin any provider). SVG isalready visible to static patterns as text; PNG/JPG cited or bundled
are unscanned hiding places for payload text in metadata layers.
Proposed scope
loose files) in the inspection ledger, no verdicts. Remote-cited
URLs are inventoried but never fetched: no network access
mid-scan, by design.
and EXIF, JPEG comments and EXIF, GIF comments) routed through
the existing static prompt-injection patterns, with findings
attributed to the image path. Bounded, never raises, no pixels,
no new dependencies.
fetching, GIF plain-text extensions, WebP EXIF, BMP/WebP text
layers. Rendered QR payloads and steganography need a vision-LLM
design, proposed separately — not claimed here.
Acceptance
cited-but-unfetched.
textless images keep existing skip accounting.