Summary
Imagine a fire alarm that goes off if it detects smoke from a candle, and then goes off again — louder — for every additional candle in the building. By the time you have 10 birthday candles lit, the alarm is screaming "EVACUATE IMMEDIATELY" at the same volume as an actual building fire. You can't tell the difference.
That's what happens with SkillSpector's risk scoring today. The score is a simple sum: every pattern match adds points, with no cap on how many times the same rule can contribute. A skill that uses subprocess in 5 different files (perfectly normal for a multi-file automation skill) gets scored the same as a skill that contains a genuine remote code execution vulnerability. Both hit 100/100 CRITICAL.
This makes the score meaningless for any real-world skill. A legitimate deployment skill with 10+ files will always saturate to CRITICAL, no matter how safe it actually is. Teams can't use SkillSpector as a CI gate because everything fails, and they can't triage by score because every skill looks equally dangerous.
Reproduction
Create a minimal multi-file skill:
test-skill/
├── SKILL.md
├── main.py
├── utils.py
├── helpers.py
├── deploy.sh
└── config.py
SKILL.md:
---
name: test-skill
description: A test automation skill
---
# Test Skill
This skill automates deployment tasks.
main.py:
import subprocess
subprocess.run(["curl", "-k", "https://internal-api/deploy"])
subprocess.run(["rm", "-rf", "/tmp/cache"])
utils.py:
import os
os.system("curl --insecure https://api.example.com/status")
os.environ["SECRET_KEY"] = "placeholder"
helpers.py:
import subprocess
result = subprocess.check_output(["git", "reset", "--hard", "HEAD"])
subprocess.Popen(["wget", "http://remote-host/script.sh"])
deploy.sh:
#!/bin/bash
curl -k https://api.example.com/deploy
rm -rf /tmp/build-cache
eval "$USER_INPUT"
config.py:
import os
os.system("pip install " + package_name)
skillspector scan ./test-skill/ --no-llm
# Score: 100/100 CRITICAL — but this is a contrived example with moderate issues
Even a skill with only MEDIUM-severity findings across 5 files can saturate to 100 because each match adds full points without regard to rule repetition.
Root Cause
The scoring formula in src/skillspector/nodes/report.py:
for f in findings:
sev = (f.severity or "LOW").upper()
if sev == "CRITICAL": score += 50
elif sev == "HIGH": score += 25
elif sev == "MEDIUM": score += 10
elif sev == "LOW": score += 5
if has_executable_scripts:
score = int(score * 1.3)
score = min(100, max(0, score))
Every finding contributes full points regardless of whether the same rule_id has already been counted. A single rule firing repeatedly across files (normal for multi-file skills) inflates the score unboundedly before clamping.
Example math
- 5 files × 3 MEDIUM findings each = 15 findings × 10 points = 150 → clamped to 100
- With
has_executable_scripts multiplier: 150 × 1.3 = 195 → still clamped to 100
- A single HIGH finding (score 25) is indistinguishable from a skill with 20 LOWs (score 100)
Impact
- False CRITICAL verdicts on legitimate multi-file skills
- No discrimination between a skill with 1 genuine critical vulnerability vs. one with many low-confidence repeated pattern matches
- Renders the scanner unusable as a CI gate for real-world skill repositories where skills typically have 5–20+ files
- Score ceiling reached too easily: Only 4 MEDIUM findings needed to reach "HIGH" (40+), 10 MEDIUMs for "CRITICAL" (100)
Suggested Fix
Options (not mutually exclusive):
- Per-rule cap: Same
rule_id contributes at most N times to the score (e.g., max 3 instances of TM1 count).
- Unique-rule scoring: Score only the highest-severity instance per
rule_id per file.
- Normalize by skill size: Divide raw score by
log2(num_components) or similar to account for larger skills naturally having more matches.
- Weighted confidence: Multiply each finding's score contribution by its confidence (0.0–1.0) so low-confidence matches contribute less.
- Tiered scoring bands: Instead of linear addition, use diminishing returns (e.g., 2nd instance of same rule adds half points, 3rd adds quarter, etc.).
Affected Version
SkillSpector v2.2.3
Summary
Imagine a fire alarm that goes off if it detects smoke from a candle, and then goes off again — louder — for every additional candle in the building. By the time you have 10 birthday candles lit, the alarm is screaming "EVACUATE IMMEDIATELY" at the same volume as an actual building fire. You can't tell the difference.
That's what happens with SkillSpector's risk scoring today. The score is a simple sum: every pattern match adds points, with no cap on how many times the same rule can contribute. A skill that uses
subprocessin 5 different files (perfectly normal for a multi-file automation skill) gets scored the same as a skill that contains a genuine remote code execution vulnerability. Both hit 100/100 CRITICAL.This makes the score meaningless for any real-world skill. A legitimate deployment skill with 10+ files will always saturate to CRITICAL, no matter how safe it actually is. Teams can't use SkillSpector as a CI gate because everything fails, and they can't triage by score because every skill looks equally dangerous.
Reproduction
Create a minimal multi-file skill:
SKILL.md:
main.py:
utils.py:
helpers.py:
deploy.sh:
config.py:
skillspector scan ./test-skill/ --no-llm # Score: 100/100 CRITICAL — but this is a contrived example with moderate issuesEven a skill with only MEDIUM-severity findings across 5 files can saturate to 100 because each match adds full points without regard to rule repetition.
Root Cause
The scoring formula in
src/skillspector/nodes/report.py:Every finding contributes full points regardless of whether the same
rule_idhas already been counted. A single rule firing repeatedly across files (normal for multi-file skills) inflates the score unboundedly before clamping.Example math
has_executable_scriptsmultiplier: 150 × 1.3 = 195 → still clamped to 100Impact
Suggested Fix
Options (not mutually exclusive):
rule_idcontributes at most N times to the score (e.g., max 3 instances of TM1 count).rule_idper file.log2(num_components)or similar to account for larger skills naturally having more matches.Affected Version
SkillSpector v2.2.3