Skip to content

[Bug] Risk score saturates to 100 on any multi-file skill due to unbounded additive scoring #134

Description

@mimran-khan

Summary

Imagine a fire alarm that goes off if it detects smoke from a candle, and then goes off again — louder — for every additional candle in the building. By the time you have 10 birthday candles lit, the alarm is screaming "EVACUATE IMMEDIATELY" at the same volume as an actual building fire. You can't tell the difference.

That's what happens with SkillSpector's risk scoring today. The score is a simple sum: every pattern match adds points, with no cap on how many times the same rule can contribute. A skill that uses subprocess in 5 different files (perfectly normal for a multi-file automation skill) gets scored the same as a skill that contains a genuine remote code execution vulnerability. Both hit 100/100 CRITICAL.

This makes the score meaningless for any real-world skill. A legitimate deployment skill with 10+ files will always saturate to CRITICAL, no matter how safe it actually is. Teams can't use SkillSpector as a CI gate because everything fails, and they can't triage by score because every skill looks equally dangerous.

Reproduction

Create a minimal multi-file skill:

test-skill/
├── SKILL.md
├── main.py
├── utils.py
├── helpers.py
├── deploy.sh
└── config.py

SKILL.md:

---
name: test-skill
description: A test automation skill
---
# Test Skill
This skill automates deployment tasks.

main.py:

import subprocess
subprocess.run(["curl", "-k", "https://internal-api/deploy"])
subprocess.run(["rm", "-rf", "/tmp/cache"])

utils.py:

import os
os.system("curl --insecure https://api.example.com/status")
os.environ["SECRET_KEY"] = "placeholder"

helpers.py:

import subprocess
result = subprocess.check_output(["git", "reset", "--hard", "HEAD"])
subprocess.Popen(["wget", "http://remote-host/script.sh"])

deploy.sh:

#!/bin/bash
curl -k https://api.example.com/deploy
rm -rf /tmp/build-cache
eval "$USER_INPUT"

config.py:

import os
os.system("pip install " + package_name)
skillspector scan ./test-skill/ --no-llm
# Score: 100/100 CRITICAL — but this is a contrived example with moderate issues

Even a skill with only MEDIUM-severity findings across 5 files can saturate to 100 because each match adds full points without regard to rule repetition.

Root Cause

The scoring formula in src/skillspector/nodes/report.py:

for f in findings:
    sev = (f.severity or "LOW").upper()
    if sev == "CRITICAL": score += 50
    elif sev == "HIGH": score += 25
    elif sev == "MEDIUM": score += 10
    elif sev == "LOW": score += 5
if has_executable_scripts:
    score = int(score * 1.3)
score = min(100, max(0, score))

Every finding contributes full points regardless of whether the same rule_id has already been counted. A single rule firing repeatedly across files (normal for multi-file skills) inflates the score unboundedly before clamping.

Example math

  • 5 files × 3 MEDIUM findings each = 15 findings × 10 points = 150 → clamped to 100
  • With has_executable_scripts multiplier: 150 × 1.3 = 195 → still clamped to 100
  • A single HIGH finding (score 25) is indistinguishable from a skill with 20 LOWs (score 100)

Impact

  • False CRITICAL verdicts on legitimate multi-file skills
  • No discrimination between a skill with 1 genuine critical vulnerability vs. one with many low-confidence repeated pattern matches
  • Renders the scanner unusable as a CI gate for real-world skill repositories where skills typically have 5–20+ files
  • Score ceiling reached too easily: Only 4 MEDIUM findings needed to reach "HIGH" (40+), 10 MEDIUMs for "CRITICAL" (100)

Suggested Fix

Options (not mutually exclusive):

  1. Per-rule cap: Same rule_id contributes at most N times to the score (e.g., max 3 instances of TM1 count).
  2. Unique-rule scoring: Score only the highest-severity instance per rule_id per file.
  3. Normalize by skill size: Divide raw score by log2(num_components) or similar to account for larger skills naturally having more matches.
  4. Weighted confidence: Multiply each finding's score contribution by its confidence (0.0–1.0) so low-confidence matches contribute less.
  5. Tiered scoring bands: Instead of linear addition, use diminishing returns (e.g., 2nd instance of same rule adds half points, 3rd adds quarter, etc.).

Affected Version

SkillSpector v2.2.3

Activity

  1. changed the title [-]Risk score saturates to 100 on any multi-file skill due to unbounded additive scoring[/-] [+][Bug] Risk score saturates to 100 on any multi-file skill due to unbounded additive scoring[/+] on Jun 22, 2026
  2. added a commit that references this issue on Jun 22, 2026
    8a97cce
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions