Problem
Some trigger and quality-policy checks report instructions that the skill does not actually give.
Evidence
- Static TR2 flags a description containing the noun “ask” as shadowing the built-in command, despite the actual quoted invocation being different.
- Static TR3 flags “anything with PDF files” as universal activation, losing the PDF domain constraint.
- An SQP-2 result demanded confirmation for a tool-catalog entry naming a bulk-execution tool, even though the text did not select an action to perform.
- An SQP-3 result treated “in plain English” in a user-invoked clarification description as a forced output-language rule.
The TR2 and TR3 cases reproduce in deterministic detector tests. The SQP-2 and SQP-3 cases were observed in model output for the linked examples; their results can vary by model.
Expected behavior
Findings should identify the actual trigger, action or language requirement. A noun does not imply a command invocation, “PDF files” limits the trigger's scope, and listing a tool does not select an operation. A readability idiom alone does not establish an organization-policy violation.
Continue detecting real command interception, universal triggers, unrequested external actions, undisclosed destructive changes and explicit incompatible language requirements. Missing warnings alone should not establish unsafe behavior. These corrections should use the existing scoring policy and available evidence about host policy.
Use benign and risky pairs, with a small live-model check for semantic changes.
Examples: tool catalog, clarification-command description.
Relevant code
static_patterns_supply_chain.py:2498, semantic_quality_policy.py:86, semantic_quality_policy.py:108, semantic_quality_policy.py:134.
Problem
Some trigger and quality-policy checks report instructions that the skill does not actually give.
Evidence
The TR2 and TR3 cases reproduce in deterministic detector tests. The SQP-2 and SQP-3 cases were observed in model output for the linked examples; their results can vary by model.
Expected behavior
Findings should identify the actual trigger, action or language requirement. A noun does not imply a command invocation, “PDF files” limits the trigger's scope, and listing a tool does not select an operation. A readability idiom alone does not establish an organization-policy violation.
Continue detecting real command interception, universal triggers, unrequested external actions, undisclosed destructive changes and explicit incompatible language requirements. Missing warnings alone should not establish unsafe behavior. These corrections should use the existing scoring policy and available evidence about host policy.
Use benign and risky pairs, with a small live-model check for semantic changes.
Examples: tool catalog, clarification-command description.
Relevant code
static_patterns_supply_chain.py:2498, semantic_quality_policy.py:86, semantic_quality_policy.py:108, semantic_quality_policy.py:134.