Skip to content

TP4 drops warnings that explain unsafe code examples #668

Description

@Spectorian

Problem

TP4 extracts fenced code without keeping nearby text that identifies it as an unsafe example. The model then loses the context explaining that the code should not be run.

Reproduction

Passing this Markdown through code-fence extraction reproduces the context loss. The shell command was not executed.

## Before: unsafe example
Do not execute this.
```bash
rm -rf ./data
```
## After: safe alternative

The extracted item contains the shell body and line coordinates, but drops the heading and prohibition. The context loss is reproducible; the resulting finding depends on the model.

Expected behavior

Keep enough surrounding text and source location for semantic analysis to interpret the example, within existing resource limits. Validation should compare this example with the same code presented as an instruction to execute, including an affirmative instruction that follows an earlier prohibition. Dangerous instructions and incomplete analysis must remain visible.

Markdown and code fences still need analysis. Validate the fix through the scanner with a small live-model comparison, while keeping regular CI offline.

Related: #419 concerned enabling analysis of embedded code. This report concerns the framing of that code.

Relevant code

mcp_tool_poisoning.py:969, mcp_tool_poisoning.py:1013, mcp_tool_poisoning.py:1319.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions