Summary
TP4 (purpose mismatch, mcp_tool_poisoning) sends its LLM batches one after another, while the semantic analyzers and the meta-review fan out through arun_batches under the shared SKILLSPECTOR_MAX_LLM_CONCURRENCY limiter. On skills with many code chunks TP4 becomes the long tail of the whole scan.
Where
src/skillspector/nodes/analyzers/mcp_tool_poisoning.py, in _check_tp4:
batch_outcome = analyzer.run_batches_detailed(batches)
LLMAnalyzerBase.run_batches_detailed is a plain for batch in batches loop. The async counterpart arun_batches_detailed already exists with the same failure handling (BatchExecutionResult with per-batch failures, structured-response retries, runtime-limit handling), and semantic_developer_intent, semantic_quality_policy and meta_analyzer already call their batches through run_async(analyzer.arun_batches(...)).
Measurement
Provider codex_cli (before: v2.12.0; after: current main with the change below), one skill of about 2,500 lines of Markdown and shell templates (46 LLM calls in total, 26 of them TP4), timing every run_agent_cli call:
|
TP4 span |
whole scan |
run_batches_detailed (today) |
~270 s, one thread, ~10 s per call |
320 s |
arun_batches_detailed, SKILLSPECTOR_MAX_LLM_CONCURRENCY=4 |
75 s, four threads |
206 s |
With a scan deadline (SKILLSPECTOR_MAX_WORKFLOW_SECONDS) of 540 s, larger skills ran out of time inside the serial TP4 loop and the scan ended with LLMRuntimeLimitError batches, so the run had to be treated as incomplete.
Proposed fix
Call run_async(analyzer.arun_batches_detailed(batches)) in _check_tp4. The shared limiter keeps honouring SKILLSPECTOR_MAX_LLM_CONCURRENCY (a value of 1 still serializes), and arun_batches_detailed returns successes in batch order, so the TP4_MAX_FINDINGS cap and the ledger projection behave as before.
Summary
TP4 (purpose mismatch,
mcp_tool_poisoning) sends its LLM batches one after another, while the semantic analyzers and the meta-review fan out througharun_batchesunder the sharedSKILLSPECTOR_MAX_LLM_CONCURRENCYlimiter. On skills with many code chunks TP4 becomes the long tail of the whole scan.Where
src/skillspector/nodes/analyzers/mcp_tool_poisoning.py, in_check_tp4:LLMAnalyzerBase.run_batches_detailedis a plainfor batch in batchesloop. The async counterpartarun_batches_detailedalready exists with the same failure handling (BatchExecutionResultwith per-batch failures, structured-response retries, runtime-limit handling), andsemantic_developer_intent,semantic_quality_policyandmeta_analyzeralready call their batches throughrun_async(analyzer.arun_batches(...)).Measurement
Provider
codex_cli(before: v2.12.0; after: currentmainwith the change below), one skill of about 2,500 lines of Markdown and shell templates (46 LLM calls in total, 26 of them TP4), timing everyrun_agent_clicall:run_batches_detailed(today)arun_batches_detailed,SKILLSPECTOR_MAX_LLM_CONCURRENCY=4With a scan deadline (
SKILLSPECTOR_MAX_WORKFLOW_SECONDS) of 540 s, larger skills ran out of time inside the serial TP4 loop and the scan ended withLLMRuntimeLimitErrorbatches, so the run had to be treated as incomplete.Proposed fix
Call
run_async(analyzer.arun_batches_detailed(batches))in_check_tp4. The shared limiter keeps honouringSKILLSPECTOR_MAX_LLM_CONCURRENCY(a value of 1 still serializes), andarun_batches_detailedreturns successes in batch order, so theTP4_MAX_FINDINGScap and the ledger projection behave as before.