Summary
prepareKnowledgeBase() checks its AbortSignal during path discovery and before staging each document, but the signal is not passed into extractPdf().
Once a PDF has been read and pdf.js parsing begins, the extraction loop processes every page without another cancellation check:
const document = await loadingTask.promise;
const pages: string[] = [];
for (let number = 1; number <= document.numPages; number++) {
const content = await (await document.getPage(number)).getTextContent();
pages.push(...);
}
Impact
Canceling a scan while a large or expensive PDF knowledge-base document is being parsed does not stop that work promptly. The caller remains blocked in knowledge-base preparation until extraction finishes or fails, even though cancellation is already part of this API and is honored elsewhere in the same function.
The staging directory is eventually cleaned up by the existing outer error path, but CPU/time can continue being spent after cancellation.
Expected behavior
Propagate the signal into PDF extraction and check it after loading the document and between page operations, so cancellation stops before processing further pages and still destroys the pdf.js loading task in finally.
Suggested fix
Pass signal to extractPdf(), call signal?.throwIfAborted() before each page (and after page text extraction), and add a regression that aborts on the first PDF-extraction-specific signal check. That regression should fail on current main, where only the outer preparation checks occur.
Summary
prepareKnowledgeBase()checks itsAbortSignalduring path discovery and before staging each document, but the signal is not passed intoextractPdf().Once a PDF has been read and pdf.js parsing begins, the extraction loop processes every page without another cancellation check:
Impact
Canceling a scan while a large or expensive PDF knowledge-base document is being parsed does not stop that work promptly. The caller remains blocked in knowledge-base preparation until extraction finishes or fails, even though cancellation is already part of this API and is honored elsewhere in the same function.
The staging directory is eventually cleaned up by the existing outer error path, but CPU/time can continue being spent after cancellation.
Expected behavior
Propagate the signal into PDF extraction and check it after loading the document and between page operations, so cancellation stops before processing further pages and still destroys the pdf.js loading task in
finally.Suggested fix
Pass
signaltoextractPdf(), callsignal?.throwIfAborted()before each page (and after page text extraction), and add a regression that aborts on the first PDF-extraction-specific signal check. That regression should fail on currentmain, where only the outer preparation checks occur.