Skip to content

fix(windows/capture): production hardening from a real 3-project deployment (closes #1, closes #2) + opt-in --global single-daemon mode - #3

Open
rui08984-dot wants to merge 20 commits into
EwanJasper:mainfrom
rui08984-dot:main
Open

rui08984-dot wants to merge 20 commits into
EwanJasper:mainfrom
rui08984-dot:main

Conversation

@rui08984-dot

Copy link
Copy Markdown

Context

This branch came from running SessionRelay 0.5.1 as the memory layer for three real projects on one heavy Windows machine (i7-14700HX, 3 project roots sharing one ZCode source db, ~25k captured messages). Two red-team rounds + a full source audit surfaced the root causes behind #1 and #2. Everything below is verified against your own test suite: 222 tests → 221 pass + 1 skipped, 0 failures (also defused a time-bomb in test/forget/functional.spec.ts whose absolute fixture date broke its own assertion past a 30-day boundary — the suite is fully green for the first time).

Fixes

#1 — Windows boot chain (src/cli/service.ts)

  • windowsRunScript: quote nodeAbs (C:\Program Files\nodejs contains spaces → 'C:\Program' is not recognized), and emit chcp 65001 >nul first — cmd.exe parses the file in the OEM codepage (GBK on Chinese systems) while we write UTF-8, so CJK project paths corrupted cd /d and the log redirect.
  • vbs is now written UTF-16LE with BOM: wscript only reads ANSI/UTF-16; UTF-8 vbs with CJK paths misreads as GBK (one layer deeper than the cmd bug).

#2 — daemon resource/lifecycle (src/capture/watch.ts, src/cli/watch.ts, src/adapters/zcode/index.ts)

  • Root cause of handle leak + 1GB memory sawtooth: detectCompaction opened a new better-sqlite3 connection per session per cycle and never closed it (the finally was empty — a stale "cache connection" comment copied from getConn). Now reuses getConn. healthCheck closes its probe too. Production evidence: handle slope ~378/h → flat.
  • Single instance: the file lock (check-then-write) races under concurrent spawns; production accumulated 18 instances. Added a per-project named-pipe lock (\\.\pipe\srelay-watch-<sha1(root)[:16]>) — kernel-owned, dies with the process, EADDRINUSE = already running. win32-only; other platforms unchanged.

Capture correctness / quality

  • extract.ts: decision extraction matched trigger words anywhere in a sentence and captured kanban table rows, markdown bold fragments and embedded-JSON residue. Added a plausibility guard (min length, trailing colon/pipe/bold rejection, table-row rejection, JSON residue rejection).
  • bin/srelay.ts unresolved: --json existed but was dropped in the action — always emitted human text. Now passed through.
  • cli/meta.ts decisions --limit: took the oldest N (ascending sort + slice(0, n)) while labelled "recent" — now takes the newest N; unresolved gets a 14-day fade.
  • capture/sync.ts runSync: added a per-session content-signature watermark (changeProbe: zcode = two aggregate MAX(rowid) queries per cycle; file sources = mtime+size) plus a "session still in db" guard (so rebuild/forget flows still re-ingest) and a 10-minute full-sweep backstop. Idle daemon CPU went from ~7.5% of a core to ~0.1%.
  • store/db.ts insertMessage: FTS index text capped at 100KB (jieba sync tokenize of multi-hundred-KB tool dumps stalled the event loop; full content is still stored).

Other warts found in the same audit

  • capture/sync.ts: ensureRegistered called twice (second call returns empty — custom adapter errors were never reported).
  • capture/archive.ts: --days 0 / --size 0mb treated as unset (falsy) → cutoff degraded to 9999-12-31, "archive nothing" archived everything (with --hard).
  • shared/config.ts: corrupt config.json produced a raw stack in every CLI command → readable error with repair hint.
  • cli/forget.ts: corrupt preview snapshot silently skipped the optimistic-lock diff before --yes → corrupt snapshot now refuses execution (fail-safe).
  • mcp annotate_session / save --tag: rewriting meta_text dropped decision texts → viaMeta search surface silently narrowed; now re-included.
  • adapters/claude-code/watcher.ts: runtime error events from fs.watch had no listener — a deleted/overflowed watched dir threw and killed the whole daemon; now degrades to the existing 1s polling fallback.

Opt-in feature: srelay watch --global

One daemon process adopts all projects from the existing registry (registryCandidates()), spawns per-project workers (each still holds its own per-project pipe/file locks), re-scans every 10 min to adopt new / prune removed projects, and dual-writes worker logs to each project's own watch.log (per-project --status tails stay fresh). Guards against a second global instance via a fixed-name pipe. Deployed on the 3-project machine: 3 daemons → 1, idle CPU ≈ 0, cross-project isolation verified.

Notes

  • Windows-only paths degrade gracefully (pipe lock and chcp are win32-guarded; polling fallback unchanged for Linux).
  • better-sqlite3 connections all carry busy_timeout already, so daemon+CLI+MCP concurrent writes are safe.
  • Happy to split this into smaller PRs if you prefer — the commits are logically grouped (see the fork's history), and the --global part is independent of the fixes.

Closes #1, closes #2.

Your Name added 15 commits September 22, 2026 13:59
…ngle-daemon mode

- zcode detectCompaction: reuse cached getConn (was new Database per session per cycle, never closed - handle leak + memory sawtooth root cause, upstream EwanJasper#2); healthCheck closes probe conn
- cli/service windowsRunScript: quote nodeAbs (upstream EwanJasper#1) + chcp 65001 for CJK paths; vbs written UTF-16LE BOM (wscript misreads UTF-8 CJK paths as GBK)
- capture/sync: per-session content-signature watermark (adapter.changeProbe: zcode MAX(rowid) aggregate, 2 SQL/cycle; file sources mtime+size) + known-in-db guard (rebuild/forget safe) + 10-min full sweep backstop -> idle daemon CPU 7.5% -> 0.08%
- capture/watch: per-project named-pipe single-instance lock (upstream EwanJasper#2 instance accumulation); worker extraction; --global mode: one process adopts all registry projects dynamically
…d --global process was adopting 0 projects and idling as zombie
… projects were unreleasable (sqlite/cmd held open by this process), now auto-pruned within one sweep
…obal lock released on exit); fix ensureRegistered double-call that swallowed custom-adapter load errors (upstream wart)
…决策'); upstream wart exposed by brief quality audit
…markdown fragments / embedded JSON residue) + unresolved 14-day fade in CLI
… opts.json - upstream wart the brief's text-fallback was papering over)
- mcp adopt(): close previous project's db on switch (handle leak + EBUSY vs forget/rebuild)
- archive: days/sizeMb 0 no longer treated as unset (0-days archive-everything footgun)
- loadConfig: corrupt config.json -> readable error instead of raw stack
- forget: corrupt preview snapshot now REJECTS --yes (fail-safe optimistic lock)
- annotate_session/save --tag: re-include decision texts in meta_text (viaMeta search surface)
…assertion broke past the 30-day boundary on Sep 19) - suite is fully green (222) for the first time
…h.log - per-project logs froze forever under the global daemon, making 'watch --status' print zombie tails in non-cwd projects
- watcher: runtime 'error' events (dir deleted / overflow / AV) had no listener -> EventEmitter throw kills the whole global daemon; now degrades to the existing 1s polling fallback (self-heal)
- insertMessage: index text capped at 100KB (jieba sync tokenize of huge tool-dump messages stalled the event loop seconds); full content still stored, only FTS coverage capped
- zcode discover: cache by source-db mtime (idle cycles = 0 SQL instead of unindexed lower() table scan per worker per event)
- watch: wire the local anonymous stats counter into the daemon cycle (status panel's 'ignored 0' lied vs 34 real blocks)
… with N workers sharing one source db, projects B/C received project A's session list and ingested it as their own (cross-project contamination, live: music 42->559, novel 15->495). Key now includes project root.
…n (4th-project field report)

From a live field report (E:/mc整合包, 4th deployment, direct SQLite forensics):
- P0 decisions froze for 20+h on long-lived sessions: extraction only ran at
  confirm (idle 10min -> pending -> 6h cooldown -> confirm); continuously-worked
  sessions never completed the loop. Fix: refreshExtraction() now runs at every
  pending transition (10-min idle), decisions lag <= 10min. confirmSession reuses it.
- P1 88% of extracted decisions were workflow-subagent fragments: zcode adapter
  now tags title-matched subagent sessions with originHint='workflow' -> origin
  column -> listDecisions excludes them by default (data preserved, queries clean).
  Live: 36 sessions auto-tagged, visible decisions 82 -> 20, residue 0.
- P2 decision text hard-cut at 90 chars produced mid-sentence/trailing-bracket
  fragments: cut at nearest pause char + trim trailing closers + reject
  leading-closer fragments.

Tests: 222 -> 221 pass + 1 skipped, 0 failures.
…text (same family as annotate/save fix - confirm/refresh path was the missed third site)
…n captured twice displayed twice in briefs)
@rui08984-dot

Copy link
Copy Markdown
Author

Follow-up: 3 more commits since opening (live field report from a 4th deployment)

A 4th project (E:/mc整合包, MC modpack) was onboarded on 09-23, and its AI operator did direct SQLite forensics (node:sqlite, bypassing the CLI) — the report caught an architectural defect the earlier rounds missed:

P0 — decisions froze for 20+ hours on long-lived sessions. Extraction only ran at the confirm transition (idle 10min → pending → 6h cooldown → confirm); a continuously-worked session never completes that loop, so sessions.decisions froze at first extraction. Fix (the reporter's suggested 1a): refreshExtraction() now runs at every pending transition — decisions lag ≤ 10 minutes. confirmSession reuses it.

P1 — 88% of expanded decisions were workflow-subagent fragments. The zcode adapter now tags title-matched subagent sessions with originHint='workflow' → origin column → listDecisions excludes them by default (data preserved, queries clean). Live on the 4th project: 36 sessions auto-tagged by the new code, visible decisions 82 → 20, residue 0.

P2 — decision text hard-cut at 90 chars produced mid-sentence / trailing-bracket fragments: cut at the nearest pause char + trim trailing closers + reject leading-closer fragments.

Plus: applyExtraction re-includes existing user_tags in meta_text (third site of the meta-rewrite family — annotate/save were fixed earlier, the confirm/refresh path was missed); listUnresolved dedupes the same question captured across sessions (briefs showed it twice).

Commits: 3e7679e (P0-P2), f17f52d (user_tags), 3b69246 (dedupe). Suite still 222 → 221 pass + 1 skipped, 0 failures. All changes verified on the live 4-project deployment (heartbeats, capture latency, brief contents).

Your Name added 5 commits September 24, 2026 20:39
…ntence start; 最终用/最终选 stay strong (they are genuine final-choice verbs - caught by the suite)
…rmark from decisions column, O(delta) instead of O(full session)); topics/summary full rebuild stays at confirm - kills the last recurring CPU/RAM burst generator on long-lived sessions
…le/bold residue rejected); briefs read unresolved entries via q field (title-only display was a silent regression from the --json fix)
…de confirm counts were invisible in the status panel)
… and incremental decision refresh (seq-watermark idempotency); guards the mc-report root-cures against silent breakage
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant