You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #445, as you asked: Phase 2 as its own discussion once Phase 1 is up (#557).
I've been running an agent inbox on my fork since mid-September, with up to five sessions (claude, codex, opencode) coordinating through it daily. Your outline in #445 differs from what I built in three places that matter. Two of them came from failures I hit in practice, the third from how I deploy, so before writing the upstream version I'd like to settle them with you.
Where we agree
No file drop, part of the wire contract. Owner-scoped through findSessionOrFail. A per-session cap (mine is 200 messages, and a full inbox refuses the post instead of dropping the oldest). ?wait= bounded (mine has its own cap of 4 parked reads per session; folding that into the session-wait-registry accounting is fine with me). And from as the resolved caller instead of a client-supplied string: mine takes the label from the body or the X-Codeman-Parent-Session header, which is spoofable in multi-user mode. You're right there and I'd change it.
1. Drain on read
I started with that (a read acknowledged what it handed out) and lost orders with it. On 2026-09-23 an opencode planner posted a task to a codex worker and nudged it with "read your mailbox". The worker ran the read, summarised the message back, and ended its turn. The read had consumed the message, so the inbox was empty and ls showed nothing pending. To the planner, "read and dropped" looked exactly like "done", and the task only got done after a human prompted the worker a second time.
What I have now: a read marks messages as seen and hands them out, and they stay until the receiver calls ack. ls shows pending mail per session, so an order that was read and not handled stays visible. A TTL on top is fine.
2. No pane write
I agree the server shouldn't push into panes by default. The problem is the opposite case: an idle receiver never looks. In a map-editing run on 2026-09-27, a codex coordinator posted tasks to two idle workers and neither read them; afterwards their completion posts sat unread in the coordinator's inbox. Every hand-off needed a manual send.
My fork answers that with a one-line nudge typed into a receiver that would never look. Each message is nudged at most once, posts within 2 s share one nudge, there's a 30 s cooldown, nothing is typed while the receiver is parked on a read, and a busy pane gets it when it goes idle, or after 5 minutes regardless. It doesn't type over a half-written prompt, as far as Codeman can see one (text that came through its own input path without an Enter). It works, but it's a pane write, which is what you ruled out for v1.
A version without pane writes that would still have caught both cases: the server already knows, per session, whether a turn has ended and how much mail is pending. If the coordinator can see "worker X is idle with unread mail" right away, it can nudge X itself with the existing input path. That's an explicit act by an agent, not the server. On my fork that's codeman agent watch, which long-polls on a per-session "turn ended at" latch and reports a reason (inbox-unread, inbox-unacked, open-todos, blocked, …) with a cursor so no turn end is missed between calls. In one test each: for claude the watch returned 6 ms after the stop hook landed, for codex the idle heuristic fired about 5 s after the rollout's task_complete. If that direction interests you, it would be its own proposal after the inbox.
3. Memory only
I restart my server on every deploy (five times today). With an in-memory inbox, every restart would drop pending orders between agents that keep running in tmux across the restart. Mine writes a snapshot on change and restores it at boot, pruning inboxes of sessions that no longer exist. Making persistence opt-in (off by default) would keep your "nothing new on disk" for everyone else.
Proposal for the upstream version
Read marks seen, ack removes, per-session cap, TTL.
from resolved server-side.
No pane writes. Pending/unseen counts and the "parked on ?wait" state are exposed in the session list or a summary route, so a coordinating agent can decide.
Persistence opt-in.
Does that shape work for you, or do you still want drain-on-read for v1?
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Follow-up to #445, as you asked: Phase 2 as its own discussion once Phase 1 is up (#557).
I've been running an agent inbox on my fork since mid-September, with up to five sessions (claude, codex, opencode) coordinating through it daily. Your outline in #445 differs from what I built in three places that matter. Two of them came from failures I hit in practice, the third from how I deploy, so before writing the upstream version I'd like to settle them with you.
Where we agree
No file drop, part of the wire contract. Owner-scoped through
findSessionOrFail. A per-session cap (mine is 200 messages, and a full inbox refuses the post instead of dropping the oldest).?wait=bounded (mine has its own cap of 4 parked reads per session; folding that into thesession-wait-registryaccounting is fine with me). Andfromas the resolved caller instead of a client-supplied string: mine takes the label from the body or theX-Codeman-Parent-Sessionheader, which is spoofable in multi-user mode. You're right there and I'd change it.1. Drain on read
I started with that (a read acknowledged what it handed out) and lost orders with it. On 2026-09-23 an opencode planner posted a task to a codex worker and nudged it with "read your mailbox". The worker ran the read, summarised the message back, and ended its turn. The read had consumed the message, so the inbox was empty and
lsshowed nothing pending. To the planner, "read and dropped" looked exactly like "done", and the task only got done after a human prompted the worker a second time.What I have now: a read marks messages as seen and hands them out, and they stay until the receiver calls
ack.lsshows pending mail per session, so an order that was read and not handled stays visible. A TTL on top is fine.2. No pane write
I agree the server shouldn't push into panes by default. The problem is the opposite case: an idle receiver never looks. In a map-editing run on 2026-09-27, a codex coordinator posted tasks to two idle workers and neither read them; afterwards their completion posts sat unread in the coordinator's inbox. Every hand-off needed a manual
send.My fork answers that with a one-line nudge typed into a receiver that would never look. Each message is nudged at most once, posts within 2 s share one nudge, there's a 30 s cooldown, nothing is typed while the receiver is parked on a read, and a busy pane gets it when it goes idle, or after 5 minutes regardless. It doesn't type over a half-written prompt, as far as Codeman can see one (text that came through its own input path without an Enter). It works, but it's a pane write, which is what you ruled out for v1.
A version without pane writes that would still have caught both cases: the server already knows, per session, whether a turn has ended and how much mail is pending. If the coordinator can see "worker X is idle with unread mail" right away, it can nudge X itself with the existing
inputpath. That's an explicit act by an agent, not the server. On my fork that'scodeman agent watch, which long-polls on a per-session "turn ended at" latch and reports a reason (inbox-unread,inbox-unacked,open-todos,blocked, …) with a cursor so no turn end is missed between calls. In one test each: for claude the watch returned 6 ms after the stop hook landed, for codex the idle heuristic fired about 5 s after the rollout'stask_complete. If that direction interests you, it would be its own proposal after the inbox.3. Memory only
I restart my server on every deploy (five times today). With an in-memory inbox, every restart would drop pending orders between agents that keep running in tmux across the restart. Mine writes a snapshot on change and restores it at boot, pruning inboxes of sessions that no longer exist. Making persistence opt-in (off by default) would keep your "nothing new on disk" for everyone else.
Proposal for the upstream version
ackremoves, per-session cap, TTL.fromresolved server-side.?wait" state are exposed in the session list or a summary route, so a coordinating agent can decide.Does that shape work for you, or do you still want drain-on-read for v1?
All reactions