Repository navigation
Keep an idle Claude session's prompt cache warm (opt-in) #554
Replies: 2 comments
|
Hey @JDProfresh, yes! Love this one, and the numbers you brought made it easy to say yes. Short answers first:
One thing changed under us though, and it matters for step 2. Over the last few days I noticed Claude Code compacting my idle sessions by itself. So I dug in, and then had a Codeman worker session go much deeper: the CC changelog and docs, the shipped 2.1.294 to 2.1.296 binaries, and 30 days of my own transcripts (48,770 API requests). Here's what came out. Claude Code is quietly testing "idle compaction"It's not in the changelog and not in the docs. It's in the binary, behind a server flag (
What it did here: 13 idle compactions between Oct 8 ~17:00 and Oct 9 05:35 UTC, on builds 2.1.289 to 2.1.295. Every one landed 54.7 to 55.7 min after the last reply. The 6 times I came back, the first turn re-wrote 19k to 106k tokens instead of 440k to 884k. The other 7 were never touched again, so for those it was pure (small) cost. Then it stopped. After Oct 9 05:35, four eligible gaps (214k to 487k of context, same builds) went fully cold. The cached flag in There's also a native keep-alive in the binary, switched offSo "Claude Code has no built-in keep-alive" is true for users today, but there's one in the binary behind a flag that's currently off: The cold re-writes are real, and bigFrom the 30 days of transcripts:
One correction to the arithmetic: on Opus 5.5 and Sonnet 5.5 cache reads are 0.05x, not 0.1x (footnote on the pricing page). The 1h write is still 2x. So nudges are even cheaper than you wrote. Step 2 as written collides with idle compactionBoth clocks start at the same moment: My first idea was "just nudge earlier", but that isn't stable. What I'd suggest instead: Codeman owns the idle policy for sessions where the keep-alive is on.
Policy simulation over my real 30 days of gaps (relative to doing nothing)166 returns after 55+ min gaps, plus sessions that never came back (counted as an upper bound). API-equivalent prices, which are only a proxy for plan usage. How subscription limits weigh reads vs writes isn't documented. Nudge output assumed at 300 tokens.
Caveat: on Asks for step 1
What the release notes do document
Other things the research turned up on our side
So: go ahead with step 1 whenever you're ready and reference this thread. For step 2, I'd love the "Codeman owns the idle policy" shape (keep warm, then Thanks a lot for this, and for #550, which shipped in 1.40.0! |
|
Thanks for digging into this, and for the 30 days of data. The idle-compaction find changes step 2 in a good way, and I'm glad step 1 lands on its own. Step 1 is up as its own PR, #607. It's just the readout plus the schema asks you listed (keep the whole On step 2, yes, let's do the "Codeman owns the idle policy" shape. Concretely that means: I pulled apart the installed binary (2.1.294, same family you read) on the live-pickup question, since that's the one that decides how the toggle behaves. The answer is that it does not pick up live, and the reason is in the resolver. The design consequence: a session started with keep-warm on gets On the nudge cost you flagged as worth measuring: I've had step 2 running on my own install, so I have real nudges to read. Three of them fired on one Opus 5.5 session in normal mode (Oct 8 to 9), and every one output exactly 4 tokens ("ok") with 0 thinking tokens, replied in 2 to 4 seconds, read the full 243k to 333k of context from cache, and created only 56 to 434 new cache tokens. So observed is 4 output tokens, not the 300 your sim assumed, which makes the keep-warm rows cheaper than the table shows. The one thing I can't close from my own data: all three were Opus 5.5 where thinking came out at 0, and I don't have a nudge captured on xhigh with thinking forced on, so I can't fully rule out your ~3k worst case there. A one-word directive produced zero thinking even on Opus 5.5, which makes me doubt it, but I'd keep 300 as the conservative default until I measure an xhigh case. For what it's worth, every binary number in your write-up held on 2.1.294: Two things where your answer changes what I build:
|
Uh oh!
There was an error while loading. Please reload this page.
When I step away from a Claude session for more than an hour, my next prompt re-writes the whole conversation into the cache. On a subscription the main conversation cache lives for one hour after the last request, and a cold re-write is charged at twice the input rate, while a warm read is a tenth. I measured my own last 14 days of transcripts (82 sessions): 255 gaps of 5 to 60 minutes all stayed warm, but 25 of the 27 gaps over an hour went cold and re-wrote 3.4M tokens between them. Six of those were overnight, which I don't think is worth chasing. The rest were the 1 to 4 hour gaps of a normal day: lunch, a meeting, working in another session.
Claude Code has no built-in keep-alive (there are feature requests for one in its tracker). But since 2.1.251 its statusline JSON reports
prompt_cache(warm,ttl,expires_at,recache_tokens_if_cold,misses,last_miss_cause), and Codeman's statusline exporter already posts that JSON to/api/status-telemetry. So Codeman knows exactly when each session's cache will expire.Proposal, in two steps
1. Show the cache state.
StatusTelemetrySchemakeeps theprompt_cacheobject instead of stripping it.Session.promptCacheridestoState(), with a change-onlysession:promptCacheevent.cache:until HH:MM/cache:coldin Codeman's footer.Display only, no behaviour change.
2. An opt-in keep-alive.
expires_at, Codeman sends one short turn throughwriteViaMux:Codeman cache keep-alive: no action needed, reply with just "ok".That read resets the cache timer.The scope decisions, all deliberate:
The arithmetic: a nudge re-reads the context at a tenth of the input rate, while a miss re-writes it at twice the rate. So a capped run of up to three nudges in a gap costs a fraction of the one miss it prevents. On a subscription both come out of plan usage rather than dollars, but the ratio is the same.
Does it work
I ran it end to end on an isolated instance with Claude Code 2.1.280:
Both steps are built and tested, and I've been running them on my own install since 2026-10-07. The toggle and readout live on the Respawn tab, which is why I hit the bug in #549 (fix in #550).
Questions
If it's a yes, I'll send step 1 as a PR referencing this thread, then step 2.
All reactions