Skip to content

Vendor-neutral conversation index: read Codex (then Gemini) sessions into cldctrl search #11

Description

@RyanSeanPhillips

Summary

Make cldctrl's conversation memory/search vendor-neutral by indexing other agents' sessions (starting with OpenAI Codex), so search_conversations — and the dashboard's recent/per-project views — span Claude and Codex (later Gemini). Because the search tool is exposed over the cldctrl MCP, this also gives the reverse: a Codex or Claude session searches one unified cross-vendor history.

This is the differentiation that first parties structurally won't build (Anthropic won't index Codex; OpenAI won't index Claude). Today cldctrl's "vendor-neutral" is only ~1 layer deep (terminals + filesystem project discovery); the brain (search/sessions/stats) reads only ~/.claude.

Codex data layout (verified on a real machine, June 2026)

  • Sessions: ~/.codex/sessions/YYYY/MM/DD/rollout-<ts>-<uuid>.jsonl (organized by DATE, not project).
  • Each rollout is JSONL of events. Useful lines:
    • {"type":"session_meta","payload":{"id","timestamp","cwd": "<PROJECT PATH>"}} → project association comes from cwd.
    • {"type":"event_msg","payload":{"type":"user_message","message":"…"}} and …"type":"agent_message","message":"…" → clean conversation text for indexing.
    • response_item (role/content) and function_call items → tool names + touched files (deeper extraction, optional).
  • ~/.codex/session_index.jsonl = {id, thread_name, updated_at} → human titles + a fast index (avoids parsing every rollout).
  • Distilled memory/state in SQLite (memories_*.sqlite, state_*.sqlite) — optional later source.
  • Gemini CLI would be analogous (~/.gemini, GEMINI.md) — not installed on this machine yet.

Design

  1. Adapter interface. Define a small SessionSource abstraction: listSessions(), readTranscript(id), projectOf(session), title(session), usage(session). Implement ClaudeSource (refactor current logic) + CodexSource.
  2. Index. Extend core/conversation-search.ts to ingest from all sources, each result tagged vendor. Cache by (mtime,size) per file as today; use session_index.jsonl to skip unchanged Codex rollouts cheaply.
  3. Surface. SearchResult gains vendor; dashboard shows a small per-vendor chip; resume/open routes to the right CLI (cockpit already vendor-aware).
  4. Reverse (mostly free). search_conversations is already an MCP tool registered in Codex's config.toml, so a cross-vendor index immediately makes unified search reachable from Codex sessions too.

Scope / phases

  1. CodexSource + search indexing (proof) — rollouts → unified search_conversations. Highest-leverage first slice.
  2. Per-project session list + recent-conversations include Codex (dashboard).
  3. Codex usage → Stats tab (needs mapping Codex token accounting; separate from Claude's usage fields).
  4. Gemini source when present.
  5. (Optional) read Codex's memories_*.sqlite as an extra recall source.

Non-goals / risks

  • Not trying to unify the agents' native memories into one store — just make cldctrl's index span them.
  • Codex rollout schema is owned by OpenAI and may change; keep the adapter defensive (tolerate unknown event types).

Relates to the strategic reframe (vendor-neutral brain = the moat) and #10 (cross-project coordination, which becomes far more useful when the brain spans vendors).

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions