Files
claude-cloak/CLAUDE.md
2026-08-26 11:17:43 +02:00

34 KiB
Raw Blame History

claude-cloak

TUI that displays Claude Code's API streams token-by-token (thinking, text, tool calls) by acting as a pass-through proxy: Claude Code points ANTHROPIC_BASE_URL at 127.0.0.1:8484, we forward everything verbatim to api.anthropic.com and tee SSE responses into the UI. Never issue API requests of our own — zero extra usage is the core constraint of this project.

Architecture

src/main.rs    entry; tokio runtime for proxy task, TUI on main thread; --headless mode
src/proxy.rs   axum fallback handler: buffers request body (for session metadata,
               tool results, and user prompts — `app::record_user_prompt` lifts
               the trailing user message into a Kind::User feed entry verbatim
               (incl. slash-command machinery — the goal is to show everything
               the model received, never filter it). On a turn-starting request
               (tools present) it also emits the system-prompt *size* as a
               Kind::System line (the prompt itself is too long to show) and the
               available tool set as Kind::ToolDefs — each once, re-emitted only
               on change (system/tools/history are re-sent every request but are
               not new data). A *side* request (no tools — topic/title haiku
               calls) is still shown, tagged with a `── side request ──` Meta
               divider. `app::extract_user_text` splits a user text block into
               its injected `<system-reminder>` spans (kept as dimmed
               Kind::Reminder entries, never discarded — the prompt survives even
               when it shares its block with a reminder, the
               first-message-after-resume case) and the real prompt;
               `strip_injected` is the label-only projection (drops reminders
               *and* slash-command machinery) used for turn-tree labels.
               Dedup drops only true resends (the just-recorded prompt is still
               the tail entry), so verbatim repeats in later turns survive.
               Also reads the `x-claude-cloak-pane` header (passed to `Tap::new`
               to bind the embedded pane — see the embed-identity invariant) and
               strips it before forwarding. Forwards via reqwest, streams the
               response back unbuffered, tees SSE
src/sse.rs     incremental SSE parser; tolerant of chunk splits mid-event/mid-UTF-8
src/app.rs     Arc<Mutex<App>> shared state; Tap = one in-flight tapped request,
               translates SSE events → session Entries (Drop closes it out).
               Each Tap belongs to a *lane* (`Lane`/`LaneId`): lane 0 is the
               main chain, every subagent gets its own. Entries stay in one
               append-only Vec tagged with `Entry::lane`; per-agent state
               (model, tokens, tool count, system/tools signatures, label,
               parent, finished) lives on `Lane`
src/ui.rs      ratatui rendering @ ~30fps; session list + scrollable feed
               (FeedCache: per-entry rendered lines + wrapped heights, only
               changed entries re-render; the viewport window of lines is
               handed to ratatui so scroll state is usize end-to-end).
               Focus accent is orange (`ACCENT` = indexed 208): borders, the
               scroll thumb and the user-prompt blocks all use it when focused,
               dim grey when not. User prompts render as full-width filled
               rectangles padded to exactly the inner width (`wrap_words` +
               exact pad, never the Paragraph's own wrap, so the box ends flush
               with the borders); the fingerprint folds feed-focus in for
               `Kind::User` only, so a focus toggle re-renders just those
               blocks. `color_on(bg)` picks black/white text by background
               luminance (used by the prompt blocks and the edit/diff blocks) so
               filled blocks stay legible under any terminal theme. Injected
               `<system-reminder>`s and the system-prompt-size `Kind::System`
               line show dim under the "system" filter; the `Kind::ToolDefs`
               tool-list line shares the "tools" filter with tool calls. The
               feed's right border doubles as a prompt
               minimap: `*` markers show where each user message sits in the
               whole conversation, with the scroll thumb drawn on top where they
               coincide.
               Subagents never touch this feed: it renders lane
               `MAIN_LANE` only, at full width, whatever the agents are doing.
               They live in the `A` popup (`popup_rect` = 80% of the *feed*
               rect, centred): `draw_agent_list` is the picker,
               `draw_feed` the chosen agent's own stream — same function as the
               main feed, own FeedCache from the `FeedCaches` pool, own
               scroll/follow from `App::lane_cols`, so it follows its own tail
               and the border title carries the identity (`⟳ Explore · find the
               retry helper · sonnet · out 2.1k · 2/3`). See the
               subagent-popup invariant.
               Sessions panel is a uniform 50% of the main area: each session
               is a multi-line item — full white title (live = first user
               prompt via `live_title`, stub = disk label, word-wrapped by
               `wrap_words`) over a dimmed id·model meta row; expanded turn
               rows are indented past the title and `truncate_str`'d to one
               line each.
src/markdown.rs wraps tui-markdown: renders GFM tables itself (box-drawing,
               width-fitted wrapped columns) and strips heading `#` markers —
               the pinned tui-markdown 0.3.5 does neither
src/sessions.rs on-disk session history (main chain *and* subagents):
               background scanner thread keeps
               App::disk_sessions fresh (~1/s poll, `read_meta` re-read only on
               mtime change — one pass yields the label *and* the session's
               last main-chain model, which `App::resume_model` turns into the
               `--model` a resume spawns with); load_view/load_history rebuild a feed Session
               from a JSONL transcript (lazily, on first view); build_tree
               parses uuid/parentUuid chains into a TurnTree (one node per
               real user prompt; rewinds leave fork points); materialize
               writes a new session file from a chosen set of turns.
               `scan_agents` reads the `<session>/subagents/agent-*.meta.json`
               sidecars (cheap: the transcripts themselves can be MBs) and
               `splice_agents` inserts each agent's entries into its own lane
               right after the `Agent` tool call that spawned it
src/term.rs    embedded claude pane: spawns `claude --session-id <uuid>` in a
               portable-pty routed through the proxy; wezterm-term models the
               screen (and answers terminal queries); renderer paints cells
               into the ratatui buffer. Each spawn injects a fresh per-pane
               token via `ANTHROPIC_CUSTOM_HEADERS` (`PANE_TOKEN_HEADER` =
               `x-claude-cloak-pane`), the correlation handle the proxy uses to
               recognise the pane's own traffic (see the embed-identity invariant).
               `cc_default_model` reads Claude Code's *own* configured default
               model out of its settings — the only persisted record of a
               `[1m]` pick (see the 1M-context invariant)

Data flow: proxy task parses SSE chunks → Tap::handle() mutates shared state → UI thread redraws on its own tick (no channel; just the mutex).

Key invariants

  • Latency-neutral pass-through: response bytes are forwarded as-is, never buffered or rewritten. Auth headers pass through untouched. If the tap code panics or misparses, the proxy must still relay bytes (tee is best-effort): chunks are try_send-cloned into a bounded channel and parsed on a separate task (dropped on overflow, never blocking the relay), and the proxy/tap side locks the app mutex poison-tolerantly (app::lock_app) so a UI panic can't kill forwarding.
  • accept-encoding is stripped from forwarded requests so the upstream sends identity encoding we can parse in transit. Don't "fix" that.
  • Hop-by-hop headers (content-length, transfer-encoding, etc.) are stripped both directions; hyper re-frames.
  • Sessions are keyed by the session UUID in request metadata.user_id — Claude Code ≥2.1.x sends a JSON blob with "session_id":"<uuid>", older builds user_…_session_<uuid>; proxy::session_key handles both. Concurrent requests (subagents) share a session but each Tap tracks its own current entry index — entries/sessions are append-only, so indices stay stable. session_key tolerates whitespace around the JSON colon (a pretty-printed blob used to fall through to the legacy session_ split and yield id": "…).
  • Subagent identity comes from Claude Code's own header, never a heuristic. A subagent's request reports the parent's session_id and no agent id in metadata, but Claude Code stamps x-claude-code-agent-id (and, from spawn depth 2, x-claude-code-parent-agent-id) on every one of them. That id is unique per agent — including byte-identical sibling prompts, which do occur and which a prompt hash cannot separate — stable across the agent's inner-loop turns, and equal to the agentId of its on-disk subagents/agent-<id>.jsonl. proxy.rs reads both headers (AGENT_ID_HEADER / PARENT_AGENT_ID_HEADER) and forwards them untouched — they are Claude Code's, not ours; only x-claude-cloak-pane is ours to consume. Session::lane_for maps the id to a lane (appended on first sight, so a LaneId stays valid forever). Labels come from a separate, later fact: Session::label_lane_from_prompt matches the subagent's opening prompt against an unclaimed Agent tool call's prompt (byte-identical on the wire) to learn subagent_type/description/ parent, and close_lane_from_result scrapes agentId: <hex> out of the Agent tool_result to tie the lane to that call and mark it finished. A lane must never wait for either: a synchronous agent's result only lands when it has already finished, and the child's first request can beat the parent's next one, so lanes are born anonymous and adopted later.
  • One entry vec, tagged with lanes. Per-lane vecs would double every index site (Tap::cur), break FeedCache's positional alignment with Session::entries, turn the viewport window into a k-way merge inside the mutex the tap shares, and lose global wire order. Lane membership is a field; showing one lane is a filter (e.lane == args.lane), which is also why there is no "show everything interleaved" mode.
  • "The agent finished" is a <task-notification>, not the tool_result. Claude Code launches every Agent call asynchronously: the tool_result comes back immediately and says so (Async agent launched successfully… \ agentId: <hex>), and the real completion is injected into the parent's next user turn as <task-notification><task-id><agent id></task-id>. So close_lane_from_result only ties the lane to its tool call (and finishes it in the non-async wording, kept for older builds), while Session::finish_lanes_from_notifications — called from record_user_prompt on the trailing user run, before its early returns — is what stamps Lane::finished_at. Background bash tasks share the notification shape with a short id that matches no lane. Reading finished off the tool_result alone is why a finished agent used to keep reading as running.
  • Subagents live in a popup; they never share the feed. The main feed always renders MAIN_LANE at full width, so how many agents run changes nothing about reading the main chain — no split, no rows, no reserved space, no interleaved entries (draw_feed filters e.lane == args.lane). A (App::toggle_agent_popupApp::agent_popup) opens the one place they are shown: AgentPopup::List picks an agent, AgentPopup::Feed gives one agent the whole popup (80% of the feed rect, ui::popup_rect). Opening takes the shortest path — a lone agent goes straight to its stream, several land on the picker with the first running one preselected — and A closes whatever is open. The popup is modal: while it is up it takes every key (and the wheel), which is why it needs no focus/column model at all. Esc unwinds one layer (feed → picker → closed), [/] step between agents from inside a feed. App::agent_list_of orders it: running first (Lane::running), then idle/finished, each group in spawn order — but every lane is listed, disk lanes included, because this popup is the only way to read a finished agent's output. State is session-local: draw clears agent_popup and lane_cols when the displayed session changes, and validate_agent_popup drops a popup whose lane the displayed session doesn't have (a rebuilt on-disk view), so the render path never sees a dangling LaneId.
  • Lane::running is a sort key, never a gate: streaming (active > 0), or no finish signal and quiet for less than LANE_IDLE_MAX (60s); a finished_at (the <task-notification>) or no traffic at all (a lane read from disk) means not running. The long idle net matters because a gap between an agent's turns (a slow local tool call) looks exactly like "done"; only the notification distinguishes them. Being wrong therefore costs an ordering and a /· mark — never a hidden stream, which is what the old row-collapse timers could do.
  • Only the main lane drives the pane and the session header. embed_grow, the ctrl-l wipe scheduled in Tap::drop, the prompt minimap, n/N and Session::model/context are gated on MAIN_LANE; last_system_len and last_tools_sig live per lane (a subagent's system prompt and restricted tool set differ, so session-wide state re-emitted both lines on every main↔subagent alternation), and the prompt dedup is scoped to the lane.
  • Embed identity is learned from traffic, never assumed from --session-id. Claude Code's interactive --session-id is not guaranteed to equal the id it reports in request metadata (and a --resume can mint a fresh one), so the pane is correlated by a token we control: term.rs injects a per-spawn x-claude-cloak-pane header (ANTHROPIC_CUSTOM_HEADERS), the proxy reads it (and strips it before forwarding), and Tap::new binds App::embed_session to whatever id that tagged request actually carries (App::bind_embed_session rebinds + renames a provisional resume row if they differ). Selection policy follows: the embed jumps the selection only on first bind; a brand-new external session auto-jumps so a fresh /clear is visible unless App::pane_focused (mirrored from the UI each frame) — never steal the selection from a pane the user is driving. This is what made an a-spawned session stream into the wrong row before.
  • One app instance = one proxy port = at most one embedded claude (EmbedUi::term / App::embed_token → learned App::embed_session). kill_current_embed is the single teardown path and bind_new_pane the single registration path, so pane identity + grow/clear flags can't drift across the spawn/replace call sites. Every other live session is an external claude pointed at our port: observable, never attachable. The pane stays visible while it holds keyboard focus even if the selection isn't on its session yet (its id is still being learned); only an intentional ctrl-↑ / tab-away hides it.
  • The session list merges live sessions (first, indices stable) with this directory's past sessions from ~/.claude/projects/<cwd with / → ->/*.jsonl as dimmed stubs (deduped by uuid — a live session's file is on disk too). Tab is viewing only, never a process operation: selecting a stub lazy-loads its transcript into App::history; tabbing off the embedded session hides the pane without killing the child (instant to come back). ctrl-↓ is the commit point that attaches the pane to the selection: reveal+focus if it's the embedded session, claude --resume <uuid> (kill + respawn) for disk stubs and dead embeds, fresh --session-id spawn when there's nothing. Live external sessions are guarded — their instance may still run elsewhere and a second --resume would fork the transcript — but a second ctrl-↓ within 3s forces it (liveness is unknowable: an idle claude sends no traffic; EmbedUi::past_embeds skips the guard for sessions whose instance we killed ourselves). --session-id cannot be combined with --resume (CLI rejects it without --fork-session); --model can, and every resume passes it.
  • A resume continues on the session's own model, not the CLI default: App::resume_model reads the model Claude Code recorded for the session's last main-chain assistant message (DiskSession::model, filled by the scanner's read_meta — subagent isSidechain records run their own model and <synthetic> error records carry no model, so both are skipped) and app::model_arg_for_id maps that id to a --model argument: a known alias (sonnet, opus, … from App::model_choices) wins over the dated snapshot id, so a retired snapshot can't pin the pane; an id with no alias inside is passed through verbatim (--model takes full names too). The transcript is authoritative, so a mid-session /model switch is honoured.
  • The 1M context window is a header, and a resume must keep it. --model opus[1m] differs from opus only by anthropic-beta: …,context-1m-… — same body model, same transcript record — so no amount of transcript reading can tell them apart. The proxy is the only place that sees it: proxy.rs reads BETA_HEADER on main-chain turn requests only (a side/title call runs haiku without the flag, a subagent runs its own model) and app::record_long_context stores it as Session::long_context. App::resume_arg then picks the window: the wire observation wins, else the [1m] in Session::spawn_model (our own spawn, while it still names the same model), and with neither — a session that predates this process — it falls back to term::cc_default_model(), Claude Code's configured default (ANTHROPIC_MODEL, then local/project/user settings.json), which is the one place a [1m] pick is persisted (/model writes it there). When that default names the same base model the resume passes no --model at all and inherits it whole, window included; any explicit knowledge overrides it, including "this session ran the short window", which is why an observed non-1m session is resumed with an explicit --model opus. The suffix is only ever added for an alias that model_choices says has a [1m] variant.
  • Turn tree / branching (lazygit/yazi-style, all in the sessions panel): space (or /l) expands the selected session's turn tree — one row per real user prompt, abandoned rewind branches indented under their fork point, trunk continuing below. j/k/↑/↓ walk sessions and turns (they never scroll the feed; the wheel and PgUp/PgDn/g/G do that). Highlighting a turn switches the feed to the on-disk transcript along the path through that turn and pins the turn's prompt to the viewport top (HistoryView caches per uuid, rebuilt when the leaf changes; FeedCache keys on leaf+live so views of the same uuid don't share slots). v anchors a contiguous visual range, b materializes a new fully decoupled session file — chain root→turn, or exactly the visual range stitched together (sessionId rewritten, each turn's head re-parented onto the previous turn's tail, our own ai-title record gives it the ⑂ … label) — injected as the selected top stub. Branching never touches a process: ctrl-↓ stays the only spawn/kill commit point.
  • The tap drives pane behavior: grows it for AskUserQuestion / ExitPlanMode (sized from the question's option count) before Claude Code renders the prompt, shrinks when the tool_result echoes back, and schedules a ctrl-l transcript wipe 400ms after each turn (the pane is prompt-only; the feed shows the context).
  • Tool input streams as raw JSON fragments; pretty-printed only on content_block_stop. Streaming text re-renders markdown on every change (FeedCache fingerprints by content length + done + result), so partial markdown self-heals; completed entries render from cache.

Gotchas

  • tui-markdown is pinned =0.3.5: 0.3.7+ moved to ratatui-core (0.30 alpha types), incompatible with ratatui 0.29 — 0.3.6 is the last 0.29-compatible release, but no version (through 0.3.8) enables pulldown-cmark's table extension, so upgrading still wouldn't render tables. Its gaps (no tables, literal heading markers) are compensated in src/markdown.rs, not by upgrading.
  • wezterm-term/wezterm-surface are not on crates.io: pinned to a git rev of the wezterm monorepo (keep both revs identical).
  • The compact pane dynamically frames Claude Code's input box rather than cropping by fixed offsets (term.rs: compact_frame + PaneView). It locates the box by its two horizontal-rule borders (text_is_rule) — the last two ──── rules on screen, since the prompt always sits at the bottom — and shows one context row above the top rule (the spinner / "✻ Worked…" row) down to the statusLine just under the bottom rule, cropping the persistent hint/token/effort chrome below it. When an @// menu is open it has replaced that chrome with a list (text_is_menu_item, CC-2.1.x glyphs — retune there if an update changes them), so the frame extends to the last non-blank row instead. The menu is detected by scanning the whole region below the bottom rule for a menu row, not just the row directly under it: the list can start after a blank/header row and only the highlighted item carries a glyph (unselected file rows are plain names), so checking one row collapsed the pane whenever that row wasn't the selected item. Above the top rule the frame also swallows an active task panel (text_is_task_row / task_block_top): Claude Code parks the N tasks (…) header + ✔ ◼ ◻ rows (and its … +N pending overflow line) directly above the input box, so walking up over that block — tolerating one blank line and single wrapped / activity rows, capped at MAX_TASK_BLOCK — makes task status visible with no extra app state. Priority when the pane can't hold everything: the panel is dropped first (CompactFrame::ess_top, the one-context-row frame) so the line you're typing and an open menu never fall off screen. The framed region drives the pane height too: compact_rows (called from ui::draw) measures box-height + tail so the pane auto-expands as the prompt gains lines or a menu opens and shrinks back when idle (floor MIN_COMPACT_INNER, cap = screen 6); PTY_PAD keeps the PTY taller than the visible window so the child can still draw the rows we crop. A cropped PTY is sized to the whole screen height, not the visible pane (ui::draw passes f.area().height to resize for every view except Full): their height is derived by measuring what Ink has already drawn, and Ink only ever draws as many rows as the PTY reports, so tying the PTY to the (small) visible height is a feedback loop — an @// menu or a big paste that suddenly needs many more rows than the current PTY+pad never gets the room to draw them, so compact_rows can't measure the growth and the pane stays stuck small. A screen-tall PTY lets Ink lay out the full box+menu in one shot; compact_view_range still shows only the cropped window. When that window is shorter than an open menu it top-anchors on the input box (crop the menu's tail, never the line you're typing) — the idle no-menu case still bottom-anchors on the statusLine. That per-frame measurement is smoothed by hysteresis (EmbedUi::compact_height / smooth_compact, seeded at DEFAULT_COMPACT_INNER): the pane grows instantly but shrinks only after the smaller height has held for SHRINK_DELAY (400ms), and compact_rows returns None on a transient mid-repaint (box border caught missing) so the last height is kept. Without this the height oscillates every frame during a subagent turn or @// menu filtering, and each change resizes the PTY → Ink repaints → flicker. PaneView::Interactive (the tap-grown AskUserQuestion / ExitPlanMode pane, whose selection box renders above the input) is measured the same way, never estimated: interactive_frame anchors on the rule above the header-chip row (← ☐ Header ✔ Submit →), else the second-to-last rule, and runs to the last non-blank row; EmbeddedTerm::interactive_rows feeds that height through the same hysteresis. App::ask_question_rows (the row guess from the tool JSON) is only the fallback for the frames before Ink has drawn the box — it can't know how far the question text wraps, which is what used to crop the first paragraph. Because the Interactive PTY is now screen-tall, Claude Code lays the prompt out in full instead of switching to its own truncated form. interactive_view_range top-anchors on that frame and slides down only far enough to keep the option on screen when the prompt overflows the pane. PaneView::Full (fullscreen) renders the child's screen verbatim from row 0 with the PTY sized exactly to the pane. Permission-prompt boxes (rounded borders, not rules, and not in the API stream) aren't expanded in the compact pane — consistent with the known "permission prompts aren't detected" limit.
  • Keybindings avoid Alt entirely: on layouts like dk_mac_fixed, Alt composes characters (alt-c = ©) and never reaches the app as a modifier. Pane keys: F2 toggle, ctrl-↓ attach pane to selected session (resume/spawn/focus), ctrl-↑ focus feed, ctrl-f fullscreen toggle (only while the pane is focused), ctrl-q quit (global; needed while the pane is focused, where plain q is forwarded to the child), c attach most-recent past session. List keys: j/k/↑/↓ move the session/turn highlight, space/→/← expand/ enter/leave the turn tree, a opens the model picker popup and spawns a brand-new claude --session-id … [--model …] (kills any current pane — show_embed_new; saves resume-then-/clear to get a fresh chat). The picker list comes from App::model_choices: seeded with default_model_choices, then replaced by term::spawn_model_discovery — a background scan that reads the live model-alias array (["sonnet","opus","haiku","fable",…]) straight out of the installed claude ELF (single self-contained binary with the JS bundle embedded). No API call, never runs claude — just resolves claude on PATH and greps its bytes for the longest lowercase-token array anchored by opus+sonnet. The list then gets the 1M-context variants appended (sonnet[1m] etc., Claude Code's --model spelling for the long context window): term::long_context_tokens collects every quoted "<token>[1m]" literal in the same byte scan, and only aliases that really have one are offered (today opus/sonnet/fablenot haiku or mythos), so the suffix is never assumed. [1m] needs no shell quoting: the pane spawns via CommandBuilder argv, not a shell. Tab/BackTab cycle sessions (p no longer mirrors BackTab). v visual range, b branch, Esc unwinds (visual → tree → quit). n/N jump the feed scroll to the next/previous user prompt (App::prompt_jump, applied in draw where entry heights are cached). The feed scrolls only via wheel / PgUp / PgDn / g / G / n / N. A toggles the subagent popup (the footer leads with A agents (N) when the displayed session has any — it is the only route to them). Inside it: j/k move the picker or scroll the agent feed one line, enter/→ opens the highlighted agent, [/] step to the previous/next agent, PgUp/PgDn/g/G scroll, Esc goes feed → picker → closed, A/q closes outright. Being modal it also owns the wheel (ui::wheel), so no pointer hit-testing is involved. Switching the displayed session clears App::lane_cols and closes the popup, so a lane id can't inherit another session's scroll offset. CT_DEBUG_KEYS=1 shows raw key events in the status bar.
  • Mouse is captured: wheel always scrolls the feed (regardless of focus), and left-drag selects screen text, copied on release via OSC 52 (like Claude Code). Native terminal selection therefore needs shift held.
  • Bracketed paste is enabled on the outer terminal (EnableBracketedPaste): a multiline paste arrives as one Event::Paste and, when the claude pane has focus, is handed to the child via EmbeddedTerm::paste (wezterm-term's send_paste re-wraps it in bracketed markers iff the child enabled them) — so Claude Code inserts it as one block instead of submitting on the first embedded newline. Paste is ignored when the pane is unfocused (nothing else takes text input).
  • Pane cursor shape mirrors the child: each frame draw records the child's DECSCUSR shape (EmbeddedTerm::cursor_shape) into EmbedUi::cursor_shape and the event loop emits SetCursorStyle only on change (so a blinking cursor isn't reset every frame), resetting to DefaultUserShape when no pane cursor is shown / on teardown. Without this the outer terminal kept a stale block cursor regardless of Claude Code's insert-vs-vim-normal state. term::cursor_style maps the child's DECSCUSR Default to a blinking bar, not DefaultUserShape: Claude Code's normal input leaves the cursor at the terminal default expecting a bar caret, so forwarding the outer terminal's own default (often a block) would wrongly show a block in insert mode; vim normal mode still sends an explicit SteadyBlock.
  • ratatui needs feature unstable-rendered-line-info for Paragraph::line_count (used to compute cached per-entry wrapped heights for follow/auto-scroll).
  • reqwest is default-features = false + rustls-tls,stream — don't enable compression features (would re-add accept-encoding).
  • The listener is bound in main before the TUI starts: prefers 8484, falls back to an OS-assigned free port so multiple instances coexist (each pane gets the actual port via ANTHROPIC_BASE_URL). CT_PORT pins the port and turns bind failure into a hard startup error.
  • Testing: SSE parser has unit tests (cargo test). For a live pass-through check: --headless (prints the bound port), then POST to 127.0.0.1:<port>/v1/messages without auth — a relayed 401 from Anthropic proves the round-trip. The TUI can't run in a non-tty. CT_UPSTREAM points the proxy at an alternative upstream (e.g. a local fake SSE server) for fully offline end-to-end tests with zero API usage. That fake server is dev/fake_upstream.py: it answers every request with canned SSE, so a real claude child can be made to render its client-side tool UIs on demand (dev/.fake_scenario = ask | plan | todo | taskupdate | agent | text, switchable mid-run) — this is how the pane's frame detector is developed against what Ink actually draws. Drive it through tmux (.claude/skills/tui-verify) and obey that skill's safety rule: never pkill/killall, tear down only your own named tmux session. The child writes real task files under ~/.claude/tasks/<its-session-id>/; delete that directory afterwards.

Not yet handled (known MVP limits)

  • Non-streaming requests pass through untapped (e.g. count_tokens).
  • Subagent popup: no per-lane prompt minimap (a subagent has no user prompts), and lanes loaded from disk have no token counts (a transcript records no usage) — the picker shows tool counts too. Only one agent is readable at a time (a modal popup, by design: the alternative was the split feed this replaced). A lane is never closed, only marked finished: a background agent (x-app: cli-bg) can wake up again long after its launch result landed, and SendMessage can revive a finished one. An agent transcript over MAX_AGENT_BYTES (8 MB) is summarised instead of parsed, because the view is built while the app mutex is held.
  • Materialized branch files carry no subagent transcripts: the Agent tool_results in them still hold the reports the parent model saw, and copying subagents/ would duplicate agentIds across two sessions and contradict the agent files' own sessionId. Deliberate — don't "fix" it by copying.
  • A subagent's first turn is what labels its lane, so if we attach mid-run (the parent's Agent call never passed through us) it stays listed as agent <id-prefix> until its result lands.
  • Sessions are never pruned (entry memory grows for the process lifetime); the same goes for viewed disk transcripts (App::history).
  • Request bodies are fully buffered (up to 512 MB) before forwarding — needed to read session metadata; adds first-byte latency on huge bodies.
  • Tool results only appear once the next request fires; if the session ends right after a tool call, that result is never seen. Output is what Claude Code sends the model (i.e. post-truncation).
  • Embedded pane: no mouse forwarding yet; no scrollback view (live screen only); shift+enter needs kitty keyboard protocol pushed on the outer terminal (not done); permission prompts aren't detected for pane growth (not visible in the API stream — would need a Notification hook hitting a local control endpoint).
  • Materialized branch files satisfy our own parser (round-trip tested) but Claude Code's loader tolerance is only verified empirically by resuming one — if a CC update changes the JSONL schema, retest b + ctrl-↓. The tree itself isn't refreshed while expanded (collapse/re-expand re-reads the file), and a highlighted turn of a live session views its on-disk transcript, which lags the in-memory feed by however much CC buffers.