13 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Status
Built and working — v0.1 (core) + v0.2 (multi-tab) + v0.3 (Claude Code cockpit) shipped. src/ (TypeScript backend), public/ (esbuild frontend), and tests all exist; 212 tests pass. v0.3 work happens on the v0.3-cockpit branch. Key docs:
docs/TECH_DOC.md— design & rationale (the why). Protocol (§4), session model (§5.2), security (§7), acceptance (§8).docs/ARCHITECTURE.md— function-signature-level contracts, dependency rules, the invariant checklist (§8). On conflict, ARCHITECTURE wins on how; the M*/L* anchors are cross-validated fixes.docs/PROGRESS_LOG.md— the memory file (cross-session). Read it first; it's the running record of what shipped (v0.1 T1–T21, then v0.2/v0.3 feature entries). Orchestrator-owned.docs/PLAN.md(v0.1 task plan) and~/.claude/plans/shimmering-wondering-island.md(v0.3 cockpit plan)..claude/agents/— project subagents:module-builder,module-reviewer(for parallel dev).
v0.3 additions (cockpit): Claude Code hooks → live status per tab (POST /hook, npm run setup-hooks); remote approve/reject (held POST /hook/permission); tmux keepalive (USE_TMUX); plus FE: themes, search, dashboard, QR, PWA, new-tab-in-cwd. The server gained a hook side-channel but the terminal stream is still a byte-shuttle.
Language decision: TypeScript (.ts), not .js — ARCHITECTURE §0 records this divergence from TECH_DOC's original .js filenames. Wherever the two docs conflict, ARCHITECTURE wins on how (it was cross-validated and corrected); TECH_DOC wins on why/scope.
Development Workflow: Plan & Progress Log (MANDATORY)
Work proceeds against a phased plan and is tracked in a progress log that acts as cross-session memory. A new Claude instance must be able to read the log and know exactly where things stand. Follow these rules:
Required reading before starting ANY task (especially sub-agents — context is isolated):
docs/TECH_DOC.md— design rationale, scope, protocol (§4), session model (§5.2), security (§7), acceptance (§8). The why.docs/ARCHITECTURE.md— function-signature-level contracts, dependency rules, the invariant checklist (§8). The how. On conflict, ARCHITECTURE wins on how, TECH_DOC on why/scope.- The task's
ARCH:field (in PLAN.md) points to the exact sections to read for that task; also read anyM*/L*anchors named in its Steps. M1–M7/L1–L5are the 17 cross-validated fixes (M=medium, L=low) embedded as same-named anchors in ARCHITECTURE.md (some echoed in TECH_DOC.md). Seeing(M2)in a task → searchARCHITECTURE.mdfor "M2" to get the full rationale and correct approach. These are the easiest things to get wrong — read them before writing.
The plan (docs/PLAN.md):
- The build is split into waves W0–W5 by dependency, with fine-grained tasks
T1–T21designed for multi-agent parallelism. Same-wave tasks are independent and own disjoint files; cross-wave follows dependencies. See PLAN §0–§1. - Every task has a stable ID (e.g.
T6) and anOwns:file list. Never edit files outside your task'sOwns:— that's how parallel agents avoid clobbering each other.src/types.ts(T2) is the frozen shared-contract source; if a new type is needed, change it there (a coordination point), don't redeclare locally. - The plan is the source of intent; do not silently deviate. If reality forces a change, update
PLAN.mdAND record the deviation in the log.
The log (docs/PROGRESS_LOG.md) — cross-session memory. WHO writes it depends on mode (G1):
PROGRESS_LOG.mdis a shared file — it is not in any task'sOwns:. To avoid concurrent-write clobbering, only one writer touches it: the orchestrator (the main session that dispatches tasks).- If you are a dispatched subagent doing one task: do NOT edit
PROGRESS_LOG.md. Instead, end your final message with a ready-to-paste log entry (the "条目模板" block from the log) describing status, files/functions touched, how it was verified (test/command + result), decisions/deviations, blockers, next step. That entry IS your deliverable back to the orchestrator. - If you are the orchestrator (or working solo): read the log at the start of every session; after each subagent returns (or after each solo sub-task), append its entry immediately and update the "当前焦点 / Current Focus" block — not in a batch at the end. An untracked completed task is a memory loss.
- Be factual — log what actually happened (tests failing, steps skipped), never aspirational. Never pre-fill future tasks as done.
Ordering — solo: consult PLAN.md → do the task (TDD per ARCHITECTURE §6/§7, stay within Owns:) → verify → append to PROGRESS_LOG.md → next task. Commit boundaries map to tasks where practical.
Multi-agent orchestration (when running tasks in parallel)
This repo is structured for orchestrator-worker parallelism (the official Claude Code pattern). The main session is the orchestrator; it dispatches PLAN tasks to subagents, waits, then synthesizes. Rules:
- Dispatch by wave: only same-wave, file-disjoint tasks run concurrently. Keep the parallel batch to ~3–5 agents (diminishing returns / token cost beyond that). PLAN §3 is the dispatch schedule; PLAN §4 gives each task's model and isolation.
- Use the project agents in
.claude/agents/:module-builder(TDD-implements one task) andmodule-reviewer(read-only review). Don't hand parallel work to a generic agent — these encode tool limits, the required-reading rule, theOwns:boundary, and the log-return contract. - Subagents cannot talk to each other or ask the user mid-task (architectural limit). So: tasks must be self-contained; if a subagent hits an ambiguity not resolved by the docs, it must stop and return a
[!] BLOCKEDentry — never guess. The orchestrator resolves it (asks you if needed) and re-dispatches. - File isolation: file-ownership (
Owns:) prevents logical conflicts. For genuinely concurrent edits, also run builders withisolation: worktree(each gets a throwaway git worktree) so paralleltsc/vitest/edits never collide on the working tree. PLAN §4 marks which tasks want it. - Verification tasks (W5) are report-only by default: a reviewer/acceptance agent finds issues and returns them; fixes are routed back to the owning module's builder (or done by the orchestrator), so no agent edits files outside its lane.
What This Is
A browser-based terminal that exposes the host machine's local shell over WebSocket, so any device on the LAN can open http://<host-ip>:3000 and get an interactive shell. The primary use case is vibe coding — sending Claude Code a task, walking away, and reconnecting from any device (phone/tablet/another computer) to check on it. This drives two non-obvious requirements: sessions must survive disconnects and the mobile UI must be usable.
Explicitly out of scope (v0.1): auth/login, multi-user isolation, SSH jump hosts, containers, and any public-internet exposure.
Planned Commands
npm install # triggers node-pty native compilation (needs Xcode CLT on macOS)
npm start # listens on 0.0.0.0:3000
npm test # unit tests (vitest, all modules)
Config is via env vars only (no hardcoding): PORT, SHELL_PATH, BIND_HOST, IDLE_TTL, SCROLLBACK_BYTES, MAX_PAYLOAD_BYTES, USE_TMUX (1/0/auto), ALLOWED_ORIGINS. Note allowedOrigins is derived from the host's network-interface IPs (not from BIND_HOST — 0.0.0.0 is never a valid Origin); see ARCHITECTURE §3.1.
Architecture (the parts that span files)
The server is a byte-shuttle, not a terminal. It does not parse ANSI/terminal semantics — xterm.js (browser) interprets escape sequences and renders; node-pty (server) provides the pseudo-terminal so the shell believes it has a real TTY. This separation is the central simplification — keep it. Don't add terminal-semantic parsing on the server.
Data flow: keypress → xterm onData → WS → pty.write() → shell stdin, and shell stdout → pty onData → WS → xterm.write() → screen.
Session/connection decoupling (the most important design point)
PTY lifecycle ≠ WebSocket lifecycle. A WS disconnect must NOT kill the PTY — the Claude Code task running inside has to keep going. Sessions are keyed by sessionId (client stores it in localStorage):
attach(null)→ spawn a new PTY + shell; PTY output goes into a ~2MB ring buffer and (if a WS is attached) forwards live.attach(sessionId)→ look up the session, replay the ring buffer, then resume the live stream. This is what makes "refresh the page and the Claude session is still there" work.- WS close →
detachone client (PTY keeps running), not kill; the idle clock starts only when the last client leaves. - Multi-device sharing (v0.4): a session may have many concurrent WS clients — a new attach JOINS (mirror), it does not kick. Output/exit/status broadcast to all; any client can type (shared control). PTY sizing = latest-writer-wins: the device that most recently fit/focused drives the size, so whichever device you're actively using is full-screen (a shared PTY can only be one size; min-sizing letterboxed the bigger screen). Attach/detach/
blurnever resize — the active device keeps its size; the frontend re-sends dims on pane-show and window-focus so switching devices reclaims full-screen. This relaxes the original "one WS per session" invariant (#5).GET /live-sessionslists running sessions so any device auto-shows them as tabs;?join=<id>and the 🔗 share-QR open a specific shared session; the 🗂 manage page (/manage.html) is a full-page grid of live preview thumbnails — each card renders the session's current screen via a read-only xterm (fed byGET /live-sessions/:id/preview→RingBuffer.tail(), no attach) so you can see what each session is doing, with open/kill (DELETE /live-sessions[/:id]). - Orphan reclaim: detached longer than
IDLE_TTL(default 24h) with no new output since detach → reclaim. (node-pty can't reliably detect a foreground child, so liveness is approximated via last-output time — see ARCHITECTURE §3.5 / M3.) On server exit,pty.kill()all sessions (no cross-restart persistence in v0.1; tmux backend is a v0.2 idea).
Represent session state as immutable snapshots (id/start-time fixed at creation); hold mutable runtime handles (pty/ws) separately.
WebSocket protocol (protocol.ts)
Single WS connection, JSON text frames. attach must be the first message. Client→server types: attach / input / resize. Server→client types: attached / output / exit. See TECH_DOC §4 for exact shapes.
resize must be its own message type (it triggers ioctl(TIOCSWINSZ) → SIGWINCH so full-screen TUIs like vim/top redraw correctly). Validate at the boundary: type must be whitelisted, resize cols/rows must be integers 1–1000, input.data must be a string passed through verbatim (it's raw keyboard bytes — do not filter content).
Security (non-negotiable)
This app hands a full shell to anyone who can reach the port. Origin-header validation on the WS handshake is the one defense that cannot be skipped — it blocks Cross-Site WebSocket Hijacking (a malicious page in your browser trying to connect to ws://<lan-ip>:3000). Reject foreign origins (401). Never port-forward / tunnel this to the public internet. Tailscale is the recommended deployment over bare LAN.
Frontend (public/)
xterm.js + FitAddon. WS URL is same-origin, with the scheme following the page protocol (wss: on HTTPS, else ws:) so no IP config is needed and Tailscale/TLS (HTTPS) deploys avoid mixed-content blocking (M6). Auto-reconnect with exponential backoff (1s/2s/4s… cap 30s), carrying the localStorage sessionId. A mobile touch key-bar (hidden >768px) sends Esc \x1b, Shift+Tab \x1b[Z, arrows, Enter \r, Ctrl+C \x03, Tab \t directly via ws.send (bypassing xterm to avoid popping the soft keyboard) — these are Claude Code's high-frequency keys that phone keyboards can't produce.
Gotchas (from the spec)
- Enter sends
\r(0x0D), not\n— easy to get wrong when synthesizing input. fit()must run after the container has real dimensions; calling it whiledisplay:noneyields NaN.- node-pty needs
npm rebuildafter a major Node version bump. - Don't double-handle IME — xterm.js already manages composition events; don't add your own keydown listener for CJK input.