RuleStack

Configs

Stacks

Compare

Diff

RuleStack

Configs

Stacks

Compare

Diff

Read API

RuleStack

Configs

Stacks

Compare

Diff

Read API

Diff/garrytan-gstack-agents ↔ garrytan-gstack-claude

Comparison

A · AGENTS.md · garrytan/gstackB · CLAUDE.md · garrytan/gstack
What each file covers, counted
DimensionSharedOnly in AOnly in BOverlap
Sections012350%
Commands523412%
Section tags80657%

What each file covers

Sections

0 shared · 12 only in A · 35 only in B
  • − gstack — AI Engineering Workflow
  • − Available skills
  • − Plan-mode reviews
  • − Implementation + review
  • − Release + deploy
  • − Operational + memory
  • − Browser + agent integration
  • − iOS QA — drive real iPhones over USB or Tailscale (v1.43.0.0+)
  • − Safety + scoping
  • − Build commands
  • − Platform support
  • − Key conventions
  • + gstack development
  • + Commands
  • + Testing
  • + Project structure
  • + SKILL.md workflow
  • + Platform-agnostic design
  • + Writing SKILL templates
  • + Writing style (V1)
  • + Browser interaction
  • + Dev symlink awareness
  • + Compiled binaries — NEVER commit browse/dist/ or design/dist/
  • + Redaction guard (PII / secrets / legal content)
  • + Commit style
  • + Slop-scan: AI code quality, not AI code hiding
  • + What to fix (genuine quality improvements)
  • + What NOT to fix (linter gaming, not quality)
  • + Utilities in `browse/src/error-handling.ts`
  • + Score tracking
  • + Community PR guardrails
  • + Checking out PRs from garrytan-agents
  • + CHANGELOG + VERSION style
  • + Release-summary format (every `## [X.Y.Z]` entry)
  • + Itemized changes (below the release summary)
  • + AI effort compression
  • + Search before building
  • + Local plans
  • + E2E eval failure blame protocol
  • + Long-running tasks: don't give up
  • + Running evals as an agent: always detach (SIGTERM-proof)
  • + E2E test fixtures: extract, don't copy
  • + Publishing native OpenClaw skills to ClawHub
  • + Deploying to the active skill
  • + Skill routing
  • + Cross-session decision memory
  • + GBrain Search Guidance (configured by /sync-gbrain)

Commands

5 shared · 2 only in A · 34 only in B
  • − bun run test:windows
  • − bun run gen:skill-docs --host codex
  • + bun run test:evals
  • + bun run test:evals:all
  • + bun run test:gate
  • + bun run test:periodic
  • + bun run test:e2e
  • + bun run test:e2e:all
  • + bun run eval:select
  • + bun run dev <cmd>
  • + bun run dev:skill
  • + bun run eval:list
  • + bun run eval:compare
  • + bun run eval:summary
  • + bun run slop
  • + bun run slop:diff
  • + npx slop-scan scan .
  • + npx slop-scan scan . --json
  • + git diff
  • + gh pr view
  • + gh repo view
  • + git pull
  • + git rm --cached
  • + git status
  • + git add file1 file2
  • + git add .
  • + git add -A
  • + git push --no-verify
  • + gh issue create
  • + git
  • + gh pr checkout
  • + gh pr checkout <N>
  • + git push origin HEAD:<branch-name>
  • + gh pr close <N> --comment "moving to base-repo branch for secret access"
  • + gh pr create --base main --head <branch-name>
  • + git log <prev-version>..HEAD --oneline
  •   bun install
  •   bun test
  •   bun run build
  •   bun run gen:skill-docs
  •   bun run skill:check

Section tags

8 shared · 0 only in A · 6 only in B
  • + lint-format
  • + architecture
  • + testing-strategy
  • + security
  • + do-not
  • + docs
  •   setup
  •   build
  •   test
  •   code-style
  •   git-pr
  •   performance
  •   deployment
  •   agent-behaviour

Line diff

+995 added−109 removed29 unchanged2.8% identical
garrytan/gstack · AGENTS.md
@@ −1 @@
1# gstack — AI Engineering Workflow
2 
3gstack is a collection of SKILL.md files that give AI agents structured roles for
4software development. Each skill is a specialist: CEO reviewer, eng manager,
5designer, QA lead, release engineer, debugger, and more.
6 
7## Available skills
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8 
9Skills live in `.agents/skills/` (or `~/.claude/skills/gstack/` on Claude Code).
10Invoke them by name (e.g., `/office-hours`).
11 
12### Plan-mode reviews
 
 
 
 
 
 
 
 
 
 
13 
14| Skill | What it does |
15|-------|-------------|
16| `/office-hours` | Start here. Reframes your product idea before you write code. |
17| `/plan-ceo-review` | CEO-level review: find the 10-star product in the request. |
18| `/plan-eng-review` | Lock architecture, data flow, edge cases, and tests. |
19| `/plan-design-review` | Rate each design dimension 0-10, explain what a 10 looks like. |
20| `/plan-devex-review` | DX-mode review: TTHW, magical moments, friction points, persona traces. |
21| `/plan-tune` | Self-tune AskUserQuestion sensitivity per question. |
22| `/autoplan` | One command runs CEO → design → eng → DX review. |
23| `/design-consultation` | Build a complete design system from scratch. |
24| `/spec` | Turn vague intent into a precise, executable spec in five phases. Files a GitHub issue, optionally spawns a Claude Code agent in a fresh worktree, and lets `/ship` close the source issue on merge. |
 
25 
26### Implementation + review
 
 
27 
28| Skill | What it does |
29|-------|-------------|
30| `/review` | Pre-landing PR review. Finds bugs that pass CI but break in prod. |
31| `/codex` | Second opinion via OpenAI Codex. Review, challenge, or consult modes. |
32| `/investigate` | Systematic root-cause debugging. No fixes without investigation. |
33| `/design-review` | Live-site visual audit + fix loop with atomic commits. |
34| `/design-shotgun` | Generate multiple AI design variants, comparison board, iterate. |
35| `/design-html` | Generate production-quality Pretext-native HTML/CSS. |
36| `/devex-review` | Live developer experience audit (TTHW measured against the real flow). |
37| `/qa` | Open a real browser, find bugs, fix them, re-verify. |
38| `/qa-only` | Same methodology as /qa but report only — no code changes. |
39| `/scrape` | Pull data from a web page. First call prototypes; codified call runs in ~200ms. |
40| `/skillify` | Codify the most recent successful `/scrape` flow into a permanent browser-skill. |
41 
42### Release + deploy
 
 
 
 
 
 
43 
44| Skill | What it does |
45|-------|-------------|
46| `/ship` | Run tests, review, push, open PR. Workspace-aware version queue. |
47| `/land-and-deploy` | Merge the PR, wait for CI and deploy, verify production health. |
48| `/canary` | Post-deploy monitoring loop using the browse daemon. |
49| `/landing-report` | Read-only dashboard for the workspace-aware ship queue. |
50| `/document-release` | Update all docs to match what you just shipped. |
51| `/document-generate` | Generate Diataxis docs (tutorial / how-to / reference / explanation) from code. |
52| `/setup-deploy` | One-time deploy config detection (Fly.io, Render, Vercel, etc.). |
53| `/gstack-upgrade` | Update gstack to the latest version. |
54 
55### Operational + memory
 
 
 
56 
57| Skill | What it does |
58|-------|-------------|
59| `/context-save` | Save working context (git state, decisions, remaining work). |
60| `/context-restore` | Resume from a saved context, even across Conductor workspaces. |
61| `/learn` | Manage what gstack learned across sessions. |
62| `/retro` | Weekly retro with per-person breakdowns and shipping streaks. |
63| `/health` | Code quality dashboard (type checker, linter, tests, dead code). |
64| `/benchmark` | Performance regression detection (page load, Core Web Vitals). |
65| `/benchmark-models` | Cross-model benchmark for skills (Claude, GPT, Gemini side-by-side). |
66| `/cso` | OWASP Top 10 + STRIDE security audit. |
67| `/setup-gbrain` | Set up gbrain for cross-machine session memory sync. |
68| `/sync-gbrain` | Keep gbrain current with this repo's code; refresh agent search guidance in CLAUDE.md. |
69 
70### Browser + agent integration
71 
72| Skill | What it does |
73|-------|-------------|
74| `/browse` | Headless browser — real Chromium, real clicks, ~100ms/command. |
75| `/open-gstack-browser` | Launch the visible GStack Browser with sidebar + stealth. |
76| `/setup-browser-cookies` | Import cookies from your real browser for authenticated testing. |
77| `/pair-agent` | Pair a remote AI agent (OpenClaw, Codex, etc.) with your browser. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78 
79### iOS QA — drive real iPhones over USB or Tailscale (v1.43.0.0+)
80 
81| Skill | What it does |
82|-------|-------------|
83| `/ios-qa` | Live-device iOS QA via USB CoreDevice tunnel + embedded StateServer. Optionally exposes the device over Tailscale so remote agents can drive it. |
84| `/ios-fix` | Autonomous iOS bug fixer with regression snapshot capture. |
85| `/ios-design-review` | Designer's-eye QA on a real iPhone — 10-dimension Apple HIG rubric. |
86| `/ios-clean` | Convenience: strip DebugBridge + #if DEBUG wiring before a Release build. |
87| `/ios-sync` | Regenerate the iOS debug bridge against the latest upstream templates. |
88 
89Companion CLIs (run on the Mac that's plugged into the device):
 
 
90 
91| Command | What it does |
92|---------|-------------|
93| `gstack-ios-qa-daemon` | Mac-side broker. Loopback by default; `--tailnet` adds a Tailscale-facing listener with capability tiers and audit logging. |
94| `gstack-ios-qa-mint` | Owner-grant CLI for the tailnet allowlist (`grant`/`revoke`/`list`). |
95| `gstack-ios-qa-regen` | Regenerate the canonical local DebugBridge package and typed accessors (`--app-source` / `--bridge-dir`). |
96 
97End-to-end walkthrough: [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-testing-with-gstack.md).
 
 
 
 
 
 
 
 
 
98 
99### Safety + scoping
 
 
 
 
100 
101| Skill | What it does |
102|-------|-------------|
103| `/careful` | Warn before destructive commands (rm -rf, DROP TABLE, force-push). |
104| `/freeze` | Lock edits to one directory. Hard block, not just a warning. |
105| `/guard` | Activate both careful + freeze at once. |
106| `/unfreeze` | Remove directory edit restrictions. |
107| `/make-pdf` | Turn any markdown file into a publication-quality PDF. |
108| `/diagram` | English in, diagram out: mermaid source + editable .excalidraw + SVG/PNG, offline. |
109 
110## Build commands
 
111 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
112```bash
113bun install # install dependencies
114bun test # run free tests (no API spend)
115bun run test:windows # curated Windows-safe subset (runs on windows-latest)
116bun run build # generate docs + compile binaries
117bun run gen:skill-docs # regenerate SKILL.md files from templates
118bun run skill:check # health dashboard for all skills
119```
120 
121## Platform support
122 
123- **macOS** + **Linux**: full test suite supported.
124- **Windows**: curated Windows-safe subset runs on `windows-latest` via the
125 `windows-free-tests` CI job. Setup script (`./setup`) requires Git Bash or
126 MSYS today; native PowerShell support is a future expansion. The `bin/gstack-paths`
127 helper resolves state roots through `CLAUDE_PLUGIN_DATA` / `GSTACK_HOME` so plugin
128 installs work on every platform.
129 
130## Key conventions
 
 
 
 
 
 
 
 
131 
132- SKILL.md files are **generated** from `.tmpl` templates. Edit the template, not the output.
133- Run `bun run gen:skill-docs --host codex` to regenerate Codex-specific output.
134- The browse binary provides headless browser access. Use `$B <command>` in skills.
135- Safety skills (careful, freeze, guard) use inline advisory prose — always confirm before destructive operations.
136- State paths resolve via `bin/gstack-paths` (sourced via `eval "$(...)"`). Honors `GSTACK_HOME`, `CLAUDE_PLUGIN_DATA`, `CLAUDE_PLANS_DIR`.
137- The `claude` CLI binary resolves via `browse/src/claude-bin.ts` (`Bun.which()` + `GSTACK_CLAUDE_BIN` override). Set `GSTACK_CLAUDE_BIN=wsl` plus `GSTACK_CLAUDE_BIN_ARGS='["claude"]'` to run Claude through WSL on Windows.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138 
garrytan/gstack · CLAUDE.md
@@ +1 @@
1# gstack development
2 
3## Commands
 
 
4 
5```bash
6bun install # install dependencies
7bun test # run free tests (browse + snapshot + skill validation)
8bun run test:evals # run paid evals: LLM judge + E2E (diff-based, ~$4/run max)
9bun run test:evals:all # run ALL paid evals regardless of diff
10bun run test:gate # run gate-tier tests only (CI default, blocks merge)
11bun run test:periodic # run periodic-tier tests only (weekly cron / manual)
12bun run test:e2e # run E2E tests only (diff-based, ~$3.85/run max)
13bun run test:e2e:all # run ALL E2E tests regardless of diff
14bun run eval:select # show which tests would run based on current diff
15bun run dev <cmd> # run CLI in dev mode, e.g. bun run dev goto https://example.com
16bun run build # gen docs + compile binaries
17bun run gen:skill-docs # regenerate SKILL.md files from templates
18bun run skill:check # health dashboard for all skills
19bun run dev:skill # watch mode: auto-regen + validate on change
20bun run eval:list # list all eval runs from ~/.gstack-dev/evals/
21bun run eval:compare # compare two eval runs (auto-picks most recent)
22bun run eval:summary # aggregate stats across all eval runs
23bun run slop # full slop-scan report (all files)
24bun run slop:diff # slop findings in files changed on this branch only
25```
26 
27`test:evals` requires `ANTHROPIC_API_KEY`. Codex E2E tests (`test/codex-e2e.test.ts`)
28use Codex's own auth from `~/.codex/` config — no `OPENAI_API_KEY` env var needed.
29 
30**Env keys in Conductor workspaces.** The `GSTACK_*` env-shim (v1.39.2.0+,
31`lib/conductor-env-shim.ts`) promotes `GSTACK_ANTHROPIC_API_KEY` /
32`GSTACK_OPENAI_API_KEY` to their canonical names inside gstack's TS binaries.
33Tests run through gstack entrypoints inherit this promotion automatically.
34Don't echo the key value to stdout, logs, or shell history. The historical
35"never pass `env:` to `runAgentSdkTest`" rule is retired: the failure was
36partial-env replacement (the SDK's `Options.env` REPLACES the child's entire
37environment, so an object without the key broke auth). The runner now always
38passes a COMPLETE hermetic env with per-test `env:` merged last, so per-test
39overrides are safe; ambient `process.env.ANTHROPIC_API_KEY` mutation also
40still works (the env builder reads process.env at call time).
41 
42**Hermetic local E2E (default).** Every E2E runner (claude -p, PTY, Agent
43SDK, codex, gemini) spawns children through `test/helpers/hermetic-env.ts`:
44allowlist-scrubbed env (operator `CONDUCTOR_*`, `CLAUDE_*`, `GSTACK_*`,
45`MCP_*`, `GBRAIN_*`, and credentials like `GH_TOKEN` never reach children),
46a fresh seeded `CLAUDE_CONFIG_DIR` (no operator `~/.claude` CLAUDE.md /
47MCP servers / skills), a temp `GSTACK_HOME`, and `--strict-mcp-config`.
48Local eval signal matches CI. Debug against real operator state with
49`EVALS_HERMETIC=0` (restores the legacy env AND drops the strict-MCP flag).
50Per-test `env:` overrides merge last, so deliberate contamination
51(`CONDUCTOR_WORKSPACE_PATH`, per-test `GSTACK_HOME`) keeps working. Wiring
52is pinned by `test/hermetic-wiring.test.ts` (static tripwire) and two
53gate-tier canaries in `test/skill-e2e-hermetic-canary.test.ts`.
54 
55E2E tests stream progress in real-time (tool-by-tool via `--output-format stream-json
56--verbose`). Results are persisted to `~/.gstack-dev/evals/` with auto-comparison
57against the previous run.
58 
59**Diff-based test selection:** `test:evals` and `test:e2e` auto-select tests based
60on `git diff` against the base branch. Each test declares its file dependencies in
61`test/helpers/touchfiles.ts`. Changes to global touchfiles (session-runner, eval-store,
62touchfiles.ts itself) trigger all tests. Use `EVALS_ALL=1` or the `:all` script
63variants to force all tests. Run `eval:select` to preview which tests would run.
 
 
 
 
 
 
 
 
64 
65**Two-tier system:** Tests are classified as `gate` or `periodic` in `E2E_TIERS`
66(in `test/helpers/touchfiles.ts`). CI runs only gate tests (`EVALS_TIER=gate`);
67periodic tests run weekly via cron or manually. Use `EVALS_TIER=gate` or
68`EVALS_TIER=periodic` to filter. When adding new E2E tests, classify them:
691. Safety guardrail or deterministic functional test? -> `gate`
702. Quality benchmark, Opus model test, or non-deterministic? -> `periodic`
713. Requires external service (Codex, Gemini)? -> `periodic`
72 
73## Testing
 
 
 
 
 
 
 
 
 
74 
75```bash
76bun test # run before every commit — free, <2s
77bun run test:evals # run before shipping — paid, diff-based (~$4/run max)
78```
79 
80`bun test` runs skill validation, gen-skill-docs quality checks, and browse
81integration tests. `bun run test:evals` runs LLM-judge quality evals and E2E
82tests via `claude -p`. Both must pass before creating a PR.
 
 
 
 
 
 
 
 
 
83 
84## Project structure
85 
86```
87gstack/
88├── browse/ # Headless browser CLI (Playwright)
89│ ├── src/ # CLI + server + commands
90│ │ ├── commands.ts # Command registry (single source of truth)
91│ │ └── snapshot.ts # SNAPSHOT_FLAGS metadata array
92│ ├── test/ # Integration tests + fixtures
93│ └── dist/ # Compiled binary
94├── hosts/ # Typed host configs (one per AI agent)
95│ ├── claude.ts # Primary host config
96│ ├── codex.ts, factory.ts, kiro.ts # Existing hosts
97│ ├── opencode.ts, slate.ts, cursor.ts, openclaw.ts # IDE hosts
98│ ├── hermes.ts, gbrain.ts # Agent runtime hosts
99│ └── index.ts # Registry: exports all, derives Host type
100├── scripts/ # Build + DX tooling
101│ ├── gen-skill-docs.ts # Template → SKILL.md generator (config-driven)
102│ ├── host-config.ts # HostConfig interface + validator
103│ ├── host-config-export.ts # Shell bridge for setup script
104│ ├── host-adapters/ # Host-specific adapters (OpenClaw tool mapping)
105│ ├── resolvers/ # Template resolver modules (preamble, design, review, gbrain, etc.)
106│ ├── skill-check.ts # Health dashboard
107│ └── dev-skill.ts # Watch mode
108├── test/ # Skill validation + eval tests
109│ ├── helpers/ # skill-parser.ts, session-runner.ts, llm-judge.ts, eval-store.ts
110│ ├── fixtures/ # Ground truth JSON, planted-bug fixtures, eval baselines
111│ ├── skill-validation.test.ts # Tier 1: static validation (free, <1s)
112│ ├── gen-skill-docs.test.ts # Tier 1: generator quality (free, <1s)
113│ ├── skill-llm-eval.test.ts # Tier 3: LLM-as-judge (~$0.15/run)
114│ └── skill-e2e-*.test.ts # Tier 2: E2E via claude -p (~$3.85/run, split by category)
115├── qa-only/ # /qa-only skill (report-only QA, no fixes)
116├── plan-design-review/ # /plan-design-review skill (report-only design audit)
117├── design-review/ # /design-review skill (design audit + fix loop)
118├── ship/ # Ship workflow skill
119├── review/ # PR review skill
120├── plan-ceo-review/ # /plan-ceo-review skill
121├── plan-eng-review/ # /plan-eng-review skill
122├── autoplan/ # /autoplan skill (auto-review pipeline: CEO → design → eng)
123├── benchmark/ # /benchmark skill (performance regression detection)
124├── canary/ # /canary skill (post-deploy monitoring loop)
125├── codex/ # /codex skill (multi-AI second opinion via OpenAI Codex CLI)
126├── land-and-deploy/ # /land-and-deploy skill (merge → deploy → canary verify)
127├── office-hours/ # /office-hours skill (YC Office Hours — startup diagnostic + builder brainstorm)
128├── investigate/ # /investigate skill (systematic root-cause debugging)
129├── spec/ # /spec skill (five-phase spec → GitHub issue, optional agent spawn, /ship auto-closes)
130├── retro/ # Retrospective skill (includes /retro global cross-project mode)
131├── bin/ # CLI utilities (gstack-repo-mode, gstack-slug, gstack-config, etc.)
132├── document-release/ # /document-release skill (post-ship doc updates + Diataxis coverage map)
133├── document-generate/ # /document-generate skill (Diataxis doc generator: tutorial/how-to/reference/explanation)
134├── cso/ # /cso skill (OWASP Top 10 + STRIDE security audit)
135├── design-consultation/ # /design-consultation skill (design system from scratch)
136├── design-shotgun/ # /design-shotgun skill (visual design exploration)
137├── open-gstack-browser/ # /open-gstack-browser skill (launch GStack Browser)
138├── connect-chrome/ # symlink → open-gstack-browser (backwards compat)
139├── design/ # Design binary CLI (GPT Image API)
140│ ├── src/ # CLI + commands (generate, variants, compare, serve, etc.)
141│ ├── test/ # Integration tests
142│ └── dist/ # Compiled binary
143├── extension/ # Chrome extension (side panel + activity feed + CSS inspector)
144├── lib/ # Shared libraries (worktree.ts)
145├── docs/designs/ # Design documents
146├── setup-deploy/ # /setup-deploy skill (one-time deploy config)
147├── .github/ # CI workflows + Docker image
148│ ├── workflows/ # evals.yml (E2E on Ubicloud), skill-docs.yml, actionlint.yml
149│ └── docker/ # Dockerfile.ci (pre-baked toolchain + Playwright/Chromium)
150├── contrib/ # Contributor-only tools (never installed for users)
151│ └── add-host/ # /gstack-contrib-add-host skill
152├── setup # One-time setup: build binary + symlink skills
153├── SKILL.md # Generated from SKILL.md.tmpl (don't edit directly)
154├── SKILL.md.tmpl # Template: edit this, run gen:skill-docs
155├── ETHOS.md # Builder philosophy (Boil the Ocean, Search Before Building)
156└── package.json # Build scripts for browse
157```
158 
159## SKILL.md workflow
160 
161SKILL.md files are **generated** from `.tmpl` templates. To update docs:
 
 
 
 
 
 
162 
1631. Edit the `.tmpl` file (e.g. `SKILL.md.tmpl` or `browse/SKILL.md.tmpl`)
1642. Run `bun run gen:skill-docs` (or `bun run build` which does it automatically)
1653. Commit both the `.tmpl` and generated `.md` files
166 
167To add a new browse command: add it to `browse/src/commands.ts` and rebuild.
168To add a snapshot flag: add it to `SNAPSHOT_FLAGS` in `browse/src/snapshot.ts` and rebuild.
 
 
 
169 
170**Token ceiling:** Generated SKILL.md files trip a warning above 160KB (~40K tokens).
171This is a "watch for feature bloat" guardrail, not a hard gate. Modern flagship
172models have 200K-1M context windows, so 40K is 4-20% of window, and prompt caching
173makes the marginal cost of larger skills small. The ceiling exists to catch runaway
174preamble/resolver growth, not to force compression on carefully-tuned big skills
175(`ship`, `plan-ceo-review`, `office-hours` legitimately pack 25-35K tokens of
176behavior). If you blow past 40K, the right fix is usually: (1) look at WHAT grew,
177(2) if one resolver added 10K+ in a single PR, question whether it belongs inline
178or as a reference doc, (3) only compress carefully-tuned prose as a last resort —
179cuts to the coverage audit, review army, or voice directive have real quality cost.
180 
181**Merge conflicts on SKILL.md files:** NEVER resolve conflicts on generated SKILL.md
182files by accepting either side. Instead: (1) resolve conflicts on the `.tmpl` templates
183and `scripts/gen-skill-docs.ts` (the sources of truth), (2) run `bun run gen:skill-docs`
184to regenerate all SKILL.md files, (3) stage the regenerated files. Accepting one side's
185generated output silently drops the other side's template changes.
186 
187## Platform-agnostic design
 
 
 
 
 
 
 
188 
189Skills must NEVER hardcode framework-specific commands, file patterns, or directory
190structures. Instead:
191 
1921. **Read CLAUDE.md** for project-specific config (test commands, eval commands, etc.)
1932. **If missing, AskUserQuestion** — let the user tell you or let gstack search the repo
1943. **Persist the answer to CLAUDE.md** so we never have to ask again
195 
196This applies to test commands, eval commands, deploy commands, and any other
197project-specific behavior. The project owns its config; gstack reads it.
198 
199## Writing SKILL templates
200 
201SKILL.md.tmpl files are **prompt templates read by Claude**, not bash scripts.
202Each bash code block runs in a separate shell — variables do not persist between blocks.
203 
204Rules:
205- **Use natural language for logic and state.** Don't use shell variables to pass
206 state between code blocks. Instead, tell Claude what to remember and reference
207 it in prose (e.g., "the base branch detected in Step 0").
208- **Don't hardcode branch names.** Detect `main`/`master`/etc dynamically via
209 `gh pr view` or `gh repo view`. Use `{{BASE_BRANCH_DETECT}}` for PR-targeting
210 skills. Use "the base branch" in prose, `<base>` in code block placeholders.
211- **Keep bash blocks self-contained.** Each code block should work independently.
212 If a block needs context from a previous step, restate it in the prose above.
213- **Express conditionals as English.** Instead of nested `if/elif/else` in bash,
214 write numbered decision steps: "1. If X, do Y. 2. Otherwise, do Z."
215 
216## Writing style (V1)
217 
218Default output from every tier-≥2 skill follows the Writing Style section in
219`scripts/resolvers/preamble.ts`: jargon glossed on first use (curated list in
220`scripts/jargon-list.json`, baked at gen-skill-docs time), questions framed in
221outcome terms ("what breaks for your users if...") not implementation terms,
222short sentences, decisions close with user impact. Power users who want the
223tighter V0 prose set `gstack-config set explain_level terse` (binary switch,
224no middle mode). See `docs/designs/PLAN_TUNING_V1.md` for the full design
225rationale. The review pacing overhaul that originally tried to ride alongside
226writing-style was extracted to V1.1 — see `docs/designs/PACING_UPDATES_V0.md`.
227 
228## Browser interaction
229 
230When you need to interact with a browser (QA, dogfooding, cookie setup), use the
231`/browse` skill or run the browse binary directly via `$B <command>`. NEVER use
232`mcp__claude-in-chrome__*` tools — they are slow, unreliable, and not what this
233project uses.
234 
235**Sidebar architecture:** Before modifying `sidepanel.js`, `background.js`,
236`content.js`, `terminal-agent.ts`, or sidebar-related server endpoints,
237read `docs/designs/SIDEBAR_MESSAGE_FLOW.md`. The sidebar has one primary
238surface — the **Terminal** pane (interactive `claude` PTY) — with
239Activity / Refs / Inspector as debug overlays behind the footer's
240`debug` toggle. The chat queue path was ripped once the PTY proved out;
241`sidebar-agent.ts` and the `/sidebar-command` / `/sidebar-chat` /
242`/sidebar-agent/event` endpoints are gone. The doc covers the WS auth
243flow, dual-token model, and threat-model boundary — silent failures
244here usually trace to not understanding the cross-component flow.
245 
246**Embedder terminal-agent ownership** (v1.42.1.0+, identity-based kill v1.44.0.0+).
247`buildFetchHandler` in `browse/src/server.ts` accepts `ServerConfig.ownsTerminalAgent?:
248boolean` (default `true`). When `true`, factory shutdown runs the full teardown:
249identity-based kill via `killAgentByRecord(readAgentRecord(stateDir))` from
250`browse/src/terminal-agent-control.ts` plus `safeUnlinkQuiet` on
251`<stateDir>/terminal-port`, `<stateDir>/terminal-internal-token`, and
252`<stateDir>/terminal-agent-pid` (the per-boot agent record introduced in v1.44).
253Embedders (e.g. the gbrowser phoenix overlay) that pre-launch their own PTY
254server must pass `false` so their discovery files survive gstack teardown cycles.
255The flag is the third caller-owned teardown gate in `ServerConfig` (alongside
256`xvfb?` and `proxyBridge?`); polarity is inverted (explicit bool vs presence) and
257documented in the field's JSDoc. CLI `start()` always passes `true` explicitly —
258the static-grep test in `browse/test/server-embedder-terminal-port.test.ts` fails
259CI if a refactor drops it. Pre-v1.44 used `pkill -f terminal-agent\.ts` (regex
260match) which would kill sibling gstack sessions on the same host; the new
261`browse/test/terminal-agent-pid-identity.test.ts` static-grep tripwire fails CI
262if any source file re-introduces `pkill ... terminal-agent` or `spawnSync('pkill', ...)`.
263 
264**WebSocket auth uses Sec-WebSocket-Protocol, not cookies.** Browsers
265can't set `Authorization` on a WebSocket upgrade, but they CAN set
266`Sec-WebSocket-Protocol` via `new WebSocket(url, [token])`. The agent
267reads it, validates against `validTokens`, and MUST echo the protocol
268back in the upgrade response — without the echo, Chromium closes the
269connection immediately. `Set-Cookie: gstack_pty=...` is kept as a
270fallback for non-browser callers (the cross-port `SameSite=Strict`
271cookie path doesn't survive from a chrome-extension origin).
272 
273**Cross-pane PTY injection.** The toolbar's Cleanup button and the
274Inspector's "Send to Code" action both pipe text into the live claude
275PTY via `window.gstackInjectToTerminal(text)`, exposed by
276`sidepanel-terminal.js`. No `/sidebar-command` POST — the live REPL is
277the only execution surface in the sidebar now.
278 
279**`/health` MUST NOT surface any shell-grant token.** It already leaks
280`AUTH_TOKEN` to localhost callers in headed mode (a v1.1+ TODO). Don't
281make that worse by adding the PTY session token there. PTY auth flows
282through `POST /pty-session` only.
283 
284**Transport-layer security** (v1.6.0.0+). When `pair-agent` starts an ngrok tunnel,
285the daemon binds two HTTP listeners: a local listener (127.0.0.1, full command
286surface, never forwarded) and a tunnel listener (locked allowlist: `/connect`,
287`/command` with a scoped token + 26-command browser-driving allowlist,
288`/sidebar-chat`). ngrok forwards only the tunnel port. Root tokens over the tunnel
289return 403. SSE endpoints use a 30-minute HttpOnly `gstack_sse` cookie minted via
290`POST /sse-session` (never valid against `/command`). Tunnel-surface rejections go
291to `~/.gstack/security/attempts.jsonl` via `tunnel-denial-log.ts`. Before editing
292`server.ts`, `sse-session-cookie.ts`, or `tunnel-denial-log.ts`, read
293[ARCHITECTURE.md](ARCHITECTURE.md#dual-listener-tunnel-architecture-v1600) —
294the module boundary (no imports from `token-registry.ts` into `sse-session-cookie.ts`)
295is load-bearing for scope isolation.
296 
297**Unicode sanitization at server egress** (v1.38.0.0+). Every server egress that
298ships page-content-derived strings MUST go through `JSON.stringify(payload,
299sanitizeReplacer)` for object payloads or `sanitizeLoneSurrogates(body)` for text
300bodies. Lone UTF-16 surrogate halves from CDP page content otherwise reach the
301Anthropic API as `\uD800`-style escapes and trigger a 400. Wired at four egress
302points today: `handleCommandInternal` (HTTP + batch via a sanitizing wrapper around
303`handleCommandInternalImpl`) and both SSE producers (`/activity/stream`,
304`/inspector/events`). Post-stringify regex is a no-op — `JSON.stringify` has
305already escaped the surrogate before regex could match, so the replacer must run
306inside the encoding pipeline. Before adding a new SSE/WebSocket writer or HTTP
307response in `server.ts`, read
308[ARCHITECTURE.md](ARCHITECTURE.md#unicode-sanitization-at-server-egress-v13800).
309`browse/test/server-sanitize-surrogates.test.ts` pins the wiring with invariant
310tests, so bypasses fail CI.
311 
312**SSE endpoint helper** (v1.51.0.0+). New SSE endpoints in `server.ts` MUST route
313through `createSseEndpoint(req, config)` from `browse/src/sse-helpers.ts`. The
314helper owns the cleanup contract (abort + enqueue-throw + heartbeat-throw, all
315idempotent) and bakes in `sanitizeLoneSurrogates` on every JSON.stringify, so
316new subscribers can't accidentally regress either invariant. Inline
317`ReadableStream` wiring leaked subscribers when the TCP connection died without
318firing `req.signal.abort` (Chromium MV3 service-worker suspend, intermediate
319proxy half-close). `/activity/stream`, `/inspector/events`, and `/memory`
320(SSE-eligible) all route through it. `browse/test/sse-helpers.test.ts` pins the
321cleanup contract.
322 
323**CDP session lifecycle** (v1.51.0.0+). Direct `page.context().newCDPSession(page)`
324calls outside `browse/src/cdp-bridge.ts` fail CI via the static-grep tripwire in
325`browse/test/cdp-session-cleanup.test.ts`. Use `withCdpSession(page, async (s) => {...})`
326for one-shot CDP work (try/finally detach) or `getOrCreateCdpSession(page, cache)`
327for cached sessions tied to a page's lifetime (close-detach via `Map<page, session>`).
328Three sites migrated: cdp-bridge frame events, write-commands archive capture,
329cdp-inspector. The helpers prevent the per-session leak class where successful-path
330detach happened but error-path detach was missed.
331 
332**Setup symlink hardening** (v1.38.0.0+). Every link site in `setup` MUST route
333through the `_link_or_copy SRC DST` helper near the `IS_WINDOWS` detection. On
334Windows without Developer Mode, plain `ln -snf` produces frozen file copies that
335don't refresh on `git pull` — silent staleness across every host adapter. The
336helper preserves `ln -snf` on Unix and switches to `cp -R` / `cp -f` on Windows.
337`test/setup-windows-fallback.test.ts` enforces a static invariant: a single raw
338`ln` call outside the helper body fails CI. Windows users get a one-line note
339from `_print_windows_copy_note_once` reminding them to re-run `./setup` after
340every `git pull`.
341 
342**Sidebar security stack** (layered defense against prompt injection):
343 
344| Layer | Module | Lives in |
345|-------|--------|----------|
346| L1-L3 | `content-security.ts` | both server and agent — datamarking, hidden element strip, ARIA regex, URL blocklist, envelope wrapping |
347| L4 | `security-classifier.ts` (TestSavantAI ONNX) | **sidebar-agent only** |
348| L4b | `security-classifier.ts` (Claude Haiku transcript) | **sidebar-agent only** |
349| L5 | `security.ts` (canary) | both — inject in compiled, check in agent |
350| L6 | `security.ts` (combineVerdict ensemble) | both |
351 
352**Critical constraint:** `security-classifier.ts` CANNOT be imported from the
353compiled browse binary. `@huggingface/transformers` v4 requires `onnxruntime-node`
354which fails to `dlopen` from Bun compile's temp extract dir. Only `security.ts`
355(pure-string operations — canary, verdict combiner, attack log, status) is safe
356for `server.ts`. See `~/.gstack/projects/garrytan-gstack/ceo-plans/2026-04-19-prompt-injection-guard.md`
357§"Pre-Impl Gate 1 Outcome" for full architectural decision.
358 
359**Thresholds** (in `security.ts`):
360- `BLOCK: 0.85` — single-layer score that would cause BLOCK if cross-confirmed
361- `WARN: 0.75` — cross-confirm threshold. When L4 AND L4b both >= 0.75 → BLOCK
362- `LOG_ONLY: 0.40` — gates transcript classifier (skip Haiku when all layers < 0.40)
363- `SOLO_CONTENT_BLOCK: 0.92` — single-layer threshold for label-less content classifiers
364 (testsavant, deberta). Intentionally higher than `BLOCK` because these layers can't
365 distinguish "this is an injection" from "this looks like phishing aimed at the user."
366 The transcript classifier keeps a separate, label-gated solo path at `BLOCK` (0.85).
367 
368**Ensemble rule:** BLOCK only when the ML content classifier AND the transcript
369classifier both report >= WARN. Single-layer high confidence degrades to WARN —
370this is the Stack Overflow instruction-writing FP mitigation. Canary leak
371always BLOCKs (deterministic).
372 
373**Env knobs:**
374- `GSTACK_SECURITY_OFF=1` — emergency kill switch. Classifier stays off even if
375 warmed. Canary is still injected; just the ML scan is skipped.
376- `GSTACK_SECURITY_ENSEMBLE=deberta` — opt-in DeBERTa-v3 ensemble. Adds
377 ProtectAI DeBERTa-v3-base-injection-onnx as L4c classifier for cross-model
378 agreement. 721MB first-run download. With ensemble enabled, BLOCK requires
379 2-of-3 ML classifiers agreeing at >= WARN (testsavant, deberta, transcript).
380 Without ensemble (default), BLOCK requires testsavant + transcript at >= WARN.
381- Classifier model cache: `~/.gstack/models/testsavant-small/` (112MB, first run only)
382 plus `~/.gstack/models/deberta-v3-injection/` (721MB, only when ensemble enabled)
383- Attack log: `~/.gstack/security/attempts.jsonl` (salted sha256 + domain only,
384 rotates at 10MB, 5 generations)
385- Per-device salt: `~/.gstack/security/device-salt` (0600)
386- Session state: `~/.gstack/security/session-state.json` (cross-process, atomic)
387 
388## Dev symlink awareness
389 
390When developing gstack, `.claude/skills/gstack` may be a symlink back to this
391working directory (gitignored). This means skill changes are **live immediately**,
392great for rapid iteration, risky during big refactors where half-written skills
393could break other Claude Code sessions using gstack concurrently.
394 
395**Check once per session:** Run `ls -la .claude/skills/gstack` to see if it's a
396symlink or a real copy. If it's a symlink to your working directory, be aware that:
397- Template changes + `bun run gen:skill-docs` immediately affect all gstack invocations
398- Breaking changes to SKILL.md.tmpl files can break concurrent gstack sessions
399- During large refactors, remove the symlink (`rm .claude/skills/gstack`) so the
400 global install at `~/.claude/skills/gstack/` is used instead
401 
402**Prefix setting:** Setup creates real directories (not symlinks) at the top level
403with a SKILL.md symlink inside (e.g., `qa/SKILL.md -> gstack/qa/SKILL.md`). This
404ensures Claude discovers them as top-level skills, not nested under `gstack/`.
405Names are either short (`qa`) or namespaced (`gstack-qa`), controlled by
406`skill_prefix` in `~/.gstack/config.yaml`. Pass `--no-prefix` or `--prefix` to
407skip the interactive prompt.
408 
409**Note:** Vendoring gstack into a project's repo is deprecated. Use global install
410+ `./setup --team` instead. See README.md for team mode instructions.
411 
412**For plan reviews:** When reviewing plans that modify skill templates or the
413gen-skill-docs pipeline, consider whether the changes should be tested in isolation
414before going live (especially if the user is actively using gstack in other windows).
415 
416**Upgrade migrations:** When a change modifies on-disk state (directory structure,
417config format, stale files) in ways that could break existing user installs, add a
418migration script to `gstack-upgrade/migrations/`. Read CONTRIBUTING.md's "Upgrade
419migrations" section for the format and testing requirements. The upgrade skill runs
420these automatically after `./setup` during `/gstack-upgrade`.
421 
422## Compiled binaries — NEVER commit browse/dist/ or design/dist/
423 
424The `browse/dist/` and `design/dist/` directories contain compiled Bun binaries
425(`browse`, `find-browse`, `design`, ~58MB each). These are Mach-O arm64 only — they
426do NOT work on Linux, Windows, or Intel Macs. The `./setup` script already builds
427from source for every platform, so the checked-in binaries are redundant. They are
428tracked by git due to a historical mistake and should eventually be removed with
429`git rm --cached`.
430 
431**NEVER stage or commit these files.** They show up as modified in `git status`
432because they're tracked despite `.gitignore` — ignore them. When staging files,
433always use specific filenames (`git add file1 file2`) — never `git add .` or
434`git add -A`, which will accidentally include the binaries.
435 
436## Redaction guard (PII / secrets / legal content)
437 
438Shared redaction engine catches credentials, PII, and legal/damaging content
439before it reaches an external sink (codex dispatch, GitHub issue/PR body, pushed
440commit). It is a **guardrail, not airtight enforcement** — `git push --no-verify`,
441direct `gh issue create`, and `GSTACK_REDACT_PREPUSH=skip` all bypass it. It
442catches accidents and carelessness, the 99% case. Do not claim it stops a
443determined leaker (a CHANGELOG line that does would fail a hostile screenshotter).
444 
445- **Engine + taxonomy:** `lib/redact-patterns.ts` (the single source of truth —
446 3 tiers; HIGH = genuinely-secret credentials that block, MEDIUM = PII/legal/
447 internal + high-FP credential shapes that confirm via AskUserQuestion, LOW =
448 FYI) and `lib/redact-engine.ts` (pure `scan()` + `applyRedactions()`).
449 Calibration matters: a gate that cries wolf gets ignored, so context-variable
450 shapes (Stripe `pk_live_`, Google `AIza`, JWT, env `*_KEY=`) sit at MEDIUM.
451- **CLI:** `bin/gstack-redact` (exit 0 clean / 2 MEDIUM / 3 HIGH; `--json`,
452 `--auto-redact`, `--repo-visibility`, `--from-file`). `bin/gstack-redact-prepush`
453 is the opt-in git hook.
454- **Skill docs are generated** from `scripts/resolvers/redact-doc.ts`
455 (`{{REDACT_TAXONOMY_TABLE}}`, `{{REDACT_INVOCATION_BLOCK:<sink>}}`) so /spec,
456 /cso, /ship, /document-release, /document-generate never drift from the engine.
457- **Scan-at-sink:** always scan the EXACT bytes that will be sent — write to a
458 temp file, scan that file, pass the SAME file to `gh`/`git`. Never scan a string
459 then re-render (that reopens a scan-vs-send gap).
460- **Visibility (no tier promotion):** resolve once per run, order = local config
461 (`gstack-config get redact_repo_visibility`, ~/.gstack so never committed) → gh
462 → glab → unknown(=public-strict). Public repos get STERNER per-finding
463 confirmation (no batch-acknowledge, no silent-proceed); MEDIUM is never
464 auto-promoted to HIGH.
465- **Tool-attributed fences:** wrap Codex/Greptile/eval output in ` ```codex-review `
466 / ` ```greptile ` fences so example credentials those tools quote WARN-degrade
467 instead of blocking. A live-format credential inside the fence still blocks.
468- **Config keys:** `redact_repo_visibility` (public|private|unknown, local-only
469 override for repos gh/glab can't read), `redact_prepush_hook` (true|false).
470 There is intentionally NO key to disable HIGH blocking.
471- **Audit:** the /spec semantic pass appends a content-free record (categories +
472 body sha256, no spec text) to `~/.gstack/security/semantic-reviews.jsonl` (0600).
473 
474## Commit style
475 
476**Always bisect commits.** Every commit should be a single logical change. When
477you've made multiple changes (e.g., a rename + a rewrite + new tests), split them
478into separate commits before pushing. Each commit should be independently
479understandable and revertable.
480 
481Examples of good bisection:
482- Rename/move separate from behavior changes
483- Test infrastructure (touchfiles, helpers) separate from test implementations
484- Template changes separate from generated file regeneration
485- Mechanical refactors separate from new features
486 
487When the user says "bisect commit" or "bisect and push," split staged/unstaged
488changes into logical commits and push.
489 
490## Slop-scan: AI code quality, not AI code hiding
491 
492We use [slop-scan](https://github.com/benvinegar/slop-scan) to catch patterns where
493AI-generated code is genuinely worse than what a human would write. We are NOT trying
494to pass as human code. We are AI-coded and proud of it. The goal is code quality.
495 
496```bash
497npx slop-scan scan . # human-readable report
498npx slop-scan scan . --json # machine-readable for diffing
 
 
 
 
499```
500 
501Config: `slop-scan.config.json` at repo root (currently excludes `**/vendor/**`).
502 
503### What to fix (genuine quality improvements)
 
 
 
 
 
504 
505- **Empty catches around file ops** — use `safeUnlink()` (ignores ENOENT, rethrows
506 EPERM/EIO). A swallowed EPERM in cleanup means silent data loss.
507- **Empty catches around process kills** — use `safeKill()` (ignores ESRCH, rethrows
508 EPERM). A swallowed EPERM means you think you killed something you didn't.
509- **Redundant `return await`** — remove when there's no enclosing try block. Saves a
510 microtask, signals intent.
511- **Typed exception catches** — `catch (err) { if (!(err instanceof TypeError)) throw err }`
512 is genuinely better than `catch {}` when the try block does URL parsing or DOM work.
513 You know what error you expect, so say so.
514 
515### What NOT to fix (linter gaming, not quality)
516 
517- **String-matching on error messages** — `err.message.includes('closed')` is brittle.
518 Playwright/Chrome can change wording anytime. If a fire-and-forget operation can fail
519 for ANY reason and you don't care, `catch {}` is the correct pattern.
520- **Adding comments to exempt pass-through wrappers** — "alias for active session" above
521 a method just to trip slop-scan's exemption rule is noise, not documentation.
522- **Converting extension catch-and-log to selective rethrow** — Chrome extensions crash
523 entirely on uncaught errors. If the catch logs and continues, that IS the right pattern
524 for extension code. Don't make it throw.
525- **Tightening best-effort cleanup paths** — shutdown, emergency cleanup, and disconnect
526 code should use `safeUnlinkQuiet()` (swallows ALL errors). A cleanup path that throws
527 on EPERM means the rest of cleanup doesn't run. That's worse.
528 
529### Utilities in `browse/src/error-handling.ts`
530 
531| Function | Use when | Behavior |
532|----------|----------|----------|
533| `safeUnlink(path)` | Normal file deletion | Ignores ENOENT, rethrows others |
534| `safeUnlinkQuiet(path)` | Shutdown/emergency cleanup | Swallows all errors |
535| `safeKill(pid, signal)` | Sending signals | Ignores ESRCH, rethrows others |
536| `isProcessAlive(pid)` | Boolean process checks | Returns true/false, never throws |
537 
538### Score tracking
539 
540Baseline (2026-04-09, before cleanup): 100 findings, 432.8 score, 2.38 score/file.
541After cleanup: 90 findings, 358.1 score, 1.96 score/file.
542 
543Don't chase the number. Fix patterns that represent actual code quality problems.
544Accept findings where the "sloppy" pattern is the correct engineering choice.
545 
546## Community PR guardrails
547 
548When reviewing or merging community PRs, **always AskUserQuestion** before accepting
549any commit that:
550 
5511. **Touches ETHOS.md** — this file is Garry's personal builder philosophy. No edits
552 from external contributors or AI agents, period.
5532. **Removes or softens promotional material** — YC references, founder perspective,
554 and product voice are intentional. PRs that frame these as "unnecessary" or
555 "too promotional" must be rejected.
5563. **Changes Garry's voice** — the tone, humor, directness, and perspective in skill
557 templates, CHANGELOG, and docs are not generic. PRs that rewrite voice to be
558 more "neutral" or "professional" must be rejected.
559 
560Even if the agent strongly believes a change improves the project, these three
561categories require explicit user approval via AskUserQuestion. No exceptions.
562No auto-merging. No "I'll just clean this up."
563 
564## Checking out PRs from garrytan-agents
565 
566When the user says "check out <PR link>" and the PR is from `garrytan-agents/gstack`
567(or any other fork that is NOT a collaborator on `garrytan/gstack`), do NOT just
568`gh pr checkout`. Fork PRs don't receive base-repo secrets (`ANTHROPIC_API_KEY`,
569`OPENAI_API_KEY`, etc.), so the eval/E2E CI jobs fail with empty-env auth errors
570regardless of what's set on the base repo.
571 
572**Workflow:** push the branch to `garrytan/gstack` (the base repo) and re-target
573the PR from there.
574 
575Concretely, after `gh pr checkout <N>`:
576 
5771. Note the original PR number and head branch name.
5782. Push the same branch to the base repo: `git push origin HEAD:<branch-name>`
579 (origin = `garrytan/gstack`, since the worktree is set up with that remote).
5803. Close the fork PR (`gh pr close <N> --comment "moving to base-repo branch for secret access"`).
5814. Open a new PR from the base-repo branch: `gh pr create --base main --head <branch-name>`.
5825. New PR's workflows will get secrets automatically.
583 
584Why not fix it on the fork side? `garrytan-agents` isn't a collaborator on
585`garrytan/gstack`. Adding it as a collaborator (option A) or flipping the
586repo-wide "send secrets to fork PRs" toggle (option B) would let secrets reach
587fork PRs from anyone — broader blast radius than just moving this one branch.
588Option C (this section) keeps secret-distribution scope tight.
589 
590If the user asks you to skip the move (e.g., "just leave it as a fork PR"),
591respect that — eval CI will fail with empty-env auth, but check-freshness,
592workflow-lint, and windows-tests will still pass on the fork PR.
593 
594## CHANGELOG + VERSION style
595 
596**Versioning invariant (workspace-aware ship).** VERSION is a monotonic ordered
597release identifier, not a strict semver commitment. The bump level
598(major/minor/patch/micro) expresses intent at ship time. Queue-advancing past a
599claimed version within the same bump level is explicitly permitted — if branch A
600claims v1.7.0.0 as a MINOR and branch B is also a MINOR, B lands at v1.8.0.0
601(still a MINOR relative to main). Downstream consumers must NOT rely on
602"MINOR = feature-only, PATCH = fix-only" as a strict contract. This is why
603`bin/gstack-next-version` advances within the chosen bump level rather than
604repicking the level when collisions happen.
605 
606**Scale-aware bumps — use common sense.** When the diff is big, bump MINOR (or
607MAJOR), not PATCH. PATCH is for bug fixes and small additions; MINOR is for
608substantial new capability or substantial reduction; MAJOR is for breaking
609changes. Rough guideposts (don't treat as rules, treat as smell-checks):
610 
611- **PATCH (X.Y.Z+1.0)**: bug fix, doc tweak, small additive change, single
612 test/file added. Net diff under ~500 lines, no new user-facing capability.
613- **MINOR (X.Y+1.0.0)**: new capability shipped (skill, harness, command, big
614 refactor), substantial code reduction (compression, migration), or coordinated
615 multi-file change. Net diff over ~2000 lines added/removed, OR a user-visible
616 feature you'd put in a tweet.
617- **MAJOR (X+1.0.0.0)**: breaking change to public surface (CLI flag rename,
618 skill removed, config format changed), OR a release big enough to be the
619 headline of a blog post.
620 
621If you find yourself debating "is 10K added + 24K removed really a PATCH?" — it
622isn't. Bump MINOR. Same for "this adds a whole new test harness with 6 new E2E
623tests + helper utilities" — MINOR. The bump level is communication to the user
624about what kind of release this is; don't undersell it.
625 
626When merging origin/main brings a higher VERSION, re-evaluate the bump level
627against the SCALE of your branch's work, not just whether main moved forward.
628If main bumped MINOR and your branch is also a substantial change, you bump
629MINOR again on top (e.g., main at v1.14.0.0, your branch lands v1.15.0.0).
630 
631**VERSION and CHANGELOG are branch-scoped.** Every feature branch that ships gets its
632own version bump and CHANGELOG entry. The entry describes what THIS branch adds —
633not what was already on main.
634 
635**The CHANGELOG entry is the diff between main and the shipping branch — what users
636get when they upgrade. NOT how the branch got there.** A reader landing on the entry
637should learn what they can do now that they couldn't before; they should not learn
638about the branch's internal version bumps, the bugs we caught and fixed mid-branch,
639the plan reviews we ran, or the commits we squashed. That is branch development
640narrative. It belongs in PR descriptions and commit messages, not CHANGELOG.
641 
642**Never reference branch-internal versions in a CHANGELOG entry.** If your branch
643bumped VERSION from v1.5.0.0 → v1.5.1.0 → v1.6.0.0 during development and only the
644final v1.6.0.0 ships to main, the entry must read as if v1.5.1.0 never existed.
645Concretely, NEVER write:
646- "v1.5.1.0 had a bug that v1.6.0.0 fixes" — readers don't know about v1.5.1.0; it's
647 a branch-internal artifact.
648- "The shipping headline of v1.5.1.0 was broken because..." — same reason. From main's
649 perspective, v1.5.1.0 was never released.
650- "Pre-fix tests encoded the broken behavior" — that's a contributor's victory lap,
651 not a user benefit.
652- "Two surgical edits, both in the dispatch path" — micro-narrative of the patch.
653 
654Instead, describe the released system: "Browser-skills run end-to-end with the
655expected tab-access semantics." If a property of the shipped system is worth calling
656out (e.g., "skill spawns get permissive tab access; pair-agent tunnel tokens require
657ownership"), document it as a property, not as a fix. The shipped system is what
658the user gets; the path to that system is invisible to them.
659 
660**When to write the CHANGELOG entry:**
661- At `/ship` time (Step 13), not during development or mid-branch.
662- The entry covers ALL commits on this branch vs the base branch.
663- Never fold new work into an existing CHANGELOG entry from a prior version that
664 already landed on main. If main has v0.10.0.0 and your branch adds features,
665 bump to v0.10.1.0 with a new entry — don't edit the v0.10.0.0 entry.
666 
667**Key questions before writing:**
6681. What branch am I on? What did THIS branch change?
6692. Is the base branch version already released? (If yes, bump and create new entry.)
6703. Does an existing entry on this branch already cover earlier work? (If yes, replace
671 it with one unified entry for the final version.)
672 
673**Merging main does NOT mean adopting main's version.** When you merge origin/main into
674a feature branch, main may bring new CHANGELOG entries and a higher VERSION. Your branch
675still needs its OWN version bump on top. If main is at v0.13.8.0 and your branch adds
676features, bump to v0.13.9.0 with a new entry. Never jam your changes into an entry that
677already landed on main. Your entry goes on top because your branch lands next.
678 
679**After merging main, always check:**
680- Does CHANGELOG have your branch's own entry separate from main's entries?
681- Is VERSION higher than main's VERSION?
682- Is your entry the topmost entry in CHANGELOG (above main's latest)?
683If any answer is no, fix it before continuing.
684 
685**After any CHANGELOG edit that moves, adds, or removes entries,** immediately run
686`grep "^## \[" CHANGELOG.md` to verify no duplicates and a sensible reverse-chronological
687order. Gaps between version numbers are fine. A branch that ships at v1.6.4.0 without
688a prior v1.5.2.0 or v1.5.3.0 entry on main is correct — those were branch-internal
689version numbers that never landed. Do not back-fill gaps with placeholder entries.
690 
691**Never orphan branch-internal versions.** If your branch bumped VERSION several times
692during development (v1.5.1.0 → v1.5.2.0 → v1.6.4.0, say) and those earlier entries were
693never released to main, the final ship consolidates ALL of them into a single entry at
694the final version (v1.6.4.0). Collapse them — delete the old entries and move their
695content into the final entry, re-version table columns accordingly. Readers see one
696release, not a branch diary. Gaps are fine (v1.6.3.0 → v1.6.4.0 with no v1.5.x
697in between on main is correct).
698 
699CHANGELOG.md is **for users**, not contributors. Write it like product release notes:
700 
701- Lead with what the user can now **do** that they couldn't before. Sell the feature.
702- Use plain language, not implementation details. "You can now..." not "Refactored the..."
703- **Never mention TODOS.md, internal tracking, eval infrastructure, or contributor-facing
704 details.** These are invisible to users and meaningless to them.
705- Put contributor/internal changes in a separate "For contributors" section at the bottom.
706- Every entry should make someone think "oh nice, I want to try that."
707- No jargon: say "every question now tells you which project and branch you're in" not
708 "AskUserQuestion format standardized across skill templates via preamble resolver."
709 
710**Only document what shipped between main and this change.** Readers do not care how
711we got here. Keep out of the CHANGELOG, always:
712 
713- Branch resyncs, merge commits with main, rebase activity.
714- Plan approvals, review outcomes (CEO / eng / design / outside-voice / codex findings),
715 AskUserQuestion decisions, scope negotiations.
716- "Work queued," "plan approved," "in-progress," "will ship later" — the CHANGELOG
717 documents what DID ship, not what MIGHT ship.
718- Version-bump housekeeping when no user-facing work actually landed.
719 
720If the diff between the base branch version and this version has no user-facing change
721(only merges, only CHANGELOG edits, only placeholder work), the honest entry is one
722sentence: "Version bump for branch-ahead discipline. No user-facing changes yet." Stop
723there. Do not pad. Do not explain the plan that will ship eventually. Do not narrate
724the branch's history. When real work lands, the entry will replace this at /ship time.
725 
726### Release-summary format (every `## [X.Y.Z]` entry)
727 
728Every version entry in `CHANGELOG.md` MUST start with a release-summary section in
729the GStack/Garry voice, one viewport's worth of prose + tables that lands like a
730verdict, not marketing. The itemized changelog (subsections, bullets, files) goes
731BELOW that summary, separated by a `### Itemized changes` header.
732 
733The release-summary section gets read by humans, by the auto-update agent, and by
734anyone deciding whether to upgrade. The itemized list is for agents that need to
735know exactly what changed.
736 
737Structure for the top of every `## [X.Y.Z]` entry:
738 
7391. **Two-line bold headline** (10-14 words total). Should land like a verdict, not
740 marketing. Sound like someone who shipped today and cares whether it works.
7412. **Lead paragraph** (3-5 sentences). What shipped, what changed for the user.
742 Specific, concrete, no AI vocabulary, no em dashes, no hype.
7433. **A "The X numbers that matter" section** with:
744 - One short setup paragraph naming the source of the numbers (real production
745 deployment OR a reproducible benchmark, name the file/command to run).
746 - A table of 3-6 key metrics with BEFORE / AFTER / Δ columns.
747 - A second optional table for per-category breakdown if relevant.
748 - 1-2 sentences interpreting the most striking number in concrete user terms.
7494. **A "What this means for [audience]" closing paragraph** (2-4 sentences) tying
750 the metrics to a real workflow shift. End with what to do.
751 
752Voice rules for the release summary:
753- No em dashes (use commas, periods, "...").
754- No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or
755 banned phrases ("here's the kicker", "the bottom line", etc.).
756- Real numbers, real file names, real commands. Not "fast" but "~30s on 30K pages."
757- Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
758- Connect to user outcomes: "the agent does ~3x less reading" beats "improved precision."
759- Be direct about quality. "Well-designed" or "this is a mess." No dancing.
760 
761Source material:
762- CHANGELOG previous entry for prior context.
763- Benchmark files or `/retro` output for headline numbers.
764- Recent commits (`git log <prev-version>..HEAD --oneline`) for what shipped.
765- Don't make up numbers. If a metric isn't in a benchmark or production data,
766 don't include it. Say "no measurement yet" if asked.
767 
768Target length: ~250-350 words for the summary. Should render as one viewport.
769 
770### Itemized changes (below the release summary)
771 
772Write `### Itemized changes` and continue with the detailed subsections (Added,
773Changed, Fixed, For contributors). Same rules as the user-facing voice guidance
774above, plus:
775 
776- **Always credit community contributions.** When an entry includes work from a
777 community PR, name the contributor with `Contributed by @username`. Contributors
778 did real work. Thank them publicly every time, no exceptions.
779 
780## AI effort compression
781 
782When estimating or discussing effort, always show both human-team and CC+gstack time:
783 
784| Task type | Human team | CC+gstack | Compression |
785|-----------|-----------|-----------|-------------|
786| Boilerplate / scaffolding | 2 days | 15 min | ~100x |
787| Test writing | 1 day | 15 min | ~50x |
788| Feature implementation | 1 week | 30 min | ~30x |
789| Bug fix + regression test | 4 hours | 15 min | ~20x |
790| Architecture / design | 2 days | 4 hours | ~5x |
791| Research / exploration | 1 day | 3 hours | ~3x |
792 
793Completeness is cheap. Don't recommend shortcuts when the complete implementation
794is achievable. Boil the ocean — the complete thing is the goal; only genuinely
795unrelated multi-quarter migrations are separate scope, never an excuse for a
796shortcut. See the Completeness Principle in the skill preamble for the full
797philosophy.
798 
799## Search before building
800 
801Before designing any solution that involves concurrency, unfamiliar patterns,
802infrastructure, or anything where the runtime/framework might have a built-in:
803 
8041. Search for "{runtime} {thing} built-in"
8052. Search for "{thing} best practice {current year}"
8063. Check official runtime/framework docs
807 
808Three layers of knowledge: tried-and-true (Layer 1), new-and-popular (Layer 2),
809first-principles (Layer 3). Prize Layer 3 above all. See ETHOS.md for the full
810builder philosophy.
811 
812## Local plans
813 
814Contributors can store long-range vision docs and design documents in `~/.gstack-dev/plans/`.
815These are local-only (not checked in). When reviewing TODOS.md, check `plans/` for candidates
816that may be ready to promote to TODOs or implement.
817 
818## E2E eval failure blame protocol
819 
820When an E2E eval fails during `/ship` or any other workflow, **never claim "not
821related to our changes" without proving it.** These systems have invisible couplings —
822a preamble text change affects agent behavior, a new helper changes timing, a
823regenerated SKILL.md shifts prompt context.
824 
825**Required before attributing a failure to "pre-existing":**
8261. Run the same eval on main (or base branch) and show it fails there too
8272. If it passes on main but fails on the branch — it IS your change. Trace the blame.
8283. If you can't run on main, say "unverified — may or may not be related" and flag it
829 as a risk in the PR body
830 
831"Pre-existing" without receipts is a lazy claim. Prove it or don't say it.
832 
833## Long-running tasks: don't give up
834 
835When running evals, E2E tests, or any long-running background task, **poll until
836completion**. Use `sleep 180 && echo "ready"` + `TaskOutput` in a loop every 3
837minutes. Never switch to blocking mode and give up when the poll times out. Never
838say "I'll be notified when it completes" and stop checking — keep the loop going
839until the task finishes or the user tells you to stop.
840 
841The full E2E suite can take 30-45 minutes. That's 10-15 polling cycles. Do all of
842them. Report progress at each check (which tests passed, which are running, any
843failures so far). The user wants to see the run complete, not a promise that
844you'll check later.
845 
846## Running evals as an agent: always detach (SIGTERM-proof)
847 
848When **you (an agent/harness)** launch a long eval/benchmark run, run it through
849`bin/gstack-detach` — NEVER as a plain backgrounded Bash task. A plain background
850task lives in the harness's process group, so a SIGTERM ("polite quit") on a turn
851boundary, a stopped Monitor, or an interruption kills the run mid-flight (observed:
852`script "test:gate" was terminated by signal SIGTERM` ~40 min into a run). On macOS
853the run can also die to idle-sleep. `gstack-detach` fixes both: a fresh session
854(escapes the group SIGTERM) wrapped in `caffeinate -i` (blocks idle-sleep).
855 
856- Use the `eval:bg*` scripts (`eval:bg`, `eval:bg:all`, `eval:bg:gate`,
857 `eval:bg:periodic`) — they wrap the eval command in `gstack-detach` with the
858 machine-wide `gstack-evals` lock (concurrent worktrees serialize instead of
859 saturating the shared model API), a per-tier watchdog, and a **run-scoped** log
860 under `~/.gstack-dev/eval-runs/` (no shared-`/tmp` collision). Each prints its
861 log path. Or call `gstack-detach [--lock NAME] [--timeout SECS] [--label LBL] --
862 <cmd>` directly for any long agent job. Export `ANTHROPIC_API_KEY` first (never
863 pass keys in argv).
864- Then **poll the printed logfile** with a death-aware watcher: break on the
865 guaranteed `### gstack-detach EXIT=<code> ###` sentinel (success AND failure are
866 both marked, so silence is never mistaken for success). The detached run survives
867 even if your watcher gets reaped, so re-checking the log always works.
868- Why the lock: a shared dev box with several Conductor worktrees will rate-limit
869 the model API if two eval suites run at once (15-way concurrency each), which
870 mass-times-out E2E tests. The lock makes the second run WAIT, not collide.
871- Humans running `bun run test:evals` foreground in their own terminal don't need
872 this — Ctrl-C is intended there. Detachment is for agent-launched runs only.
873 
874## E2E test fixtures: extract, don't copy
875 
876**NEVER copy a full SKILL.md file into an E2E test fixture.** SKILL.md files are
8771500-2000 lines. When `claude -p` reads a file that large, context bloat causes
878timeouts, flaky turn limits, and tests that take 5-10x longer than necessary.
879 
880Instead, extract only the section the test actually needs:
881 
882```typescript
883// BAD — agent reads 1900 lines, burns tokens on irrelevant sections
884fs.copyFileSync(path.join(ROOT, 'ship', 'SKILL.md'), path.join(dir, 'ship-SKILL.md'));
885 
886// GOOD — agent reads ~60 lines, finishes in 38s instead of timing out
887const full = fs.readFileSync(path.join(ROOT, 'ship', 'SKILL.md'), 'utf-8');
888const start = full.indexOf('## Review Readiness Dashboard');
889const end = full.indexOf('\n---\n', start);
890fs.writeFileSync(path.join(dir, 'ship-SKILL.md'), full.slice(start, end > start ? end : undefined));
891```
892 
893Also when running targeted E2E tests to debug failures:
894- Run in **foreground** (`bun test ...`), not background with `&` and `tee`
895- Never `pkill` running eval processes and restart — you lose results and waste money
896- One clean run beats three killed-and-restarted runs
897 
898## Publishing native OpenClaw skills to ClawHub
899 
900Native OpenClaw skills live in `openclaw/skills/gstack-openclaw-*/SKILL.md`. These are
901hand-crafted methodology skills (not generated by the pipeline) published to ClawHub
902so any OpenClaw user can install them.
903 
904**Publishing:** The command is `clawhub publish` (NOT `clawhub skill publish`):
905 
906```bash
907clawhub publish openclaw/skills/gstack-openclaw-office-hours \
908 --slug gstack-openclaw-office-hours --name "gstack Office Hours" \
909 --version 1.0.0 --changelog "description of changes"
910```
911 
912Repeat for each skill: `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`,
913`gstack-openclaw-retro`. Bump `--version` on each update.
914 
915**Auth:** `clawhub login` (opens browser for GitHub auth). `clawhub whoami` to verify.
916 
917**Updating:** Same `clawhub publish` command with a higher `--version` and `--changelog`.
918 
919**Verification:** `clawhub search gstack` to confirm they're live.
920 
921## Deploying to the active skill
922 
923The active skill lives at `~/.claude/skills/gstack/`. After making changes:
924 
9251. Push your branch
9262. Fetch and reset in the skill directory: `cd ~/.claude/skills/gstack && git fetch origin && git reset --hard origin/main`
9273. Rebuild: `cd ~/.claude/skills/gstack && bun run build`
928 
929**If you use gbrain:** the `git reset --hard` in step 2 reverts the brain-aware
930(`GBRAIN_CONTEXT_LOAD` / `GBRAIN_SAVE_RESULTS`) blocks that `gstack-config
931gbrain-refresh` renders into the install (those generated blocks differ from
932`main` by design). After deploying, re-run `gstack-config gbrain-refresh` to
933restore them across all your projects' Claude sessions. It's idempotent.
934 
935Or copy the binaries directly:
936- `cp browse/dist/browse ~/.claude/skills/gstack/browse/dist/browse`
937- `cp design/dist/design ~/.claude/skills/gstack/design/dist/design`
938 
939## Skill routing
940 
941When the user's request matches an available skill, invoke it via the Skill tool. When in doubt, invoke the skill.
942 
943Key routing rules:
944- Product ideas/brainstorming → invoke /office-hours
945- Strategy/scope → invoke /plan-ceo-review
946- Architecture → invoke /plan-eng-review
947- Design system/plan review → invoke /design-consultation or /plan-design-review
948- Full review pipeline → invoke /autoplan
949- Bugs/errors → invoke /investigate
950- QA/testing site behavior → invoke /qa or /qa-only
951- Code review/diff check → invoke /review
952- Visual polish → invoke /design-review
953- Ship/deploy/PR → invoke /ship or /land-and-deploy
954- Save progress → invoke /context-save
955- Resume context → invoke /context-restore
956 
957## Cross-session decision memory
958 
959Durable decisions and their rationale are captured in an append-only, event-sourced
960store at `~/.gstack/projects/<slug>/decisions.jsonl` so neither you nor the user
961re-litigates a settled call or loses the "why" across sessions. This is the reliable,
962file-only path: it works with gbrain OFF. (gbrain semantic recall is an optional
963enhancement layered on top, never a dependency.)
964 
965- **Resurface** active decisions before re-deciding: `bin/gstack-decision-search`
966 (`--recent N`, `--scope repo|branch|issue`, `--query KW`, `--all`, `--json`).
967 Add `--semantic` (with `--query`) to append related hits from gbrain memory when
968 it's up; it degrades silently to the reliable file results when gbrain is off.
969 Session start already surfaces scope-relevant active decisions via Context Recovery.
970 If a decision is listed, treat it as settled with its rationale; if you're about to
971 reverse it, say so explicitly.
972- **Capture** a DURABLE decision when you or the user make one:
973 `bin/gstack-decision-log '{"decision":"...","rationale":"...","scope":"repo|branch|issue","source":"user|skill|agent","confidence":1-10}'`.
974 Reverse a prior call with `--supersede <id>`; expunge an accidental secret with
975 `--redact <id>`; rewrite the log to the active set with `--compact`. Non-interactive
976 (never prompts), injection-sanitized, and HIGH-secret-blocking on write.
977- **Durable means:** architecture choice, scope cut, tool/vendor choice, or a reversal
978 of a prior call. NOT a turn-level edit, a phrasing tweak, or anything trivially
979 re-derivable. Capture is curated at the source — log durable decisions only, or the
980 store becomes noise.
981 
982## GBrain Search Guidance (configured by /sync-gbrain)
983<!-- gstack-gbrain-search-guidance:start -->
984 
985GBrain is set up and synced on this machine. The agent should prefer gbrain
986over Grep when the question is semantic or when you don't know the exact
987identifier yet.
988 
989**This worktree is pinned to a worktree-scoped code source** via the
990`.gbrain-source` file in the repo root (kubectl-style context). Any
991`gbrain code-def`, `code-refs`, `code-callers`, `code-callees`, or `query`
992call from anywhere under this worktree routes to that source by default —
993no `--source` flag needed. Conductor sibling worktrees of the same repo
994each have their own pin and their own indexed pages, so semantic results
995match the actual code on disk in this worktree.
996 
997Two indexed corpora available via the `gbrain` CLI:
998- This worktree's code (auto-pinned via `.gbrain-source`).
999- `~/.gstack/` curated memory (registered as `gstack-brain-<user>` source via
1000 the existing federation pipeline).
1001 
1002Prefer gbrain when:
1003- "Where is X handled?" / semantic intent, no exact string yet:
1004 `gbrain search "<terms>"` or `gbrain query "<question>"`
1005- "Where is symbol Y defined?" / symbol-based code questions:
1006 `gbrain code-def <symbol>` or `gbrain code-refs <symbol>`
1007- "What calls Y?" / "What does Y depend on?":
1008 `gbrain code-callers <symbol>` / `gbrain code-callees <symbol>`
1009- "What did we decide last time?" / past plans, retros, learnings:
1010 `gbrain search "<terms>" --source gstack-brain-<user>`
1011 
1012Grep is still right for known exact strings, regex, multiline patterns, and
1013file globs. Run `/sync-gbrain` after meaningful code changes; for ongoing
1014auto-sync across all worktrees, run `gbrain autopilot --install` once per
1015machine — gbrain's daemon handles incremental refresh on a schedule.
1016 
1017Safety: don't run `/sync-gbrain` while `gbrain autopilot` is active — the
1018orchestrator refuses destructive source ops when it detects a running autopilot
1019to avoid racing it (#1734). Prefer registering user repos with `gbrain sources
1020add --path <dir>` (no `--url`): URL-managed sources can auto-reclone, and the
1021sync code walk for them requires an explicit `--allow-reclone` opt-in.
1022 
1023<!-- gstack-gbrain-search-guidance:end -->
1024 
@@ −1 +1 @@
1−# gstack — AI Engineering Workflow
1+# gstack development
22  
3−gstack is a collection of SKILL.md files that give AI agents structured roles for
4−software development. Each skill is a specialist: CEO reviewer, eng manager,
5−designer, QA lead, release engineer, debugger, and more.
3+## Commands
64  
7−## Available skills
5+```bash
6+bun install # install dependencies
7+bun test # run free tests (browse + snapshot + skill validation)
8+bun run test:evals # run paid evals: LLM judge + E2E (diff-based, ~$4/run max)
9+bun run test:evals:all # run ALL paid evals regardless of diff
10+bun run test:gate # run gate-tier tests only (CI default, blocks merge)
11+bun run test:periodic # run periodic-tier tests only (weekly cron / manual)
12+bun run test:e2e # run E2E tests only (diff-based, ~$3.85/run max)
13+bun run test:e2e:all # run ALL E2E tests regardless of diff
14+bun run eval:select # show which tests would run based on current diff
15+bun run dev <cmd> # run CLI in dev mode, e.g. bun run dev goto https://example.com
16+bun run build # gen docs + compile binaries
17+bun run gen:skill-docs # regenerate SKILL.md files from templates
18+bun run skill:check # health dashboard for all skills
19+bun run dev:skill # watch mode: auto-regen + validate on change
20+bun run eval:list # list all eval runs from ~/.gstack-dev/evals/
21+bun run eval:compare # compare two eval runs (auto-picks most recent)
22+bun run eval:summary # aggregate stats across all eval runs
23+bun run slop # full slop-scan report (all files)
24+bun run slop:diff # slop findings in files changed on this branch only
25+```
826  
9−Skills live in `.agents/skills/` (or `~/.claude/skills/gstack/` on Claude Code).
10−Invoke them by name (e.g., `/office-hours`).
27+`test:evals` requires `ANTHROPIC_API_KEY`. Codex E2E tests (`test/codex-e2e.test.ts`)
28+use Codex's own auth from `~/.codex/` config — no `OPENAI_API_KEY` env var needed.
1129  
12−### Plan-mode reviews
30+**Env keys in Conductor workspaces.** The `GSTACK_*` env-shim (v1.39.2.0+,
31+`lib/conductor-env-shim.ts`) promotes `GSTACK_ANTHROPIC_API_KEY` /
32+`GSTACK_OPENAI_API_KEY` to their canonical names inside gstack's TS binaries.
33+Tests run through gstack entrypoints inherit this promotion automatically.
34+Don't echo the key value to stdout, logs, or shell history. The historical
35+"never pass `env:` to `runAgentSdkTest`" rule is retired: the failure was
36+partial-env replacement (the SDK's `Options.env` REPLACES the child's entire
37+environment, so an object without the key broke auth). The runner now always
38+passes a COMPLETE hermetic env with per-test `env:` merged last, so per-test
39+overrides are safe; ambient `process.env.ANTHROPIC_API_KEY` mutation also
40+still works (the env builder reads process.env at call time).
1341  
14−| Skill | What it does |
15−|-------|-------------|
16−| `/office-hours` | Start here. Reframes your product idea before you write code. |
17−| `/plan-ceo-review` | CEO-level review: find the 10-star product in the request. |
18−| `/plan-eng-review` | Lock architecture, data flow, edge cases, and tests. |
19−| `/plan-design-review` | Rate each design dimension 0-10, explain what a 10 looks like. |
20−| `/plan-devex-review` | DX-mode review: TTHW, magical moments, friction points, persona traces. |
21−| `/plan-tune` | Self-tune AskUserQuestion sensitivity per question. |
22−| `/autoplan` | One command runs CEO → design → eng → DX review. |
23−| `/design-consultation` | Build a complete design system from scratch. |
24−| `/spec` | Turn vague intent into a precise, executable spec in five phases. Files a GitHub issue, optionally spawns a Claude Code agent in a fresh worktree, and lets `/ship` close the source issue on merge. |
42+**Hermetic local E2E (default).** Every E2E runner (claude -p, PTY, Agent
43+SDK, codex, gemini) spawns children through `test/helpers/hermetic-env.ts`:
44+allowlist-scrubbed env (operator `CONDUCTOR_*`, `CLAUDE_*`, `GSTACK_*`,
45+`MCP_*`, `GBRAIN_*`, and credentials like `GH_TOKEN` never reach children),
46+a fresh seeded `CLAUDE_CONFIG_DIR` (no operator `~/.claude` CLAUDE.md /
47+MCP servers / skills), a temp `GSTACK_HOME`, and `--strict-mcp-config`.
48+Local eval signal matches CI. Debug against real operator state with
49+`EVALS_HERMETIC=0` (restores the legacy env AND drops the strict-MCP flag).
50+Per-test `env:` overrides merge last, so deliberate contamination
51+(`CONDUCTOR_WORKSPACE_PATH`, per-test `GSTACK_HOME`) keeps working. Wiring
52+is pinned by `test/hermetic-wiring.test.ts` (static tripwire) and two
53+gate-tier canaries in `test/skill-e2e-hermetic-canary.test.ts`.
2554  
26−### Implementation + review
55+E2E tests stream progress in real-time (tool-by-tool via `--output-format stream-json
56+--verbose`). Results are persisted to `~/.gstack-dev/evals/` with auto-comparison
57+against the previous run.
2758  
28−| Skill | What it does |
29−|-------|-------------|
30−| `/review` | Pre-landing PR review. Finds bugs that pass CI but break in prod. |
31−| `/codex` | Second opinion via OpenAI Codex. Review, challenge, or consult modes. |
32−| `/investigate` | Systematic root-cause debugging. No fixes without investigation. |
33−| `/design-review` | Live-site visual audit + fix loop with atomic commits. |
34−| `/design-shotgun` | Generate multiple AI design variants, comparison board, iterate. |
35−| `/design-html` | Generate production-quality Pretext-native HTML/CSS. |
36−| `/devex-review` | Live developer experience audit (TTHW measured against the real flow). |
37−| `/qa` | Open a real browser, find bugs, fix them, re-verify. |
38−| `/qa-only` | Same methodology as /qa but report only — no code changes. |
39−| `/scrape` | Pull data from a web page. First call prototypes; codified call runs in ~200ms. |
40−| `/skillify` | Codify the most recent successful `/scrape` flow into a permanent browser-skill. |
59+**Diff-based test selection:** `test:evals` and `test:e2e` auto-select tests based
60+on `git diff` against the base branch. Each test declares its file dependencies in
61+`test/helpers/touchfiles.ts`. Changes to global touchfiles (session-runner, eval-store,
62+touchfiles.ts itself) trigger all tests. Use `EVALS_ALL=1` or the `:all` script
63+variants to force all tests. Run `eval:select` to preview which tests would run.
4164  
42−### Release + deploy
65+**Two-tier system:** Tests are classified as `gate` or `periodic` in `E2E_TIERS`
66+(in `test/helpers/touchfiles.ts`). CI runs only gate tests (`EVALS_TIER=gate`);
67+periodic tests run weekly via cron or manually. Use `EVALS_TIER=gate` or
68+`EVALS_TIER=periodic` to filter. When adding new E2E tests, classify them:
69+1. Safety guardrail or deterministic functional test? -> `gate`
70+2. Quality benchmark, Opus model test, or non-deterministic? -> `periodic`
71+3. Requires external service (Codex, Gemini)? -> `periodic`
4372  
44−| Skill | What it does |
45−|-------|-------------|
46−| `/ship` | Run tests, review, push, open PR. Workspace-aware version queue. |
47−| `/land-and-deploy` | Merge the PR, wait for CI and deploy, verify production health. |
48−| `/canary` | Post-deploy monitoring loop using the browse daemon. |
49−| `/landing-report` | Read-only dashboard for the workspace-aware ship queue. |
50−| `/document-release` | Update all docs to match what you just shipped. |
51−| `/document-generate` | Generate Diataxis docs (tutorial / how-to / reference / explanation) from code. |
52−| `/setup-deploy` | One-time deploy config detection (Fly.io, Render, Vercel, etc.). |
53−| `/gstack-upgrade` | Update gstack to the latest version. |
73+## Testing
5474  
55−### Operational + memory
75+```bash
76+bun test # run before every commit — free, <2s
77+bun run test:evals # run before shipping — paid, diff-based (~$4/run max)
78+```
5679  
57−| Skill | What it does |
58−|-------|-------------|
59−| `/context-save` | Save working context (git state, decisions, remaining work). |
60−| `/context-restore` | Resume from a saved context, even across Conductor workspaces. |
61−| `/learn` | Manage what gstack learned across sessions. |
62−| `/retro` | Weekly retro with per-person breakdowns and shipping streaks. |
63−| `/health` | Code quality dashboard (type checker, linter, tests, dead code). |
64−| `/benchmark` | Performance regression detection (page load, Core Web Vitals). |
65−| `/benchmark-models` | Cross-model benchmark for skills (Claude, GPT, Gemini side-by-side). |
66−| `/cso` | OWASP Top 10 + STRIDE security audit. |
67−| `/setup-gbrain` | Set up gbrain for cross-machine session memory sync. |
68−| `/sync-gbrain` | Keep gbrain current with this repo's code; refresh agent search guidance in CLAUDE.md. |
80+`bun test` runs skill validation, gen-skill-docs quality checks, and browse
81+integration tests. `bun run test:evals` runs LLM-judge quality evals and E2E
82+tests via `claude -p`. Both must pass before creating a PR.
6983  
70−### Browser + agent integration
84+## Project structure
7185  
72−| Skill | What it does |
73−|-------|-------------|
74−| `/browse` | Headless browser — real Chromium, real clicks, ~100ms/command. |
75−| `/open-gstack-browser` | Launch the visible GStack Browser with sidebar + stealth. |
76−| `/setup-browser-cookies` | Import cookies from your real browser for authenticated testing. |
77−| `/pair-agent` | Pair a remote AI agent (OpenClaw, Codex, etc.) with your browser. |
86+```
87+gstack/
88+├── browse/ # Headless browser CLI (Playwright)
89+│ ├── src/ # CLI + server + commands
90+│ │ ├── commands.ts # Command registry (single source of truth)
91+│ │ └── snapshot.ts # SNAPSHOT_FLAGS metadata array
92+│ ├── test/ # Integration tests + fixtures
93+│ └── dist/ # Compiled binary
94+├── hosts/ # Typed host configs (one per AI agent)
95+│ ├── claude.ts # Primary host config
96+│ ├── codex.ts, factory.ts, kiro.ts # Existing hosts
97+│ ├── opencode.ts, slate.ts, cursor.ts, openclaw.ts # IDE hosts
98+│ ├── hermes.ts, gbrain.ts # Agent runtime hosts
99+│ └── index.ts # Registry: exports all, derives Host type
100+├── scripts/ # Build + DX tooling
101+│ ├── gen-skill-docs.ts # Template → SKILL.md generator (config-driven)
102+│ ├── host-config.ts # HostConfig interface + validator
103+│ ├── host-config-export.ts # Shell bridge for setup script
104+│ ├── host-adapters/ # Host-specific adapters (OpenClaw tool mapping)
105+│ ├── resolvers/ # Template resolver modules (preamble, design, review, gbrain, etc.)
106+│ ├── skill-check.ts # Health dashboard
107+│ └── dev-skill.ts # Watch mode
108+├── test/ # Skill validation + eval tests
109+│ ├── helpers/ # skill-parser.ts, session-runner.ts, llm-judge.ts, eval-store.ts
110+│ ├── fixtures/ # Ground truth JSON, planted-bug fixtures, eval baselines
111+│ ├── skill-validation.test.ts # Tier 1: static validation (free, <1s)
112+│ ├── gen-skill-docs.test.ts # Tier 1: generator quality (free, <1s)
113+│ ├── skill-llm-eval.test.ts # Tier 3: LLM-as-judge (~$0.15/run)
114+│ └── skill-e2e-*.test.ts # Tier 2: E2E via claude -p (~$3.85/run, split by category)
115+├── qa-only/ # /qa-only skill (report-only QA, no fixes)
116+├── plan-design-review/ # /plan-design-review skill (report-only design audit)
117+├── design-review/ # /design-review skill (design audit + fix loop)
118+├── ship/ # Ship workflow skill
119+├── review/ # PR review skill
120+├── plan-ceo-review/ # /plan-ceo-review skill
121+├── plan-eng-review/ # /plan-eng-review skill
122+├── autoplan/ # /autoplan skill (auto-review pipeline: CEO → design → eng)
123+├── benchmark/ # /benchmark skill (performance regression detection)
124+├── canary/ # /canary skill (post-deploy monitoring loop)
125+├── codex/ # /codex skill (multi-AI second opinion via OpenAI Codex CLI)
126+├── land-and-deploy/ # /land-and-deploy skill (merge → deploy → canary verify)
127+├── office-hours/ # /office-hours skill (YC Office Hours — startup diagnostic + builder brainstorm)
128+├── investigate/ # /investigate skill (systematic root-cause debugging)
129+├── spec/ # /spec skill (five-phase spec → GitHub issue, optional agent spawn, /ship auto-closes)
130+├── retro/ # Retrospective skill (includes /retro global cross-project mode)
131+├── bin/ # CLI utilities (gstack-repo-mode, gstack-slug, gstack-config, etc.)
132+├── document-release/ # /document-release skill (post-ship doc updates + Diataxis coverage map)
133+├── document-generate/ # /document-generate skill (Diataxis doc generator: tutorial/how-to/reference/explanation)
134+├── cso/ # /cso skill (OWASP Top 10 + STRIDE security audit)
135+├── design-consultation/ # /design-consultation skill (design system from scratch)
136+├── design-shotgun/ # /design-shotgun skill (visual design exploration)
137+├── open-gstack-browser/ # /open-gstack-browser skill (launch GStack Browser)
138+├── connect-chrome/ # symlink → open-gstack-browser (backwards compat)
139+├── design/ # Design binary CLI (GPT Image API)
140+│ ├── src/ # CLI + commands (generate, variants, compare, serve, etc.)
141+│ ├── test/ # Integration tests
142+│ └── dist/ # Compiled binary
143+├── extension/ # Chrome extension (side panel + activity feed + CSS inspector)
144+├── lib/ # Shared libraries (worktree.ts)
145+├── docs/designs/ # Design documents
146+├── setup-deploy/ # /setup-deploy skill (one-time deploy config)
147+├── .github/ # CI workflows + Docker image
148+│ ├── workflows/ # evals.yml (E2E on Ubicloud), skill-docs.yml, actionlint.yml
149+│ └── docker/ # Dockerfile.ci (pre-baked toolchain + Playwright/Chromium)
150+├── contrib/ # Contributor-only tools (never installed for users)
151+│ └── add-host/ # /gstack-contrib-add-host skill
152+├── setup # One-time setup: build binary + symlink skills
153+├── SKILL.md # Generated from SKILL.md.tmpl (don't edit directly)
154+├── SKILL.md.tmpl # Template: edit this, run gen:skill-docs
155+├── ETHOS.md # Builder philosophy (Boil the Ocean, Search Before Building)
156+└── package.json # Build scripts for browse
157+```
78158  
79−### iOS QA — drive real iPhones over USB or Tailscale (v1.43.0.0+)
159+## SKILL.md workflow
80160  
81−| Skill | What it does |
82−|-------|-------------|
83−| `/ios-qa` | Live-device iOS QA via USB CoreDevice tunnel + embedded StateServer. Optionally exposes the device over Tailscale so remote agents can drive it. |
84−| `/ios-fix` | Autonomous iOS bug fixer with regression snapshot capture. |
85−| `/ios-design-review` | Designer's-eye QA on a real iPhone — 10-dimension Apple HIG rubric. |
86−| `/ios-clean` | Convenience: strip DebugBridge + #if DEBUG wiring before a Release build. |
87−| `/ios-sync` | Regenerate the iOS debug bridge against the latest upstream templates. |
161+SKILL.md files are **generated** from `.tmpl` templates. To update docs:
88162  
89−Companion CLIs (run on the Mac that's plugged into the device):
163+1. Edit the `.tmpl` file (e.g. `SKILL.md.tmpl` or `browse/SKILL.md.tmpl`)
164+2. Run `bun run gen:skill-docs` (or `bun run build` which does it automatically)
165+3. Commit both the `.tmpl` and generated `.md` files
90166  
91−| Command | What it does |
92−|---------|-------------|
93−| `gstack-ios-qa-daemon` | Mac-side broker. Loopback by default; `--tailnet` adds a Tailscale-facing listener with capability tiers and audit logging. |
94−| `gstack-ios-qa-mint` | Owner-grant CLI for the tailnet allowlist (`grant`/`revoke`/`list`). |
95−| `gstack-ios-qa-regen` | Regenerate the canonical local DebugBridge package and typed accessors (`--app-source` / `--bridge-dir`). |
167+To add a new browse command: add it to `browse/src/commands.ts` and rebuild.
168+To add a snapshot flag: add it to `SNAPSHOT_FLAGS` in `browse/src/snapshot.ts` and rebuild.
96169  
97−End-to-end walkthrough: [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-testing-with-gstack.md).
170+**Token ceiling:** Generated SKILL.md files trip a warning above 160KB (~40K tokens).
171+This is a "watch for feature bloat" guardrail, not a hard gate. Modern flagship
172+models have 200K-1M context windows, so 40K is 4-20% of window, and prompt caching
173+makes the marginal cost of larger skills small. The ceiling exists to catch runaway
174+preamble/resolver growth, not to force compression on carefully-tuned big skills
175+(`ship`, `plan-ceo-review`, `office-hours` legitimately pack 25-35K tokens of
176+behavior). If you blow past 40K, the right fix is usually: (1) look at WHAT grew,
177+(2) if one resolver added 10K+ in a single PR, question whether it belongs inline
178+or as a reference doc, (3) only compress carefully-tuned prose as a last resort —
179+cuts to the coverage audit, review army, or voice directive have real quality cost.
98180  
99−### Safety + scoping
181+**Merge conflicts on SKILL.md files:** NEVER resolve conflicts on generated SKILL.md
182+files by accepting either side. Instead: (1) resolve conflicts on the `.tmpl` templates
183+and `scripts/gen-skill-docs.ts` (the sources of truth), (2) run `bun run gen:skill-docs`
184+to regenerate all SKILL.md files, (3) stage the regenerated files. Accepting one side's
185+generated output silently drops the other side's template changes.
100186  
101−| Skill | What it does |
102−|-------|-------------|
103−| `/careful` | Warn before destructive commands (rm -rf, DROP TABLE, force-push). |
104−| `/freeze` | Lock edits to one directory. Hard block, not just a warning. |
105−| `/guard` | Activate both careful + freeze at once. |
106−| `/unfreeze` | Remove directory edit restrictions. |
107−| `/make-pdf` | Turn any markdown file into a publication-quality PDF. |
108−| `/diagram` | English in, diagram out: mermaid source + editable .excalidraw + SVG/PNG, offline. |
187+## Platform-agnostic design
109188  
110−## Build commands
189+Skills must NEVER hardcode framework-specific commands, file patterns, or directory
190+structures. Instead:
111191  
192+1. **Read CLAUDE.md** for project-specific config (test commands, eval commands, etc.)
193+2. **If missing, AskUserQuestion** — let the user tell you or let gstack search the repo
194+3. **Persist the answer to CLAUDE.md** so we never have to ask again
195+ 
196+This applies to test commands, eval commands, deploy commands, and any other
197+project-specific behavior. The project owns its config; gstack reads it.
198+ 
199+## Writing SKILL templates
200+ 
201+SKILL.md.tmpl files are **prompt templates read by Claude**, not bash scripts.
202+Each bash code block runs in a separate shell — variables do not persist between blocks.
203+ 
204+Rules:
205+- **Use natural language for logic and state.** Don't use shell variables to pass
206+ state between code blocks. Instead, tell Claude what to remember and reference
207+ it in prose (e.g., "the base branch detected in Step 0").
208+- **Don't hardcode branch names.** Detect `main`/`master`/etc dynamically via
209+ `gh pr view` or `gh repo view`. Use `{{BASE_BRANCH_DETECT}}` for PR-targeting
210+ skills. Use "the base branch" in prose, `<base>` in code block placeholders.
211+- **Keep bash blocks self-contained.** Each code block should work independently.
212+ If a block needs context from a previous step, restate it in the prose above.
213+- **Express conditionals as English.** Instead of nested `if/elif/else` in bash,
214+ write numbered decision steps: "1. If X, do Y. 2. Otherwise, do Z."
215+ 
216+## Writing style (V1)
217+ 
218+Default output from every tier-≥2 skill follows the Writing Style section in
219+`scripts/resolvers/preamble.ts`: jargon glossed on first use (curated list in
220+`scripts/jargon-list.json`, baked at gen-skill-docs time), questions framed in
221+outcome terms ("what breaks for your users if...") not implementation terms,
222+short sentences, decisions close with user impact. Power users who want the
223+tighter V0 prose set `gstack-config set explain_level terse` (binary switch,
224+no middle mode). See `docs/designs/PLAN_TUNING_V1.md` for the full design
225+rationale. The review pacing overhaul that originally tried to ride alongside
226+writing-style was extracted to V1.1 — see `docs/designs/PACING_UPDATES_V0.md`.
227+ 
228+## Browser interaction
229+ 
230+When you need to interact with a browser (QA, dogfooding, cookie setup), use the
231+`/browse` skill or run the browse binary directly via `$B <command>`. NEVER use
232+`mcp__claude-in-chrome__*` tools — they are slow, unreliable, and not what this
233+project uses.
234+ 
235+**Sidebar architecture:** Before modifying `sidepanel.js`, `background.js`,
236+`content.js`, `terminal-agent.ts`, or sidebar-related server endpoints,
237+read `docs/designs/SIDEBAR_MESSAGE_FLOW.md`. The sidebar has one primary
238+surface — the **Terminal** pane (interactive `claude` PTY) — with
239+Activity / Refs / Inspector as debug overlays behind the footer's
240+`debug` toggle. The chat queue path was ripped once the PTY proved out;
241+`sidebar-agent.ts` and the `/sidebar-command` / `/sidebar-chat` /
242+`/sidebar-agent/event` endpoints are gone. The doc covers the WS auth
243+flow, dual-token model, and threat-model boundary — silent failures
244+here usually trace to not understanding the cross-component flow.
245+ 
246+**Embedder terminal-agent ownership** (v1.42.1.0+, identity-based kill v1.44.0.0+).
247+`buildFetchHandler` in `browse/src/server.ts` accepts `ServerConfig.ownsTerminalAgent?:
248+boolean` (default `true`). When `true`, factory shutdown runs the full teardown:
249+identity-based kill via `killAgentByRecord(readAgentRecord(stateDir))` from
250+`browse/src/terminal-agent-control.ts` plus `safeUnlinkQuiet` on
251+`<stateDir>/terminal-port`, `<stateDir>/terminal-internal-token`, and
252+`<stateDir>/terminal-agent-pid` (the per-boot agent record introduced in v1.44).
253+Embedders (e.g. the gbrowser phoenix overlay) that pre-launch their own PTY
254+server must pass `false` so their discovery files survive gstack teardown cycles.
255+The flag is the third caller-owned teardown gate in `ServerConfig` (alongside
256+`xvfb?` and `proxyBridge?`); polarity is inverted (explicit bool vs presence) and
257+documented in the field's JSDoc. CLI `start()` always passes `true` explicitly —
258+the static-grep test in `browse/test/server-embedder-terminal-port.test.ts` fails
259+CI if a refactor drops it. Pre-v1.44 used `pkill -f terminal-agent\.ts` (regex
260+match) which would kill sibling gstack sessions on the same host; the new
261+`browse/test/terminal-agent-pid-identity.test.ts` static-grep tripwire fails CI
262+if any source file re-introduces `pkill ... terminal-agent` or `spawnSync('pkill', ...)`.
263+ 
264+**WebSocket auth uses Sec-WebSocket-Protocol, not cookies.** Browsers
265+can't set `Authorization` on a WebSocket upgrade, but they CAN set
266+`Sec-WebSocket-Protocol` via `new WebSocket(url, [token])`. The agent
267+reads it, validates against `validTokens`, and MUST echo the protocol
268+back in the upgrade response — without the echo, Chromium closes the
269+connection immediately. `Set-Cookie: gstack_pty=...` is kept as a
270+fallback for non-browser callers (the cross-port `SameSite=Strict`
271+cookie path doesn't survive from a chrome-extension origin).
272+ 
273+**Cross-pane PTY injection.** The toolbar's Cleanup button and the
274+Inspector's "Send to Code" action both pipe text into the live claude
275+PTY via `window.gstackInjectToTerminal(text)`, exposed by
276+`sidepanel-terminal.js`. No `/sidebar-command` POST — the live REPL is
277+the only execution surface in the sidebar now.
278+ 
279+**`/health` MUST NOT surface any shell-grant token.** It already leaks
280+`AUTH_TOKEN` to localhost callers in headed mode (a v1.1+ TODO). Don't
281+make that worse by adding the PTY session token there. PTY auth flows
282+through `POST /pty-session` only.
283+ 
284+**Transport-layer security** (v1.6.0.0+). When `pair-agent` starts an ngrok tunnel,
285+the daemon binds two HTTP listeners: a local listener (127.0.0.1, full command
286+surface, never forwarded) and a tunnel listener (locked allowlist: `/connect`,
287+`/command` with a scoped token + 26-command browser-driving allowlist,
288+`/sidebar-chat`). ngrok forwards only the tunnel port. Root tokens over the tunnel
289+return 403. SSE endpoints use a 30-minute HttpOnly `gstack_sse` cookie minted via
290+`POST /sse-session` (never valid against `/command`). Tunnel-surface rejections go
291+to `~/.gstack/security/attempts.jsonl` via `tunnel-denial-log.ts`. Before editing
292+`server.ts`, `sse-session-cookie.ts`, or `tunnel-denial-log.ts`, read
293+[ARCHITECTURE.md](ARCHITECTURE.md#dual-listener-tunnel-architecture-v1600) —
294+the module boundary (no imports from `token-registry.ts` into `sse-session-cookie.ts`)
295+is load-bearing for scope isolation.
296+ 
297+**Unicode sanitization at server egress** (v1.38.0.0+). Every server egress that
298+ships page-content-derived strings MUST go through `JSON.stringify(payload,
299+sanitizeReplacer)` for object payloads or `sanitizeLoneSurrogates(body)` for text
300+bodies. Lone UTF-16 surrogate halves from CDP page content otherwise reach the
301+Anthropic API as `\uD800`-style escapes and trigger a 400. Wired at four egress
302+points today: `handleCommandInternal` (HTTP + batch via a sanitizing wrapper around
303+`handleCommandInternalImpl`) and both SSE producers (`/activity/stream`,
304+`/inspector/events`). Post-stringify regex is a no-op — `JSON.stringify` has
305+already escaped the surrogate before regex could match, so the replacer must run
306+inside the encoding pipeline. Before adding a new SSE/WebSocket writer or HTTP
307+response in `server.ts`, read
308+[ARCHITECTURE.md](ARCHITECTURE.md#unicode-sanitization-at-server-egress-v13800).
309+`browse/test/server-sanitize-surrogates.test.ts` pins the wiring with invariant
310+tests, so bypasses fail CI.
311+ 
312+**SSE endpoint helper** (v1.51.0.0+). New SSE endpoints in `server.ts` MUST route
313+through `createSseEndpoint(req, config)` from `browse/src/sse-helpers.ts`. The
314+helper owns the cleanup contract (abort + enqueue-throw + heartbeat-throw, all
315+idempotent) and bakes in `sanitizeLoneSurrogates` on every JSON.stringify, so
316+new subscribers can't accidentally regress either invariant. Inline
317+`ReadableStream` wiring leaked subscribers when the TCP connection died without
318+firing `req.signal.abort` (Chromium MV3 service-worker suspend, intermediate
319+proxy half-close). `/activity/stream`, `/inspector/events`, and `/memory`
320+(SSE-eligible) all route through it. `browse/test/sse-helpers.test.ts` pins the
321+cleanup contract.
322+ 
323+**CDP session lifecycle** (v1.51.0.0+). Direct `page.context().newCDPSession(page)`
324+calls outside `browse/src/cdp-bridge.ts` fail CI via the static-grep tripwire in
325+`browse/test/cdp-session-cleanup.test.ts`. Use `withCdpSession(page, async (s) => {...})`
326+for one-shot CDP work (try/finally detach) or `getOrCreateCdpSession(page, cache)`
327+for cached sessions tied to a page's lifetime (close-detach via `Map<page, session>`).
328+Three sites migrated: cdp-bridge frame events, write-commands archive capture,
329+cdp-inspector. The helpers prevent the per-session leak class where successful-path
330+detach happened but error-path detach was missed.
331+ 
332+**Setup symlink hardening** (v1.38.0.0+). Every link site in `setup` MUST route
333+through the `_link_or_copy SRC DST` helper near the `IS_WINDOWS` detection. On
334+Windows without Developer Mode, plain `ln -snf` produces frozen file copies that
335+don't refresh on `git pull` — silent staleness across every host adapter. The
336+helper preserves `ln -snf` on Unix and switches to `cp -R` / `cp -f` on Windows.
337+`test/setup-windows-fallback.test.ts` enforces a static invariant: a single raw
338+`ln` call outside the helper body fails CI. Windows users get a one-line note
339+from `_print_windows_copy_note_once` reminding them to re-run `./setup` after
340+every `git pull`.
341+ 
342+**Sidebar security stack** (layered defense against prompt injection):
343+ 
344+| Layer | Module | Lives in |
345+|-------|--------|----------|
346+| L1-L3 | `content-security.ts` | both server and agent — datamarking, hidden element strip, ARIA regex, URL blocklist, envelope wrapping |
347+| L4 | `security-classifier.ts` (TestSavantAI ONNX) | **sidebar-agent only** |
348+| L4b | `security-classifier.ts` (Claude Haiku transcript) | **sidebar-agent only** |
349+| L5 | `security.ts` (canary) | both — inject in compiled, check in agent |
350+| L6 | `security.ts` (combineVerdict ensemble) | both |
351+ 
352+**Critical constraint:** `security-classifier.ts` CANNOT be imported from the
353+compiled browse binary. `@huggingface/transformers` v4 requires `onnxruntime-node`
354+which fails to `dlopen` from Bun compile's temp extract dir. Only `security.ts`
355+(pure-string operations — canary, verdict combiner, attack log, status) is safe
356+for `server.ts`. See `~/.gstack/projects/garrytan-gstack/ceo-plans/2026-04-19-prompt-injection-guard.md`
357+§"Pre-Impl Gate 1 Outcome" for full architectural decision.
358+ 
359+**Thresholds** (in `security.ts`):
360+- `BLOCK: 0.85` — single-layer score that would cause BLOCK if cross-confirmed
361+- `WARN: 0.75` — cross-confirm threshold. When L4 AND L4b both >= 0.75 → BLOCK
362+- `LOG_ONLY: 0.40` — gates transcript classifier (skip Haiku when all layers < 0.40)
363+- `SOLO_CONTENT_BLOCK: 0.92` — single-layer threshold for label-less content classifiers
364+ (testsavant, deberta). Intentionally higher than `BLOCK` because these layers can't
365+ distinguish "this is an injection" from "this looks like phishing aimed at the user."
366+ The transcript classifier keeps a separate, label-gated solo path at `BLOCK` (0.85).
367+ 
368+**Ensemble rule:** BLOCK only when the ML content classifier AND the transcript
369+classifier both report >= WARN. Single-layer high confidence degrades to WARN —
370+this is the Stack Overflow instruction-writing FP mitigation. Canary leak
371+always BLOCKs (deterministic).
372+ 
373+**Env knobs:**
374+- `GSTACK_SECURITY_OFF=1` — emergency kill switch. Classifier stays off even if
375+ warmed. Canary is still injected; just the ML scan is skipped.
376+- `GSTACK_SECURITY_ENSEMBLE=deberta` — opt-in DeBERTa-v3 ensemble. Adds
377+ ProtectAI DeBERTa-v3-base-injection-onnx as L4c classifier for cross-model
378+ agreement. 721MB first-run download. With ensemble enabled, BLOCK requires
379+ 2-of-3 ML classifiers agreeing at >= WARN (testsavant, deberta, transcript).
380+ Without ensemble (default), BLOCK requires testsavant + transcript at >= WARN.
381+- Classifier model cache: `~/.gstack/models/testsavant-small/` (112MB, first run only)
382+ plus `~/.gstack/models/deberta-v3-injection/` (721MB, only when ensemble enabled)
383+- Attack log: `~/.gstack/security/attempts.jsonl` (salted sha256 + domain only,
384+ rotates at 10MB, 5 generations)
385+- Per-device salt: `~/.gstack/security/device-salt` (0600)
386+- Session state: `~/.gstack/security/session-state.json` (cross-process, atomic)
387+ 
388+## Dev symlink awareness
389+ 
390+When developing gstack, `.claude/skills/gstack` may be a symlink back to this
391+working directory (gitignored). This means skill changes are **live immediately**,
392+great for rapid iteration, risky during big refactors where half-written skills
393+could break other Claude Code sessions using gstack concurrently.
394+ 
395+**Check once per session:** Run `ls -la .claude/skills/gstack` to see if it's a
396+symlink or a real copy. If it's a symlink to your working directory, be aware that:
397+- Template changes + `bun run gen:skill-docs` immediately affect all gstack invocations
398+- Breaking changes to SKILL.md.tmpl files can break concurrent gstack sessions
399+- During large refactors, remove the symlink (`rm .claude/skills/gstack`) so the
400+ global install at `~/.claude/skills/gstack/` is used instead
401+ 
402+**Prefix setting:** Setup creates real directories (not symlinks) at the top level
403+with a SKILL.md symlink inside (e.g., `qa/SKILL.md -> gstack/qa/SKILL.md`). This
404+ensures Claude discovers them as top-level skills, not nested under `gstack/`.
405+Names are either short (`qa`) or namespaced (`gstack-qa`), controlled by
406+`skill_prefix` in `~/.gstack/config.yaml`. Pass `--no-prefix` or `--prefix` to
407+skip the interactive prompt.
408+ 
409+**Note:** Vendoring gstack into a project's repo is deprecated. Use global install
410++ `./setup --team` instead. See README.md for team mode instructions.
411+ 
412+**For plan reviews:** When reviewing plans that modify skill templates or the
413+gen-skill-docs pipeline, consider whether the changes should be tested in isolation
414+before going live (especially if the user is actively using gstack in other windows).
415+ 
416+**Upgrade migrations:** When a change modifies on-disk state (directory structure,
417+config format, stale files) in ways that could break existing user installs, add a
418+migration script to `gstack-upgrade/migrations/`. Read CONTRIBUTING.md's "Upgrade
419+migrations" section for the format and testing requirements. The upgrade skill runs
420+these automatically after `./setup` during `/gstack-upgrade`.
421+ 
422+## Compiled binaries — NEVER commit browse/dist/ or design/dist/
423+ 
424+The `browse/dist/` and `design/dist/` directories contain compiled Bun binaries
425+(`browse`, `find-browse`, `design`, ~58MB each). These are Mach-O arm64 only — they
426+do NOT work on Linux, Windows, or Intel Macs. The `./setup` script already builds
427+from source for every platform, so the checked-in binaries are redundant. They are
428+tracked by git due to a historical mistake and should eventually be removed with
429+`git rm --cached`.
430+ 
431+**NEVER stage or commit these files.** They show up as modified in `git status`
432+because they're tracked despite `.gitignore` — ignore them. When staging files,
433+always use specific filenames (`git add file1 file2`) — never `git add .` or
434+`git add -A`, which will accidentally include the binaries.
435+ 
436+## Redaction guard (PII / secrets / legal content)
437+ 
438+Shared redaction engine catches credentials, PII, and legal/damaging content
439+before it reaches an external sink (codex dispatch, GitHub issue/PR body, pushed
440+commit). It is a **guardrail, not airtight enforcement** — `git push --no-verify`,
441+direct `gh issue create`, and `GSTACK_REDACT_PREPUSH=skip` all bypass it. It
442+catches accidents and carelessness, the 99% case. Do not claim it stops a
443+determined leaker (a CHANGELOG line that does would fail a hostile screenshotter).
444+ 
445+- **Engine + taxonomy:** `lib/redact-patterns.ts` (the single source of truth —
446+ 3 tiers; HIGH = genuinely-secret credentials that block, MEDIUM = PII/legal/
447+ internal + high-FP credential shapes that confirm via AskUserQuestion, LOW =
448+ FYI) and `lib/redact-engine.ts` (pure `scan()` + `applyRedactions()`).
449+ Calibration matters: a gate that cries wolf gets ignored, so context-variable
450+ shapes (Stripe `pk_live_`, Google `AIza`, JWT, env `*_KEY=`) sit at MEDIUM.
451+- **CLI:** `bin/gstack-redact` (exit 0 clean / 2 MEDIUM / 3 HIGH; `--json`,
452+ `--auto-redact`, `--repo-visibility`, `--from-file`). `bin/gstack-redact-prepush`
453+ is the opt-in git hook.
454+- **Skill docs are generated** from `scripts/resolvers/redact-doc.ts`
455+ (`{{REDACT_TAXONOMY_TABLE}}`, `{{REDACT_INVOCATION_BLOCK:<sink>}}`) so /spec,
456+ /cso, /ship, /document-release, /document-generate never drift from the engine.
457+- **Scan-at-sink:** always scan the EXACT bytes that will be sent — write to a
458+ temp file, scan that file, pass the SAME file to `gh`/`git`. Never scan a string
459+ then re-render (that reopens a scan-vs-send gap).
460+- **Visibility (no tier promotion):** resolve once per run, order = local config
461+ (`gstack-config get redact_repo_visibility`, ~/.gstack so never committed) → gh
462+ → glab → unknown(=public-strict). Public repos get STERNER per-finding
463+ confirmation (no batch-acknowledge, no silent-proceed); MEDIUM is never
464+ auto-promoted to HIGH.
465+- **Tool-attributed fences:** wrap Codex/Greptile/eval output in ` ```codex-review `
466+ / ` ```greptile ` fences so example credentials those tools quote WARN-degrade
467+ instead of blocking. A live-format credential inside the fence still blocks.
468+- **Config keys:** `redact_repo_visibility` (public|private|unknown, local-only
469+ override for repos gh/glab can't read), `redact_prepush_hook` (true|false).
470+ There is intentionally NO key to disable HIGH blocking.
471+- **Audit:** the /spec semantic pass appends a content-free record (categories +
472+ body sha256, no spec text) to `~/.gstack/security/semantic-reviews.jsonl` (0600).
473+ 
474+## Commit style
475+ 
476+**Always bisect commits.** Every commit should be a single logical change. When
477+you've made multiple changes (e.g., a rename + a rewrite + new tests), split them
478+into separate commits before pushing. Each commit should be independently
479+understandable and revertable.
480+ 
481+Examples of good bisection:
482+- Rename/move separate from behavior changes
483+- Test infrastructure (touchfiles, helpers) separate from test implementations
484+- Template changes separate from generated file regeneration
485+- Mechanical refactors separate from new features
486+ 
487+When the user says "bisect commit" or "bisect and push," split staged/unstaged
488+changes into logical commits and push.
489+ 
490+## Slop-scan: AI code quality, not AI code hiding
491+ 
492+We use [slop-scan](https://github.com/benvinegar/slop-scan) to catch patterns where
493+AI-generated code is genuinely worse than what a human would write. We are NOT trying
494+to pass as human code. We are AI-coded and proud of it. The goal is code quality.
495+ 
112496 ```bash
113−bun install # install dependencies
114−bun test # run free tests (no API spend)
115−bun run test:windows # curated Windows-safe subset (runs on windows-latest)
116−bun run build # generate docs + compile binaries
117−bun run gen:skill-docs # regenerate SKILL.md files from templates
118−bun run skill:check # health dashboard for all skills
497+npx slop-scan scan . # human-readable report
498+npx slop-scan scan . --json # machine-readable for diffing
119499 ```
120500  
121−## Platform support
501+Config: `slop-scan.config.json` at repo root (currently excludes `**/vendor/**`).
122502  
123−- **macOS** + **Linux**: full test suite supported.
124−- **Windows**: curated Windows-safe subset runs on `windows-latest` via the
125− `windows-free-tests` CI job. Setup script (`./setup`) requires Git Bash or
126− MSYS today; native PowerShell support is a future expansion. The `bin/gstack-paths`
127− helper resolves state roots through `CLAUDE_PLUGIN_DATA` / `GSTACK_HOME` so plugin
128− installs work on every platform.
503+### What to fix (genuine quality improvements)
129504  
130−## Key conventions
505+- **Empty catches around file ops** — use `safeUnlink()` (ignores ENOENT, rethrows
506+ EPERM/EIO). A swallowed EPERM in cleanup means silent data loss.
507+- **Empty catches around process kills** — use `safeKill()` (ignores ESRCH, rethrows
508+ EPERM). A swallowed EPERM means you think you killed something you didn't.
509+- **Redundant `return await`** — remove when there's no enclosing try block. Saves a
510+ microtask, signals intent.
511+- **Typed exception catches** — `catch (err) { if (!(err instanceof TypeError)) throw err }`
512+ is genuinely better than `catch {}` when the try block does URL parsing or DOM work.
513+ You know what error you expect, so say so.
131514  
132−- SKILL.md files are **generated** from `.tmpl` templates. Edit the template, not the output.
133−- Run `bun run gen:skill-docs --host codex` to regenerate Codex-specific output.
134−- The browse binary provides headless browser access. Use `$B <command>` in skills.
135−- Safety skills (careful, freeze, guard) use inline advisory prose — always confirm before destructive operations.
136−- State paths resolve via `bin/gstack-paths` (sourced via `eval "$(...)"`). Honors `GSTACK_HOME`, `CLAUDE_PLUGIN_DATA`, `CLAUDE_PLANS_DIR`.
137−- The `claude` CLI binary resolves via `browse/src/claude-bin.ts` (`Bun.which()` + `GSTACK_CLAUDE_BIN` override). Set `GSTACK_CLAUDE_BIN=wsl` plus `GSTACK_CLAUDE_BIN_ARGS='["claude"]'` to run Claude through WSL on Windows.
515+### What NOT to fix (linter gaming, not quality)
516+ 
517+- **String-matching on error messages** — `err.message.includes('closed')` is brittle.
518+ Playwright/Chrome can change wording anytime. If a fire-and-forget operation can fail
519+ for ANY reason and you don't care, `catch {}` is the correct pattern.
520+- **Adding comments to exempt pass-through wrappers** — "alias for active session" above
521+ a method just to trip slop-scan's exemption rule is noise, not documentation.
522+- **Converting extension catch-and-log to selective rethrow** — Chrome extensions crash
523+ entirely on uncaught errors. If the catch logs and continues, that IS the right pattern
524+ for extension code. Don't make it throw.
525+- **Tightening best-effort cleanup paths** — shutdown, emergency cleanup, and disconnect
526+ code should use `safeUnlinkQuiet()` (swallows ALL errors). A cleanup path that throws
527+ on EPERM means the rest of cleanup doesn't run. That's worse.
528+ 
529+### Utilities in `browse/src/error-handling.ts`
530+ 
531+| Function | Use when | Behavior |
532+|----------|----------|----------|
533+| `safeUnlink(path)` | Normal file deletion | Ignores ENOENT, rethrows others |
534+| `safeUnlinkQuiet(path)` | Shutdown/emergency cleanup | Swallows all errors |
535+| `safeKill(pid, signal)` | Sending signals | Ignores ESRCH, rethrows others |
536+| `isProcessAlive(pid)` | Boolean process checks | Returns true/false, never throws |
537+ 
538+### Score tracking
539+ 
540+Baseline (2026-04-09, before cleanup): 100 findings, 432.8 score, 2.38 score/file.
541+After cleanup: 90 findings, 358.1 score, 1.96 score/file.
542+ 
543+Don't chase the number. Fix patterns that represent actual code quality problems.
544+Accept findings where the "sloppy" pattern is the correct engineering choice.
545+ 
546+## Community PR guardrails
547+ 
548+When reviewing or merging community PRs, **always AskUserQuestion** before accepting
549+any commit that:
550+ 
551+1. **Touches ETHOS.md** — this file is Garry's personal builder philosophy. No edits
552+ from external contributors or AI agents, period.
553+2. **Removes or softens promotional material** — YC references, founder perspective,
554+ and product voice are intentional. PRs that frame these as "unnecessary" or
555+ "too promotional" must be rejected.
556+3. **Changes Garry's voice** — the tone, humor, directness, and perspective in skill
557+ templates, CHANGELOG, and docs are not generic. PRs that rewrite voice to be
558+ more "neutral" or "professional" must be rejected.
559+ 
560+Even if the agent strongly believes a change improves the project, these three
561+categories require explicit user approval via AskUserQuestion. No exceptions.
562+No auto-merging. No "I'll just clean this up."
563+ 
564+## Checking out PRs from garrytan-agents
565+ 
566+When the user says "check out <PR link>" and the PR is from `garrytan-agents/gstack`
567+(or any other fork that is NOT a collaborator on `garrytan/gstack`), do NOT just
568+`gh pr checkout`. Fork PRs don't receive base-repo secrets (`ANTHROPIC_API_KEY`,
569+`OPENAI_API_KEY`, etc.), so the eval/E2E CI jobs fail with empty-env auth errors
570+regardless of what's set on the base repo.
571+ 
572+**Workflow:** push the branch to `garrytan/gstack` (the base repo) and re-target
573+the PR from there.
574+ 
575+Concretely, after `gh pr checkout <N>`:
576+ 
577+1. Note the original PR number and head branch name.
578+2. Push the same branch to the base repo: `git push origin HEAD:<branch-name>`
579+ (origin = `garrytan/gstack`, since the worktree is set up with that remote).
580+3. Close the fork PR (`gh pr close <N> --comment "moving to base-repo branch for secret access"`).
581+4. Open a new PR from the base-repo branch: `gh pr create --base main --head <branch-name>`.
582+5. New PR's workflows will get secrets automatically.
583+ 
584+Why not fix it on the fork side? `garrytan-agents` isn't a collaborator on
585+`garrytan/gstack`. Adding it as a collaborator (option A) or flipping the
586+repo-wide "send secrets to fork PRs" toggle (option B) would let secrets reach
587+fork PRs from anyone — broader blast radius than just moving this one branch.
588+Option C (this section) keeps secret-distribution scope tight.
589+ 
590+If the user asks you to skip the move (e.g., "just leave it as a fork PR"),
591+respect that — eval CI will fail with empty-env auth, but check-freshness,
592+workflow-lint, and windows-tests will still pass on the fork PR.
593+ 
594+## CHANGELOG + VERSION style
595+ 
596+**Versioning invariant (workspace-aware ship).** VERSION is a monotonic ordered
597+release identifier, not a strict semver commitment. The bump level
598+(major/minor/patch/micro) expresses intent at ship time. Queue-advancing past a
599+claimed version within the same bump level is explicitly permitted — if branch A
600+claims v1.7.0.0 as a MINOR and branch B is also a MINOR, B lands at v1.8.0.0
601+(still a MINOR relative to main). Downstream consumers must NOT rely on
602+"MINOR = feature-only, PATCH = fix-only" as a strict contract. This is why
603+`bin/gstack-next-version` advances within the chosen bump level rather than
604+repicking the level when collisions happen.
605+ 
606+**Scale-aware bumps — use common sense.** When the diff is big, bump MINOR (or
607+MAJOR), not PATCH. PATCH is for bug fixes and small additions; MINOR is for
608+substantial new capability or substantial reduction; MAJOR is for breaking
609+changes. Rough guideposts (don't treat as rules, treat as smell-checks):
610+ 
611+- **PATCH (X.Y.Z+1.0)**: bug fix, doc tweak, small additive change, single
612+ test/file added. Net diff under ~500 lines, no new user-facing capability.
613+- **MINOR (X.Y+1.0.0)**: new capability shipped (skill, harness, command, big
614+ refactor), substantial code reduction (compression, migration), or coordinated
615+ multi-file change. Net diff over ~2000 lines added/removed, OR a user-visible
616+ feature you'd put in a tweet.
617+- **MAJOR (X+1.0.0.0)**: breaking change to public surface (CLI flag rename,
618+ skill removed, config format changed), OR a release big enough to be the
619+ headline of a blog post.
620+ 
621+If you find yourself debating "is 10K added + 24K removed really a PATCH?" — it
622+isn't. Bump MINOR. Same for "this adds a whole new test harness with 6 new E2E
623+tests + helper utilities" — MINOR. The bump level is communication to the user
624+about what kind of release this is; don't undersell it.
625+ 
626+When merging origin/main brings a higher VERSION, re-evaluate the bump level
627+against the SCALE of your branch's work, not just whether main moved forward.
628+If main bumped MINOR and your branch is also a substantial change, you bump
629+MINOR again on top (e.g., main at v1.14.0.0, your branch lands v1.15.0.0).
630+ 
631+**VERSION and CHANGELOG are branch-scoped.** Every feature branch that ships gets its
632+own version bump and CHANGELOG entry. The entry describes what THIS branch adds —
633+not what was already on main.
634+ 
635+**The CHANGELOG entry is the diff between main and the shipping branch — what users
636+get when they upgrade. NOT how the branch got there.** A reader landing on the entry
637+should learn what they can do now that they couldn't before; they should not learn
638+about the branch's internal version bumps, the bugs we caught and fixed mid-branch,
639+the plan reviews we ran, or the commits we squashed. That is branch development
640+narrative. It belongs in PR descriptions and commit messages, not CHANGELOG.
641+ 
642+**Never reference branch-internal versions in a CHANGELOG entry.** If your branch
643+bumped VERSION from v1.5.0.0 → v1.5.1.0 → v1.6.0.0 during development and only the
644+final v1.6.0.0 ships to main, the entry must read as if v1.5.1.0 never existed.
645+Concretely, NEVER write:
646+- "v1.5.1.0 had a bug that v1.6.0.0 fixes" — readers don't know about v1.5.1.0; it's
647+ a branch-internal artifact.
648+- "The shipping headline of v1.5.1.0 was broken because..." — same reason. From main's
649+ perspective, v1.5.1.0 was never released.
650+- "Pre-fix tests encoded the broken behavior" — that's a contributor's victory lap,
651+ not a user benefit.
652+- "Two surgical edits, both in the dispatch path" — micro-narrative of the patch.
653+ 
654+Instead, describe the released system: "Browser-skills run end-to-end with the
655+expected tab-access semantics." If a property of the shipped system is worth calling
656+out (e.g., "skill spawns get permissive tab access; pair-agent tunnel tokens require
657+ownership"), document it as a property, not as a fix. The shipped system is what
658+the user gets; the path to that system is invisible to them.
659+ 
660+**When to write the CHANGELOG entry:**
661+- At `/ship` time (Step 13), not during development or mid-branch.
662+- The entry covers ALL commits on this branch vs the base branch.
663+- Never fold new work into an existing CHANGELOG entry from a prior version that
664+ already landed on main. If main has v0.10.0.0 and your branch adds features,
665+ bump to v0.10.1.0 with a new entry — don't edit the v0.10.0.0 entry.
666+ 
667+**Key questions before writing:**
668+1. What branch am I on? What did THIS branch change?
669+2. Is the base branch version already released? (If yes, bump and create new entry.)
670+3. Does an existing entry on this branch already cover earlier work? (If yes, replace
671+ it with one unified entry for the final version.)
672+ 
673+**Merging main does NOT mean adopting main's version.** When you merge origin/main into
674+a feature branch, main may bring new CHANGELOG entries and a higher VERSION. Your branch
675+still needs its OWN version bump on top. If main is at v0.13.8.0 and your branch adds
676+features, bump to v0.13.9.0 with a new entry. Never jam your changes into an entry that
677+already landed on main. Your entry goes on top because your branch lands next.
678+ 
679+**After merging main, always check:**
680+- Does CHANGELOG have your branch's own entry separate from main's entries?
681+- Is VERSION higher than main's VERSION?
682+- Is your entry the topmost entry in CHANGELOG (above main's latest)?
683+If any answer is no, fix it before continuing.
684+ 
685+**After any CHANGELOG edit that moves, adds, or removes entries,** immediately run
686+`grep "^## \[" CHANGELOG.md` to verify no duplicates and a sensible reverse-chronological
687+order. Gaps between version numbers are fine. A branch that ships at v1.6.4.0 without
688+a prior v1.5.2.0 or v1.5.3.0 entry on main is correct — those were branch-internal
689+version numbers that never landed. Do not back-fill gaps with placeholder entries.
690+ 
691+**Never orphan branch-internal versions.** If your branch bumped VERSION several times
692+during development (v1.5.1.0 → v1.5.2.0 → v1.6.4.0, say) and those earlier entries were
693+never released to main, the final ship consolidates ALL of them into a single entry at
694+the final version (v1.6.4.0). Collapse them — delete the old entries and move their
695+content into the final entry, re-version table columns accordingly. Readers see one
696+release, not a branch diary. Gaps are fine (v1.6.3.0 → v1.6.4.0 with no v1.5.x
697+in between on main is correct).
698+ 
699+CHANGELOG.md is **for users**, not contributors. Write it like product release notes:
700+ 
701+- Lead with what the user can now **do** that they couldn't before. Sell the feature.
702+- Use plain language, not implementation details. "You can now..." not "Refactored the..."
703+- **Never mention TODOS.md, internal tracking, eval infrastructure, or contributor-facing
704+ details.** These are invisible to users and meaningless to them.
705+- Put contributor/internal changes in a separate "For contributors" section at the bottom.
706+- Every entry should make someone think "oh nice, I want to try that."
707+- No jargon: say "every question now tells you which project and branch you're in" not
708+ "AskUserQuestion format standardized across skill templates via preamble resolver."
709+ 
710+**Only document what shipped between main and this change.** Readers do not care how
711+we got here. Keep out of the CHANGELOG, always:
712+ 
713+- Branch resyncs, merge commits with main, rebase activity.
714+- Plan approvals, review outcomes (CEO / eng / design / outside-voice / codex findings),
715+ AskUserQuestion decisions, scope negotiations.
716+- "Work queued," "plan approved," "in-progress," "will ship later" — the CHANGELOG
717+ documents what DID ship, not what MIGHT ship.
718+- Version-bump housekeeping when no user-facing work actually landed.
719+ 
720+If the diff between the base branch version and this version has no user-facing change
721+(only merges, only CHANGELOG edits, only placeholder work), the honest entry is one
722+sentence: "Version bump for branch-ahead discipline. No user-facing changes yet." Stop
723+there. Do not pad. Do not explain the plan that will ship eventually. Do not narrate
724+the branch's history. When real work lands, the entry will replace this at /ship time.
725+ 
726+### Release-summary format (every `## [X.Y.Z]` entry)
727+ 
728+Every version entry in `CHANGELOG.md` MUST start with a release-summary section in
729+the GStack/Garry voice, one viewport's worth of prose + tables that lands like a
730+verdict, not marketing. The itemized changelog (subsections, bullets, files) goes
731+BELOW that summary, separated by a `### Itemized changes` header.
732+ 
733+The release-summary section gets read by humans, by the auto-update agent, and by
734+anyone deciding whether to upgrade. The itemized list is for agents that need to
735+know exactly what changed.
736+ 
737+Structure for the top of every `## [X.Y.Z]` entry:
738+ 
739+1. **Two-line bold headline** (10-14 words total). Should land like a verdict, not
740+ marketing. Sound like someone who shipped today and cares whether it works.
741+2. **Lead paragraph** (3-5 sentences). What shipped, what changed for the user.
742+ Specific, concrete, no AI vocabulary, no em dashes, no hype.
743+3. **A "The X numbers that matter" section** with:
744+ - One short setup paragraph naming the source of the numbers (real production
745+ deployment OR a reproducible benchmark, name the file/command to run).
746+ - A table of 3-6 key metrics with BEFORE / AFTER / Δ columns.
747+ - A second optional table for per-category breakdown if relevant.
748+ - 1-2 sentences interpreting the most striking number in concrete user terms.
749+4. **A "What this means for [audience]" closing paragraph** (2-4 sentences) tying
750+ the metrics to a real workflow shift. End with what to do.
751+ 
752+Voice rules for the release summary:
753+- No em dashes (use commas, periods, "...").
754+- No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or
755+ banned phrases ("here's the kicker", "the bottom line", etc.).
756+- Real numbers, real file names, real commands. Not "fast" but "~30s on 30K pages."
757+- Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
758+- Connect to user outcomes: "the agent does ~3x less reading" beats "improved precision."
759+- Be direct about quality. "Well-designed" or "this is a mess." No dancing.
760+ 
761+Source material:
762+- CHANGELOG previous entry for prior context.
763+- Benchmark files or `/retro` output for headline numbers.
764+- Recent commits (`git log <prev-version>..HEAD --oneline`) for what shipped.
765+- Don't make up numbers. If a metric isn't in a benchmark or production data,
766+ don't include it. Say "no measurement yet" if asked.
767+ 
768+Target length: ~250-350 words for the summary. Should render as one viewport.
769+ 
770+### Itemized changes (below the release summary)
771+ 
772+Write `### Itemized changes` and continue with the detailed subsections (Added,
773+Changed, Fixed, For contributors). Same rules as the user-facing voice guidance
774+above, plus:
775+ 
776+- **Always credit community contributions.** When an entry includes work from a
777+ community PR, name the contributor with `Contributed by @username`. Contributors
778+ did real work. Thank them publicly every time, no exceptions.
779+ 
780+## AI effort compression
781+ 
782+When estimating or discussing effort, always show both human-team and CC+gstack time:
783+ 
784+| Task type | Human team | CC+gstack | Compression |
785+|-----------|-----------|-----------|-------------|
786+| Boilerplate / scaffolding | 2 days | 15 min | ~100x |
787+| Test writing | 1 day | 15 min | ~50x |
788+| Feature implementation | 1 week | 30 min | ~30x |
789+| Bug fix + regression test | 4 hours | 15 min | ~20x |
790+| Architecture / design | 2 days | 4 hours | ~5x |
791+| Research / exploration | 1 day | 3 hours | ~3x |
792+ 
793+Completeness is cheap. Don't recommend shortcuts when the complete implementation
794+is achievable. Boil the ocean — the complete thing is the goal; only genuinely
795+unrelated multi-quarter migrations are separate scope, never an excuse for a
796+shortcut. See the Completeness Principle in the skill preamble for the full
797+philosophy.
798+ 
799+## Search before building
800+ 
801+Before designing any solution that involves concurrency, unfamiliar patterns,
802+infrastructure, or anything where the runtime/framework might have a built-in:
803+ 
804+1. Search for "{runtime} {thing} built-in"
805+2. Search for "{thing} best practice {current year}"
806+3. Check official runtime/framework docs
807+ 
808+Three layers of knowledge: tried-and-true (Layer 1), new-and-popular (Layer 2),
809+first-principles (Layer 3). Prize Layer 3 above all. See ETHOS.md for the full
810+builder philosophy.
811+ 
812+## Local plans
813+ 
814+Contributors can store long-range vision docs and design documents in `~/.gstack-dev/plans/`.
815+These are local-only (not checked in). When reviewing TODOS.md, check `plans/` for candidates
816+that may be ready to promote to TODOs or implement.
817+ 
818+## E2E eval failure blame protocol
819+ 
820+When an E2E eval fails during `/ship` or any other workflow, **never claim "not
821+related to our changes" without proving it.** These systems have invisible couplings —
822+a preamble text change affects agent behavior, a new helper changes timing, a
823+regenerated SKILL.md shifts prompt context.
824+ 
825+**Required before attributing a failure to "pre-existing":**
826+1. Run the same eval on main (or base branch) and show it fails there too
827+2. If it passes on main but fails on the branch — it IS your change. Trace the blame.
828+3. If you can't run on main, say "unverified — may or may not be related" and flag it
829+ as a risk in the PR body
830+ 
831+"Pre-existing" without receipts is a lazy claim. Prove it or don't say it.
832+ 
833+## Long-running tasks: don't give up
834+ 
835+When running evals, E2E tests, or any long-running background task, **poll until
836+completion**. Use `sleep 180 && echo "ready"` + `TaskOutput` in a loop every 3
837+minutes. Never switch to blocking mode and give up when the poll times out. Never
838+say "I'll be notified when it completes" and stop checking — keep the loop going
839+until the task finishes or the user tells you to stop.
840+ 
841+The full E2E suite can take 30-45 minutes. That's 10-15 polling cycles. Do all of
842+them. Report progress at each check (which tests passed, which are running, any
843+failures so far). The user wants to see the run complete, not a promise that
844+you'll check later.
845+ 
846+## Running evals as an agent: always detach (SIGTERM-proof)
847+ 
848+When **you (an agent/harness)** launch a long eval/benchmark run, run it through
849+`bin/gstack-detach` — NEVER as a plain backgrounded Bash task. A plain background
850+task lives in the harness's process group, so a SIGTERM ("polite quit") on a turn
851+boundary, a stopped Monitor, or an interruption kills the run mid-flight (observed:
852+`script "test:gate" was terminated by signal SIGTERM` ~40 min into a run). On macOS
853+the run can also die to idle-sleep. `gstack-detach` fixes both: a fresh session
854+(escapes the group SIGTERM) wrapped in `caffeinate -i` (blocks idle-sleep).
855+ 
856+- Use the `eval:bg*` scripts (`eval:bg`, `eval:bg:all`, `eval:bg:gate`,
857+ `eval:bg:periodic`) — they wrap the eval command in `gstack-detach` with the
858+ machine-wide `gstack-evals` lock (concurrent worktrees serialize instead of
859+ saturating the shared model API), a per-tier watchdog, and a **run-scoped** log
860+ under `~/.gstack-dev/eval-runs/` (no shared-`/tmp` collision). Each prints its
861+ log path. Or call `gstack-detach [--lock NAME] [--timeout SECS] [--label LBL] --
862+ <cmd>` directly for any long agent job. Export `ANTHROPIC_API_KEY` first (never
863+ pass keys in argv).
864+- Then **poll the printed logfile** with a death-aware watcher: break on the
865+ guaranteed `### gstack-detach EXIT=<code> ###` sentinel (success AND failure are
866+ both marked, so silence is never mistaken for success). The detached run survives
867+ even if your watcher gets reaped, so re-checking the log always works.
868+- Why the lock: a shared dev box with several Conductor worktrees will rate-limit
869+ the model API if two eval suites run at once (15-way concurrency each), which
870+ mass-times-out E2E tests. The lock makes the second run WAIT, not collide.
871+- Humans running `bun run test:evals` foreground in their own terminal don't need
872+ this — Ctrl-C is intended there. Detachment is for agent-launched runs only.
873+ 
874+## E2E test fixtures: extract, don't copy
875+ 
876+**NEVER copy a full SKILL.md file into an E2E test fixture.** SKILL.md files are
877+1500-2000 lines. When `claude -p` reads a file that large, context bloat causes
878+timeouts, flaky turn limits, and tests that take 5-10x longer than necessary.
879+ 
880+Instead, extract only the section the test actually needs:
881+ 
882+```typescript
883+// BAD — agent reads 1900 lines, burns tokens on irrelevant sections
884+fs.copyFileSync(path.join(ROOT, 'ship', 'SKILL.md'), path.join(dir, 'ship-SKILL.md'));
885+ 
886+// GOOD — agent reads ~60 lines, finishes in 38s instead of timing out
887+const full = fs.readFileSync(path.join(ROOT, 'ship', 'SKILL.md'), 'utf-8');
888+const start = full.indexOf('## Review Readiness Dashboard');
889+const end = full.indexOf('\n---\n', start);
890+fs.writeFileSync(path.join(dir, 'ship-SKILL.md'), full.slice(start, end > start ? end : undefined));
891+```
892+ 
893+Also when running targeted E2E tests to debug failures:
894+- Run in **foreground** (`bun test ...`), not background with `&` and `tee`
895+- Never `pkill` running eval processes and restart — you lose results and waste money
896+- One clean run beats three killed-and-restarted runs
897+ 
898+## Publishing native OpenClaw skills to ClawHub
899+ 
900+Native OpenClaw skills live in `openclaw/skills/gstack-openclaw-*/SKILL.md`. These are
901+hand-crafted methodology skills (not generated by the pipeline) published to ClawHub
902+so any OpenClaw user can install them.
903+ 
904+**Publishing:** The command is `clawhub publish` (NOT `clawhub skill publish`):
905+ 
906+```bash
907+clawhub publish openclaw/skills/gstack-openclaw-office-hours \
908+ --slug gstack-openclaw-office-hours --name "gstack Office Hours" \
909+ --version 1.0.0 --changelog "description of changes"
910+```
911+ 
912+Repeat for each skill: `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`,
913+`gstack-openclaw-retro`. Bump `--version` on each update.
914+ 
915+**Auth:** `clawhub login` (opens browser for GitHub auth). `clawhub whoami` to verify.
916+ 
917+**Updating:** Same `clawhub publish` command with a higher `--version` and `--changelog`.
918+ 
919+**Verification:** `clawhub search gstack` to confirm they're live.
920+ 
921+## Deploying to the active skill
922+ 
923+The active skill lives at `~/.claude/skills/gstack/`. After making changes:
924+ 
925+1. Push your branch
926+2. Fetch and reset in the skill directory: `cd ~/.claude/skills/gstack && git fetch origin && git reset --hard origin/main`
927+3. Rebuild: `cd ~/.claude/skills/gstack && bun run build`
928+ 
929+**If you use gbrain:** the `git reset --hard` in step 2 reverts the brain-aware
930+(`GBRAIN_CONTEXT_LOAD` / `GBRAIN_SAVE_RESULTS`) blocks that `gstack-config
931+gbrain-refresh` renders into the install (those generated blocks differ from
932+`main` by design). After deploying, re-run `gstack-config gbrain-refresh` to
933+restore them across all your projects' Claude sessions. It's idempotent.
934+ 
935+Or copy the binaries directly:
936+- `cp browse/dist/browse ~/.claude/skills/gstack/browse/dist/browse`
937+- `cp design/dist/design ~/.claude/skills/gstack/design/dist/design`
938+ 
939+## Skill routing
940+ 
941+When the user's request matches an available skill, invoke it via the Skill tool. When in doubt, invoke the skill.
942+ 
943+Key routing rules:
944+- Product ideas/brainstorming → invoke /office-hours
945+- Strategy/scope → invoke /plan-ceo-review
946+- Architecture → invoke /plan-eng-review
947+- Design system/plan review → invoke /design-consultation or /plan-design-review
948+- Full review pipeline → invoke /autoplan
949+- Bugs/errors → invoke /investigate
950+- QA/testing site behavior → invoke /qa or /qa-only
951+- Code review/diff check → invoke /review
952+- Visual polish → invoke /design-review
953+- Ship/deploy/PR → invoke /ship or /land-and-deploy
954+- Save progress → invoke /context-save
955+- Resume context → invoke /context-restore
956+ 
957+## Cross-session decision memory
958+ 
959+Durable decisions and their rationale are captured in an append-only, event-sourced
960+store at `~/.gstack/projects/<slug>/decisions.jsonl` so neither you nor the user
961+re-litigates a settled call or loses the "why" across sessions. This is the reliable,
962+file-only path: it works with gbrain OFF. (gbrain semantic recall is an optional
963+enhancement layered on top, never a dependency.)
964+ 
965+- **Resurface** active decisions before re-deciding: `bin/gstack-decision-search`
966+ (`--recent N`, `--scope repo|branch|issue`, `--query KW`, `--all`, `--json`).
967+ Add `--semantic` (with `--query`) to append related hits from gbrain memory when
968+ it's up; it degrades silently to the reliable file results when gbrain is off.
969+ Session start already surfaces scope-relevant active decisions via Context Recovery.
970+ If a decision is listed, treat it as settled with its rationale; if you're about to
971+ reverse it, say so explicitly.
972+- **Capture** a DURABLE decision when you or the user make one:
973+ `bin/gstack-decision-log '{"decision":"...","rationale":"...","scope":"repo|branch|issue","source":"user|skill|agent","confidence":1-10}'`.
974+ Reverse a prior call with `--supersede <id>`; expunge an accidental secret with
975+ `--redact <id>`; rewrite the log to the active set with `--compact`. Non-interactive
976+ (never prompts), injection-sanitized, and HIGH-secret-blocking on write.
977+- **Durable means:** architecture choice, scope cut, tool/vendor choice, or a reversal
978+ of a prior call. NOT a turn-level edit, a phrasing tweak, or anything trivially
979+ re-derivable. Capture is curated at the source — log durable decisions only, or the
980+ store becomes noise.
981+ 
982+## GBrain Search Guidance (configured by /sync-gbrain)
983+<!-- gstack-gbrain-search-guidance:start -->
984+ 
985+GBrain is set up and synced on this machine. The agent should prefer gbrain
986+over Grep when the question is semantic or when you don't know the exact
987+identifier yet.
988+ 
989+**This worktree is pinned to a worktree-scoped code source** via the
990+`.gbrain-source` file in the repo root (kubectl-style context). Any
991+`gbrain code-def`, `code-refs`, `code-callers`, `code-callees`, or `query`
992+call from anywhere under this worktree routes to that source by default —
993+no `--source` flag needed. Conductor sibling worktrees of the same repo
994+each have their own pin and their own indexed pages, so semantic results
995+match the actual code on disk in this worktree.
996+ 
997+Two indexed corpora available via the `gbrain` CLI:
998+- This worktree's code (auto-pinned via `.gbrain-source`).
999+- `~/.gstack/` curated memory (registered as `gstack-brain-<user>` source via
1000+ the existing federation pipeline).
1001+ 
1002+Prefer gbrain when:
1003+- "Where is X handled?" / semantic intent, no exact string yet:
1004+ `gbrain search "<terms>"` or `gbrain query "<question>"`
1005+- "Where is symbol Y defined?" / symbol-based code questions:
1006+ `gbrain code-def <symbol>` or `gbrain code-refs <symbol>`
1007+- "What calls Y?" / "What does Y depend on?":
1008+ `gbrain code-callers <symbol>` / `gbrain code-callees <symbol>`
1009+- "What did we decide last time?" / past plans, retros, learnings:
1010+ `gbrain search "<terms>" --source gstack-brain-<user>`
1011+ 
1012+Grep is still right for known exact strings, regex, multiline patterns, and
1013+file globs. Run `/sync-gbrain` after meaningful code changes; for ongoing
1014+auto-sync across all worktrees, run `gbrain autopilot --install` once per
1015+machine — gbrain's daemon handles incremental refresh on a schedule.
1016+ 
1017+Safety: don't run `/sync-gbrain` while `gbrain autopilot` is active — the
1018+orchestrator refuses destructive source ops when it detects a running autopilot
1019+to avoid racing it (#1734). Prefer registering user repos with `gbrain sources
1020+add --path <dir>` (no `--url`): URL-managed sources can auto-reclone, and the
1021+sync code walk for them requires an explicit `--allow-reclone` opt-in.
1022+ 
1023+<!-- gstack-gbrain-search-guidance:end -->
1381024  
RuleStack

Built by

Kynth Studio

Directory

Configs
Stacks
Compare formats
Diff two configs
Best AGENTS.md examples

Formats

AGENTS.md
CLAUDE.md
Cursor rules
Copilot instructions

Reference

Read API
Corpus health
Privacy Policy
Terms

RuleStack

RuleStack

Built by

Kynth Studio

Directory

Configs
Stacks
Compare formats
Diff two configs
Best AGENTS.md examples

Formats

AGENTS.md
CLAUDE.md
Cursor rules
Copilot instructions

Reference

Read API
Corpus health
Privacy Policy
Terms

RuleStack

RuleStack

Built by

Kynth Studio

Directory

Configs
Stacks
Compare formats
Diff two configs
Best AGENTS.md examples

Formats

AGENTS.md
CLAUDE.md
Cursor rules
Copilot instructions

Reference

Read API
Corpus health
Privacy Policy
Terms

RuleStack