## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Humans sign in through the auth layer; every authenticated request
parses the session/user profile with `currentUserProfileSchema` in
`packages/shared`
> - The schema requires `name` to be `null` or a non-empty string, but
some identity providers hand back `name: ""` for users who never set a
display name
> - For those users the session payload fails validation on every
request, so the app treats them as unauthenticated and bounces them to
`/auth` in a loop — they can never get in
> - This pull request preprocesses empty/whitespace-only names to `null`
before validation, so the existing `min(1).max(120).nullable()` rule
still holds for real names
> - Review found the sibling `email` field has the same failure mode
(the DB `auth` schema declares `email` as `notNull`, so a provider that
supplies no email stores `""`, which `z.string().email()` rejects); the
same preprocess is applied there
> - The benefit is that users whose provider reports an empty name (or
email) can sign in normally instead of being locked out, with no change
in behavior for anyone else
## Linked Issues or Issue Description
No existing issue; described in-PR following the bug report template:
**What happened?**
Users whose auth provider returns `name: ""` (empty string) in the
profile payload fail `currentUserProfileSchema` / `authSessionSchema`
parsing (`name: z.string().min(1)...`). The parse failure makes the
session look invalid and the UI redirects to `/auth` on every attempt —
an endless sign-in loop. The `email` field has the same failure mode
(`z.string().email()` rejects `""`).
**Expected behavior:**
An empty display name (or email) should be treated the same as a missing
one (`null`); the user should be signed in normally.
**Steps to reproduce:**
Sign in with an account whose upstream identity record has an
empty-string name (or set a user's `name` column to `''` directly), then
load the app: session parse fails and you are bounced back to `/auth`.
**Adapter(s) involved:**
Not adapter-specific (core bug).
**Deployment mode / version:**
Any; reproduces on current `master`.
## What Changed
- `packages/shared/src/validators/access.ts`:
`currentUserProfileSchema.name` now runs through `z.preprocess` that
coerces empty or whitespace-only strings to `null` before the existing
`z.string().min(1).max(120).nullable()` validation.
- `packages/shared/src/validators/access.ts`: the same preprocess is
applied to `email` (review follow-up): `users.email` is `notNull` in the
DB schema, so a provider without an email stores `""`, which
`z.string().email()` rejects — the identical lockout loop.
Empty/whitespace-only emails now coerce to `null` (the field was already
nullable); malformed non-empty emails are still rejected.
- `packages/shared/src/validators/access.test.ts` (new): covers
empty-string → `null`, whitespace-only → `null`, real values preserved,
`null` preserved, and malformed non-empty email still rejected — for
both `name` and `email`, and the same cases through `authSessionSchema`.
## Verification
- `vitest run src/validators/access.test.ts` in `packages/shared` — 13
tests pass.
- `tsc --noEmit -p packages/shared` passes.
- Manual: parse `{ id, email: "", name: "", image: null }` with
`currentUserProfileSchema` — succeeds with `name: null` and `email:
null` instead of failing validation.
## Risks
- Low risk. The change only widens accepted input (empty/whitespace
string → `null` for `name` and `email`); every previously valid payload
parses identically. `updateCurrentUserProfileSchema` (user-initiated
rename) is untouched and still rejects empty names.
## Model Used
Claude Fable 5 (Anthropic, `claude-fable-5`, agentic coding harness via
Claude Code, extended reasoning enabled). Original fix drafted with
Claude Sonnet 4.6.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents execute inside sandboxed runtime containers built from
`docker/agent-runtime/Dockerfile.base`
> - OpenCode's skill tooling shells out to ripgrep; when `rg` is not on
PATH it tries to download a pinned build from
`github.com/BurntSushi/ripgrep/releases` at run time
> - In a sandbox with locked-down egress that download hangs ~127s and
then fails, burning run budget on every agent run before the agent
reaches its actual work
> - The root repo `Dockerfile` already installs ripgrep; the
agent-runtime base image drifted without it
> - This pull request adds `ripgrep` to the base image's apt install so
OpenCode uses the system binary and never reaches for the network
> - The benefit is that every sandboxed agent run stops wasting ~2
minutes on a doomed download and spends its budget on real work
## Linked Issues or Issue Description
No existing issue — problem described here per the bug report template:
- **What happened:** Sandboxed agent runs using OpenCode stall for ~127
seconds at startup, then log `Transport error ...
BurntSushi/ripgrep/releases/download/...` before continuing degraded.
- **Expected behavior:** The agent starts working immediately; skill
tooling finds `rg` on PATH.
- **Root cause:** The agent-runtime base image
(`docker/agent-runtime/Dockerfile.base`) does not ship ripgrep, so
OpenCode falls back to downloading a pinned build at run time, which
egress-restricted sandboxes block.
- **Reproduction:** Run any OpenCode-backed agent in a sandbox with
locked-down egress using the current agent-runtime image; observe the
startup hang and transport error.
Supersedes #8859.
## What Changed
- Added `ripgrep` to the existing `apt-get install
--no-install-recommends` list in `docker/agent-runtime/Dockerfile.base`
- Added an explanatory comment documenting why ripgrep must be present
(run-time download fallback + egress-restricted sandboxes), restoring
parity with the root repo `Dockerfile`
## Verification
- `docker build -f docker/agent-runtime/Dockerfile.base .` then `docker
run --rm <image> rg --version` — prints the ripgrep version from the
system package
- Run an OpenCode-backed agent in an egress-restricted sandbox on the
new image: no `BurntSushi/ripgrep` download attempt, no ~127s startup
stall
## Risks
- Low risk: no behavior change beyond shipping one additional apt
package in the base image; slightly larger image size
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude (Anthropic), model ID `claude-fable-5` (Fable 5), agentic
coding via Claude Code with tool use; original change authored with
Claude Opus 4.8 (1M context) and re-based/re-verified with Fable 5
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents and humans coordinate on issue threads, where
`request_confirmation` cards capture pending decisions; a genuine human
comment on the thread is meant to supersede (cancel) a card.
> - Supersession is keyed on `!comment.authorUserId` — the guard assumes
only real human comments carry a user id.
> - But local-CLI agent heartbeats post comments under user auth, so a
machine comment's `authorUserId` is populated **nondeterministically per
run** (the same agent resolves as `agent` on one run and `user` on
another).
> - As a result an agent's own on-thread comment — or a teammate's, from
a different run — can carry `authorUserId` and silently expire a pending
decision card. A card was observed expiring 7ms after its own automated
comment landed, stranding the decision with no live approval path.
> - This PR switches the discriminator to a durable, deterministic
signal already persisted on every comment — `created_by_run_id` — so
only comments with **no run context** (genuine board-UI comments)
supersede.
> - The benefit: machine-authored comments can never again expire
decision cards, while real human supersession is preserved exactly.
## Linked Issues or Issue Description
No public GitHub issue — describing the bug in-PR.
- **What happened:** A pending `request_confirmation` decision card was
expired by an automated, machine-authored comment on the same thread.
Supersession is keyed on `!comment.authorUserId`, but local-CLI agent
heartbeats post under user auth, so a machine comment's `authorUserId`
is set nondeterministically per run. An agent's own comment (or a
teammate's, from a different run) can therefore carry a user id and
expire a pending card — one was observed expiring 7ms after its own
automated comment landed.
- **Expected behavior:** Only genuine interactive human (board-UI)
comments should supersede pending decision cards. Machine-authored
comments must never expire them, regardless of how the adapter's auth
resolves.
- **Steps to reproduce:** With a pending `request_confirmation` card
(`supersedeOnUserComment: true`), post a comment via a local-CLI agent
run whose actor resolves to `user`; the card expires with outcome
`superseded_by_comment`.
- **Deployment mode:** server (self-hosted), reproduced against
`master`.
Related PRs (same lifecycle area, not duplicates): #6094 (auto-resolve
stale `request_confirmation` interactions) and #8799 (expire ask-user
questions superseded by comments, merged).
## What Changed
- Supersession now fires **only on comments with no run context**
(`created_by_run_id` is null), in both paths:
- `expireRequestConfirmationsSupersededByComment` (live post path) —
early-return when `comment.createdByRunId` is set.
- `expireRequestConfirmationsSupersededByHistoricalComments` (repair
sweep) — query filters `isNull(created_by_run_id)`.
- Mirrors the existing `shouldImplicitlyMoveCommentedIssueToTodo` reopen
guard, which already uses run context to solve the same
nondeterministic-identity problem.
- Adds live + historical regression tests asserting a run-originated
comment under user auth does not supersede a pending card.
## Verification
- Interactions service suite: **27 tests pass (1 file)**, including the
two new regression tests.
- CI: all substantive gates green (Build, General tests, serialized
server suites, Typecheck, e2e, verify, security-review, policy).
- Manual: with a pending card, a comment carrying `created_by_run_id`
leaves it `pending`; a comment with null run context still supersedes
it.
## Risks
- Low risk, narrowly scoped to the supersession discriminator. Human
supersession is preserved (comments with no run context still cancel
cards); only the machine-authored case is closed.
- No schema migration — `created_by_run_id` is already persisted by
`addComment`.
- Alternatives considered: (a) ignore only the assignee's own run —
misses cross-run machine comments; (b) default `supersedeOnUserComment:
false` for agent-created cards — would drop the legitimate "human
comment redirects → cancel the card" behavior. The run-context guard
covers all machine comments while preserving human supersession.
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
reasoning + tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
Paperclip ticket id — **not yet met**; renaming an open PR's branch
risks closing this PR, so it's flagged for a maintainer to rename safely
(or via the GitHub rename-branch API).
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — N/A
(internal behavior fix, no user-facing docs)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (the only red check is the
automated PR-review template gate this revision addresses)
- [ ] Greptile is 5/5 with no open P2s — re-review requested after this
revision
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 8/8 and focuses on end-to-end coverage,
operator docs, evals, and release notes
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: The complete stack needs discoverable browser scenarios,
operator guidance, threat modeling, eval coverage, and a parity proof
before merge.
- Proposed solution: Adds MCP user-story and Smoke Lab e2e suites,
docs/evals/release notes, the skill update, and the root e2e driver
script registration.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/07-ui-apps-activation`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for flag audit and e2e/browser acceptance;
Greptile on every PR.
## What Changed
- Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release
notes, the skill update, and the root e2e driver script registration.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- `node --check scripts/e2e-mcp-user-stories.mjs`
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
--list` — 43 tests discovered
- `git diff pap10341-split/08-e2e-docs
6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes)
## Risks
- Browser suites depend on runtime services and environment setup; this
PR validates discovery locally while QA owns full flag-on/flag-off
execution.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 7/8 and focuses on Apps/Gateways UI and
feature-flagged activation
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: The standalone UI foundation needs feature-flagged routes,
navigation, app detail flows, gateways, Storybook scenarios, and QA
configuration.
- Proposed solution: Adds Apps/Gateways pages and components,
navigation/route activation, remaining page integrations, Storybook
stories, and QA Vite configuration.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is
`pap10341-split/06-ui-tools-foundation`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: UXDesigner sanity pass on Apps/Gateways flows and
flag-off behavior; Greptile on every PR.
## What Changed
- Adds Apps/Gateways pages and components, navigation/route activation,
remaining page integrations, Storybook stories, and QA Vite
configuration.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- `pnpm check:token-gates` — all gates clean
- Focused UI Vitest run with `NODE_ENV=test` — 17 files, 147 tests
passed
## Risks
- Navigation or flag regressions could expose incomplete experiences;
activation remains controlled by existing experimental settings.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
## UI Evidence
QA captured these from the live Garden MCP split stack at 1440px and
verified clean rendering:



---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 6/8 and focuses on UI API, shared
components, and Tools/Profile surfaces
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: Operators need typed clients and administration surfaces that
compile independently before navigation exposes them.
- Proposed solution: Adds UI APIs, hooks, libraries, shared components,
Tools/Profiles pages, and the plugin settings consumer required by the
new company-scoped API.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is
`pap10341-split/05-runtime-integration`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: UXDesigner sanity pass on Tools/Profile surfaces;
Greptile on every PR.
## What Changed
- Adds UI APIs, hooks, libraries, shared components, Tools/Profiles
pages, and the plugin settings consumer required by the new
company-scoped API.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- `pnpm check:token-gates` — all gates clean
- Focused UI Vitest run with `NODE_ENV=test` — 28 files, 183 tests
passed
## Risks
- Large dead-code UI additions can drift from activation routes; PR 7
supplies the registration layer and top-level parity catches omissions.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
## UI Evidence
QA captured these from the live Garden MCP split stack at 1440px and
verified clean rendering:



---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 5/8 and focuses on remaining adapters,
CLI, plugin examples, and deployment packaging
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: The backend runtime needs packaging, CLI propagation,
worktree provisioning, release manifests, and remaining adapter/plugin
consumers.
- Proposed solution: Adds the remaining runtime/deployment integration
after compile-required contracts and concrete MCP injection moved into
lower server levels.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is
`pap10341-split/04-server-runtime-wiring`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for CLI, packaging, and worktree behavior;
Greptile on every PR.
## What Changed
- Adds the remaining runtime/deployment integration after
compile-required contracts and concrete MCP injection moved into lower
server levels.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- Focused CLI Vitest run — 4 files, 49 tests passed
## Risks
- Packaging omissions could make the feature work in source but fail in
Docker, worktrees, or release assembly.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 4/8 and focuses on gateway runtime, Smoke
Lab, plugins, and server wiring
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: The policy core needs runtime execution, endpoint guards,
route registration, heartbeat integration, and adapter MCP injection to
become operational.
- Proposed solution: Adds the remaining server routes/wiring/consumers,
runtime tests, adapter-utils MCP contracts, and Claude/Codex injection
implementations required by the server layer.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/03-server-tool-access`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: SecurityEngineer for gateway, endpoint guard, token
issuance, and runtime wiring; Greptile on every PR.
## What Changed
- Adds the remaining server routes/wiring/consumers, runtime tests,
adapter-utils MCP contracts, and Claude/Codex injection implementations
required by the server layer.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- Changed server test set — 26 files, 382 tests passed
- Affected server adapter tests — 38 tests passed after concrete adapter
boundary move
- Adapter-utils and Codex focused tests — 76 tests passed
## Risks
- Remote endpoint validation, token handling, and runtime supervision
are security-sensitive and can fail closed or deny legitimate access if
misconfigured.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 3/8 and focuses on tool-access policy and
authorization core
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: Authorization, OAuth binding, secret projection, content
guards, and policy evaluation need a security-reviewable server
boundary.
- Proposed solution: Adds tool-access services/routes/tests plus the
runtime service dependencies directly imported by the core, without
registering the routes in the application.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/02-schema-shared`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: SecurityEngineer for authz, OAuth, secrets, and
content guards; Greptile on every PR.
## What Changed
- Adds tool-access services/routes/tests plus the runtime service
dependencies directly imported by the core, without registering the
routes in the application.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- Focused server Vitest run — 4 files, 143 tests passed
## Risks
- Authorization bugs could permit cross-company or over-broad tool
access; the PR remains inert until PR 4 wiring and requires dedicated
security review.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Its server and UI test suites protect company-scoped plugin access
and instance settings behavior
> - Recent governed-access contracts intentionally added company
invocation scope and new experimental-setting defaults
> - Four existing tests were not updated consistently with those
contracts, causing current-master CI failures unrelated to the changes
under review
> - The runtime behavior is intentional, so changing production code
would weaken the new authorization and settings contracts
> - This pull request aligns the stale tests with current behavior and
removes one UI assertion accidentally pulled forward from a later
stacked feature
> - The benefit is a focused, low-risk repair that restores master CI
without changing application behavior
## Linked Issues or Issue Description
- **Bug:** Current master has four regression failures in plugin
authorization, plugin execution-workspace bridging, instance settings
normalization, and experimental settings UI tests.
- **Expected behavior:** Tests provide required company/invocation
scope, use the governed object-shaped secret reference contract, include
all current defaults, and only assert UI controls implemented at this
stack level.
- **Actual behavior:** Tests exercised obsolete request shapes or
expected a later-stack Apps toggle that is not present on current
master.
- **Reproduction:** Run the four test files listed in the Verification
section on master before this commit.
## What Changed
- Updates plugin config authorization coverage to include company scope
and an object-shaped `secret_ref` binding.
- Supplies invocation company scope to execution-workspace host-client
tests.
- Adds `enableApps` and `enableSmokeLab` to normalized settings
expectations.
- Removes the premature Apps toggle UI test introduced without its
later-stack implementation.
## Verification
- `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts
server/src/__tests__/plugin-execution-workspace-bridge.test.ts
server/src/__tests__/instance-settings-service.test.ts
ui/src/pages/InstanceExperimentalSettings.test.tsx` — 73 tests passed.
- `pnpm exec vitest run
packages/plugins/sdk/tests/host-client-factory.test.ts
server/src/__tests__/plugin-secrets-handler.test.ts
server/src/__tests__/instance-settings-routes.test.ts
ui/src/lib/instance-settings.test.ts` — 39 tests passed.
- `git diff --check` — passed.
## Risks
- Low risk: test-only changes with no production runtime, schema, API,
or UI behavior changes.
- The removed Apps toggle assertion should return in the later stacked
change that introduces the actual control.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.
Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 2/8 and focuses on database schema and
shared governance contracts
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: The governed access model needs additive persistence and
synchronized shared types before server enforcement can compile.
- Proposed solution: Adds migrations 0148–0169, tool-access and Smoke
Lab schema, shared types/validators/gallery helpers, and the minimal
compile-required contract consumers identified by boundary testing.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/01-demo-servers`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for migrations/validators; Greptile on every PR.
## What Changed
- Adds migrations 0148–0169, tool-access and Smoke Lab schema, shared
types/validators/gallery helpers, and the minimal compile-required
contract consumers identified by boundary testing.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck` — passed, including migration numbering and safety
checks
- `pnpm --filter @paperclipai/db test` — passed
- `pnpm --filter @paperclipai/shared test` — passed
## Risks
- Migration or contract mistakes could affect every upper layer; all
migrations are additive/idempotent and compile consumers are included in
this boundary.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 1/8 and focuses on fixture and demo MCP
servers
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack
## Linked Issues or Issue Description
- Related parity reference: #9534
- Problem: Developers need deterministic local MCP fixtures and visible
demo servers without pulling in the governed production runtime.
- Proposed solution: Adds the Google Sheets and KV demo MCP packages,
fixture catalog/servers, smoke harness, guide, and the root
smoke/typecheck registration hunks.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `master`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for fixture and smoke coverage; Greptile on every
PR.
## What Changed
- Adds the Google Sheets and KV demo MCP packages, fixture
catalog/servers, smoke harness, guide, and the root smoke/typecheck
registration hunks.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.
## Verification
- `pnpm typecheck`
- `pnpm --filter @paperclipai/google-sheets-mcp-server test` — 27 tests
passed
- `pnpm --filter @paperclipai/kv-demo-mcp-server test` — 12 tests passed
## Risks
- The new packages add dependencies that are intentionally not committed
to `pnpm-lock.yaml`, per repository policy.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Stack Coordination
- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents stream their run output live into the web UI, viewed per-run
in the `AgentDetail` transcript viewer
> - Browser tabs holding these views were sitting at 8–16 GB of memory
footprint while their JS heap stayed at ~256 MB — a 60×+ gap, meaning
the cost is in retained DOM / render objects, not JS objects
> - The `LogViewer` in `AgentDetail.tsx` kept every streamed
stdout/stderr line and structured event in unbounded React state and
rendered each into a rich DOM block (the default "nice" mode has no
virtualization), so a run streaming for hours grew an unbounded live DOM
tree
> - The sibling `useLiveRunTranscripts` hook (used by `IssueChatThread`)
already bounds its buffers and virtualizes; `LogViewer` bypassed it and
managed its own uncapped state — that inconsistency is the bug
> - This pull request caps the live buffers and bounds the live DOM
render, aligning `LogViewer` with the already-bounded path
> - The benefit is that long-lived streaming tabs no longer grow without
bound, cutting multi-GB tabs back to a bounded footprint
## Linked Issues or Issue Description
No public GitHub issue exists; describing inline per CONTRIBUTING.md →
"Link Issues or Describe Them In-PR", following the bug report template.
**What happened?**
Chrome's Task Manager showed multiple long-lived Paperclip tabs
(agent-run / task views) each consuming 8–16 GB of memory footprint,
while each tab's JS heap stayed at only ~150–256 MB. Memory grew
monotonically the longer a run streamed.
**Expected behavior**
A tab viewing a live agent run should hold a bounded amount of memory
regardless of how long the run streams.
**Steps to reproduce**
Open an agent run with a long-running / high-volume stream in
`AgentDetail`, leave the tab open while output streams for an extended
period, and watch the tab's memory footprint climb without bound in
Chrome's Task Manager.
**Paperclip version or commit**
`3991a19a` (branch `fix/agent-run-transcript-memory`, off `master`).
**Deployment mode**
Local dev (`pnpm dev`), web UI. Not adapter-specific — core UI bug in
the shared transcript viewer.
## What Changed
- Add `ui/src/lib/live-log-buffer.ts`: a pure `appendCapped(prev,
additions, max)` helper plus caps `MAX_LIVE_LOG_LINES=5000`,
`MAX_LIVE_EVENTS=2000`, and `LIVE_TRANSCRIPT_RENDER_LIMIT=1500`, with
rationale documented in the module.
- Add `ui/src/lib/live-log-buffer.test.ts`: 6 unit tests (append,
trim-to-cap, oversized batch, exact-cap, no-mutation, referential
bail-out).
- `ui/src/pages/AgentDetail.tsx` (`LogViewer`): route all four
live-append sites (WebSocket log / progress / event, plus the poll
fallback) through `appendCapped`, and pass
`limit={LIVE_TRANSCRIPT_RENDER_LIMIT}` to `RunTranscriptView` for live
runs so the "nice" view mounts only the most recent blocks.
- The terminated-run "Load more log" pagination is deliberately left
**uncapped** (guarded by `isLive`), so no historical output is lost —
older output remains on the server and reachable there.
## Verification
- `vitest run src/lib/live-log-buffer.test.ts` → 6/6 pass.
- Existing suites `RunTranscriptView.test.tsx`,
`AgentDetail.instructions.test.tsx`, `useLiveRunTranscripts.test.tsx` →
22/22 pass.
- `tsc -b` (UI) → clean.
- Manual/behavioral: live runs tail the last ~1500 blocks; terminated
runs still render full history via "Load more log". Follow-up planned to
profile before/after with the Chrome DevTools MCP.
## Risks
Low risk. Changes only bound **in-memory state for live runs**; the
terminated-run paginated path is untouched (still uncapped, guarded by
`isLive`). No API, schema, or persistence changes. Worst case for a live
run is that only the most recent 5000 lines / 1500 rendered blocks are
visible in the tab — which is the intended "tail" behavior, and full
history remains on the server.
## Model Used
- **Provider:** Anthropic, via the Claude Code CLI.
- **Model:** Claude Opus 4.8 (`claude-opus-4-8`).
- **Reasoning mode:** Extended thinking enabled.
- **Capabilities used:** tool use (shell execution, file editing),
sub-agent fan-out for the codebase memory sweep, and the Chrome DevTools
MCP for the diagnosis phase.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work (the ROADMAP "Memory" item is about company/agent
knowledge, unrelated to this browser-tab memory fix)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found among open PRs)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered and documented any risks above
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no user-facing docs affected; rationale is documented inline in
`live-log-buffer.ts`)
- [ ] All Paperclip CI gates are green (in progress at time of writing)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(the only open P2 was this missing template, which this update resolves;
awaiting re-review)
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Bumps [react-i18next](https://github.com/i18next/react-i18next) from
17.0.8 to 17.0.9.
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/i18next/react-i18next/blob/master/CHANGELOG.md">react-i18next's
changelog</a>.</em></p>
<blockquote>
<h2>17.0.9</h2>
<ul>
<li>fix: allow TypeScript 7 in the optional <code>typescript</code> peer
dependency range (<code>^5 || ^6 || ^7</code>). With
<code>typescript@7.0.2</code> in a project, <code>npm install</code>
failed with an <code>ERESOLVE</code> peer conflict. Fixes <a
href="https://redirect.github.com/i18next/react-i18next/issues/1927">#1927</a>,
thanks <a
href="https://github.com/andikapradanaarif"><code>@andikapradanaarif</code></a>.</li>
<li>fix(types): <code><Trans t={t} ns="ns" …></code>
with a <code>t</code> from <code>useTranslation(['ns'])</code> now
typechecks under TypeScript 7. TS7 intersects the <code>Ns</code>
inference candidates coming from the <code>t</code> prop (<code>readonly
['ns']</code>) and the <code>ns</code> prop (<code>'ns'</code>) into an
unsatisfiable <code>'ns' & readonly ['ns']</code>, where TS6
resolved them. The <code>ns</code> prop on <code>TransProps</code>,
<code>TransSelectorProps</code> and
<code>IcuTransWithoutContextProps</code> now also accepts a single
namespace out of an array-typed <code>Ns</code> (<code>Ns | (Ns extends
readonly (infer S extends string)[] ? S : never)</code>) — which matches
runtime behavior and is unchanged under TS5/TS6.</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="8b4a9ea139"><code>8b4a9ea</code></a>
17.0.9</li>
<li><a
href="422bab13d4"><code>422bab1</code></a>
fix: support typescript 7 — widen peer range and fix Trans ns inference
under...</li>
<li><a
href="6e18aa95b5"><code>6e18aa9</code></a>
README: mention npx i18next-cli localize as the zero-to-localized
path</li>
<li>See full diff in <a
href="https://github.com/i18next/react-i18next/compare/v17.0.8...v17.0.9">compare
view</a></li>
</ul>
</details>
<br />
[](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)
Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.
[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)
---
<details>
<summary>Dependabot commands and options</summary>
<br />
You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)
</details>
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
[//]: # (dependabot-start)
⚠️ **Dependabot is rebasing this PR** ⚠️
Rebasing might not happen immediately, so don't worry if this takes some
time.
Note: if you make any changes to this PR yourself, they will take
precedence over the rebase.
---
[//]: # (dependabot-end)
Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.8 to
3.4.12.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/cure53/DOMPurify/releases">dompurify's
releases</a>.</em></p>
<blockquote>
<h2>DOMPurify 3.4.12</h2>
<ul>
<li>Fixed an issue where a hook would not get called for custom
elements, thanks <a
href="https://github.com/Rikuxx0"><code>@Rikuxx0</code></a></li>
<li>Hardened the handling of hooks removing elements, <a
href="https://github.com/mkrause-bee360"><code>@mkrause-bee360</code></a></li>
<li>Added support for a few new SVG attributes, thanks <a
href="https://github.com/cbn-falias"><code>@cbn-falias</code></a> &
<a
href="https://github.com/Develop-KIM"><code>@Develop-KIM</code></a></li>
<li>Hardened the handling of declarative partial updates</li>
<li>Updated the documentation is several spots, README, wiki, etc.</li>
<li>Bumped several dependencies where possible</li>
</ul>
<h2>DOMPurify 3.4.11</h2>
<ul>
<li>Fixed an issue with a leaky config for hooks via
<code>setConfig</code>, thanks <a
href="https://github.com/trace37labs"><code>@trace37labs</code></a></li>
<li>Bumped vulnerable development dependencies to arrive at plain 0 with
<code>npm audit</code></li>
<li>Updated the <code>osv-scanner</code> suppression list as no
vulnerable dependencies are left for now</li>
<li>Updated up the linting tool-chain and removed now-redundant lint
directives</li>
<li>Updated the documentation is several spots, README, wiki, etc.</li>
<li>Bumped several dependencies where possible</li>
</ul>
<h2>DOMPurify 3.4.10</h2>
<ul>
<li>Refactored codebase for clarity: extracted the public type
declarations into <code>types.ts</code></li>
<li>Decomposed the three largest sanitizer functions into focused
helpers</li>
<li>Removed duplicated defaults and dead branches, consolidated
<code>SAFE_FOR_TEMPLATES</code> scrubbing into single shared path</li>
<li>Improved per-node performance by hoisting the mXSS probe regexes and
testing <code>textContent</code> before <code>innerHTML</code></li>
<li>Added a deterministic micro-benchmark harness (<code>npm run
bench</code>) with a <code>--compare</code> mode</li>
<li>Reduced CI cost by running the full three-engine browser suite once
per PR</li>
<li>Refreshed the <code>demos/</code> folder so every demo runs again,
and added a SVG-via-<code><img></code> demo</li>
<li>Documented the bench and <code>test:happydom</code> scripts in the
README</li>
<li>Completed the Attack Classes & Bypass History wiki page</li>
<li>Bumped several dependencies where possible</li>
</ul>
<h2>DOMPurify 3.4.9</h2>
<ul>
<li>Further improved the handling of Trusted Types config options,
thanks <a
href="https://github.com/offset"><code>@offset</code></a></li>
<li>Further improved the handling of <code>IN_PLACE</code> sanitization,
thanks <a
href="https://github.com/mozfreddyb"><code>@mozfreddyb</code></a></li>
<li>Added more test coverage for <code>IN_PLACE</code> and Trusted Types
related usage</li>
<li>Bumped several dependencies where possible</li>
<li>Updated README and wiki with more accurate documentation &
attack samples</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="a9ca1e5374"><code>a9ca1e5</code></a>
release: 3.4.12 (<a
href="https://redirect.github.com/cure53/DOMPurify/issues/1537">#1537</a>)</li>
<li><a
href="0cae518740"><code>0cae518</code></a>
release: 3.4.11 (<a
href="https://redirect.github.com/cure53/DOMPurify/issues/1494">#1494</a>)</li>
<li><a
href="6ee5716f83"><code>6ee5716</code></a>
release: 3.4.10 (<a
href="https://redirect.github.com/cure53/DOMPurify/issues/1478">#1478</a>)</li>
<li><a
href="52102472d4"><code>5210247</code></a>
release: 3.4.9 (<a
href="https://redirect.github.com/cure53/DOMPurify/issues/1459">#1459</a>)</li>
<li>See full diff in <a
href="https://github.com/cure53/DOMPurify/compare/3.4.8...3.4.12">compare
view</a></li>
</ul>
</details>
<br />
[](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)
Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.
[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)
---
<details>
<summary>Dependabot commands and options</summary>
<br />
You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)
</details>
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps
[@clack/prompts](https://github.com/bombshell-dev/clack/tree/HEAD/packages/prompts)
from 0.10.1 to 0.11.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/bombshell-dev/clack/releases">@clack/prompts's
releases</a>.</em></p>
<blockquote>
<h2><code>@clack/prompts</code><a
href="https://github.com/0"><code>@0</code></a>.11.0</h2>
<h3>Minor Changes</h3>
<ul>
<li>07ca32d: Reverted a change where placeholders were being set as
values on return.</li>
</ul>
<h3>Patch Changes</h3>
<ul>
<li>Updated dependencies [07ca32d]
<ul>
<li><code>@clack/core</code><a
href="https://github.com/0"><code>@0</code></a>.5.0</li>
</ul>
</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/bombshell-dev/clack/blob/@clack/prompts@0.11.0/packages/prompts/CHANGELOG.md">@clack/prompts's
changelog</a>.</em></p>
<blockquote>
<h2>0.11.0</h2>
<h3>Minor Changes</h3>
<ul>
<li>07ca32d: Reverted a change where placeholders were being set as
values on return.</li>
</ul>
<h3>Patch Changes</h3>
<ul>
<li>Updated dependencies [07ca32d]
<ul>
<li><code>@clack/core</code><a
href="https://github.com/0"><code>@0</code></a>.5.0</li>
</ul>
</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="737f172569"><code>737f172</code></a>
[ci] release (<a
href="https://github.com/bombshell-dev/clack/tree/HEAD/packages/prompts/issues/325">#325</a>)</li>
<li><a
href="07ca32dcfc"><code>07ca32d</code></a>
fix: revert placeholder-on-return change (<a
href="https://github.com/bombshell-dev/clack/tree/HEAD/packages/prompts/issues/324">#324</a>)</li>
<li>See full diff in <a
href="https://github.com/bombshell-dev/clack/commits/@clack/prompts@0.11.0/packages/prompts">compare
view</a></li>
</ul>
</details>
<br />
[](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)
Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.
[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)
---
<details>
<summary>Dependabot commands and options</summary>
<br />
You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)
</details>
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps
[@storybook/react-vite](https://github.com/storybookjs/storybook/tree/HEAD/code/frameworks/react-vite)
from 10.4.6 to 10.5.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/storybookjs/storybook/releases">@storybook/react-vite's
releases</a>.</em></p>
<blockquote>
<h2>v10.5.0</h2>
<h2>10.5.0</h2>
<blockquote>
<p><em>Foundational changes for new AI workflows</em></p>
</blockquote>
<p>Storybook 10.5 contains hundreds of fixes and improvements:</p>
<ul>
<li>⚡️ Angular-vite framework: Modern, fast dev, docs, and test
(preview)</li>
<li>🌈 Vitest initialGlobals: Test across themes, viewports, locales</li>
<li>🤖 Agentic review: AI-curated visual changesets and search results
(experimental)</li>
<li>⚛️ React docgen service: Unified metadata across MCP, Docs, and
Controls (experimental)</li>
<li>🧑💻 Claude / Codex plugins: One-click ADE integration
(experimental)</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/storybookjs/storybook/blob/next/CHANGELOG.md">@storybook/react-vite's
changelog</a>.</em></p>
<blockquote>
<h2>10.5.0</h2>
<blockquote>
<p><em>Foundational changes for new AI workflows</em></p>
</blockquote>
<p>Storybook 10.5 contains hundreds of fixes and improvements:</p>
<ul>
<li>⚡️ Angular-vite framework: Modern, fast dev, docs, and test
(preview)</li>
<li>🌈 Vitest initialGlobals: Test across themes, viewports, locales</li>
<li>🤖 Agentic review: AI-curated visual changesets and search results
(experimental)</li>
<li>⚛️ React docgen service: Unified metadata across MCP, Docs, and
Controls (experimental)</li>
<li>🧑💻 Claude / Codex plugins: One-click ADE integration
(experimental)</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="9dafcd22ed"><code>9dafcd2</code></a>
Bump version from "10.5.0-beta.2" to "10.5.0" [skip
ci]</li>
<li><a
href="448db85e65"><code>448db85</code></a>
Bump version from "10.5.0-beta.1" to "10.5.0-beta.2"
[skip ci]</li>
<li><a
href="a4ce9790a7"><code>a4ce979</code></a>
Bump version from "10.5.0-beta.0" to "10.5.0-beta.1"
[skip ci]</li>
<li><a
href="f0bf138a0a"><code>f0bf138</code></a>
Bump version from "10.5.0-alpha.11" to
"10.5.0-beta.0" [skip ci]</li>
<li><a
href="fac05a5741"><code>fac05a5</code></a>
Bump version from "10.5.0-alpha.10" to
"10.5.0-alpha.11" [skip ci]</li>
<li><a
href="4057c4169f"><code>4057c41</code></a>
Bump version from "10.5.0-alpha.9" to
"10.5.0-alpha.10" [skip ci]</li>
<li><a
href="da84210b49"><code>da84210</code></a>
Bump version from "10.5.0-alpha.8" to
"10.5.0-alpha.9" [skip ci]</li>
<li><a
href="c347410b9f"><code>c347410</code></a>
Bump version from "10.5.0-alpha.7" to
"10.5.0-alpha.8" [skip ci]</li>
<li><a
href="d16ab3008a"><code>d16ab30</code></a>
React: Align react-docgen versions</li>
<li><a
href="c9a1ac9f72"><code>c9a1ac9</code></a>
Bump version from "10.5.0-alpha.6" to
"10.5.0-alpha.7" [skip ci]</li>
<li>Additional commits viewable in <a
href="https://github.com/storybookjs/storybook/commits/v10.5.0/code/frameworks/react-vite">compare
view</a></li>
</ul>
</details>
<br />
[](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)
Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.
[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)
---
<details>
<summary>Dependabot commands and options</summary>
<br />
You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)
</details>
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
## Thinking Path
> - Paperclip is the open source control plane people use to manage AI
agents for work.
> - Local coding adapters currently spawn their CLI processes directly
on the Paperclip host.
> - CLI-native approval and sandbox flags do not provide a reliable host
filesystem or network boundary.
> - An agent can therefore inspect unrelated host files or fetch
external material when an operator needs stronger isolation.
> - The confinement must stay opt-in so existing local adapter behavior
does not change unexpectedly.
> - This pull request adds a shared Linux Bubblewrap spawn layer for
workspace filesystem and deny/allowlist network scopes.
> - The benefit is enforceable defense in depth around Codex and Claude
local runs while preserving explicit provider connectivity.
## Linked Issues or Issue Description
### What happened?
`codex_local` and `claude_local` processes could read arbitrary host
paths and make unrestricted outbound network requests because Paperclip
did not impose a spawn-level boundary.
### Expected behavior
Operators can opt into a workspace-only filesystem view and either deny
network egress or allow exact provider/API hosts, independently of CLI
approval flags.
### Steps to reproduce
1. Run current `master` on Linux and configure a Codex or Claude local
adapter.
2. Ask the agent to read a canary file outside its active workspace.
3. Ask the agent to `curl` a public host.
4. Observe that both operations succeed without a Paperclip-level
confinement option.
### Environment
- Paperclip commit: `c36f1a4af` / current `master` base.
- Deployment mode: Linux local dev or self-hosted server.
- Installation: built from source.
- Adapters: Codex and Claude Code.
- Database: not related.
## What Changed
- Added a shared Bubblewrap process wrapper with opt-in
`filesystemScope: "workspace"`, managed/extra path mounts, private
`/tmp`, and Linux-only validation.
- Added `networkScope: "deny" | "allowlist"`; both use a private network
namespace, while allowlist mode exposes an exact-host HTTP(S) proxy over
a Unix-socket bridge.
- Wired Codex and Claude local CLI execution through the wrapper and
forced scoped auto runs onto the CLI lane because ACP processes are not
covered.
- Added unit and gated Bubblewrap canaries for outside-file denial,
workspace writes, direct network denial, allowlisted forwarding, and
rejected destinations.
- Documented both scopes, provider allowlist examples, Bubblewrap
requirements, and default-off behavior.
## Verification
- `pnpm exec vitest run
packages/adapter-utils/src/local-process-sandbox.test.ts
packages/adapters/codex-local/src/server/acp.test.ts
packages/adapters/claude-local/src/server/acp.test.ts` — 34 passed, 4
gated Bubblewrap tests skipped by default.
- `pnpm --filter @paperclipai/adapter-utils typecheck`
- `pnpm --filter @paperclipai/adapter-codex-local typecheck`
- `pnpm --filter @paperclipai/adapter-claude-local typecheck`
- `pnpm --filter @paperclipai/adapter-utils build`
- `pnpm --filter @paperclipai/adapter-codex-local build`
- `pnpm --filter @paperclipai/adapter-claude-local build`
- Attempted the gated tests with a vendored Bubblewrap binary; this
container blocks unprivileged namespace setup (`setting up uid map:
Permission denied` / loopback `RTM_NEWADDR: Operation not permitted`),
so kernel-level execution remains for CI or a namespace-enabled Linux
host.
## Risks
- Bubblewrap must be installed and unprivileged user/mount/network
namespaces must be enabled on the host; scoped runs fail clearly if the
prerequisite is missing.
- Allowlist mode depends on the coding CLI honoring standard
`HTTP_PROXY` / `HTTPS_PROXY` variables; custom providers must list every
required exact hostname and port.
- Exact-host allowlists intentionally reject wildcards, which is safer
but may require operators to enumerate multi-host provider setups.
- No behavior changes unless an operator enables `filesystemScope` or
`networkScope`.
> This aligns with the ROADMAP direction toward safer remote and
sandboxed agent environments and does not duplicate an open PR or issue
found in the repository search.
## Model Used
- OpenAI GPT-5.5 (`gpt-5.5`) via Codex CLI, with reasoning, repository
tool use, shell execution, code editing, and test execution. The serving
context-window size is not exposed to the agent.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane people use to manage AI
agents for work
> - Local agent execution uses isolated git worktrees with
worktree-specific config, environment, storage, and ports
> - Legacy worktree repair and runtime-port persistence must mutate only
the worktree they are serving
> - A leaked ambient `PAPERCLIP_IN_WORKTREE=true` could be combined with
config resolution pointing at the default instance
> - The server test suite reproduced that combination and repeatedly
rewrote the live default instance `.env` with its old fixture name
> - Existing PR #3071 guards configs under Paperclip home, but does not
require the target itself to attest worktree ownership and does not
cover runtime-port persistence
> - This pull request requires both a worktree config layout and
target-local persisted worktree attestation before either writer adopts
the target
> - The benefit is that ambient process state can never turn the main
instance into a worktree on its next restart
## Linked Issues or Issue Description
Related implementation: Refs #3071
**Pre-submission checklist**
- [x] Searched open and closed issues and pull requests; #3071 is the
only direct related implementation.
- [x] Reproduced on current `master` before applying the fix.
- [x] Confirmed the mutation originates in Paperclip's worktree config
repair path.
**What happened?**
A process with leaked `PAPERCLIP_IN_WORKTREE=true` could resolve
`PAPERCLIP_CONFIG` to the default instance and cause worktree repair to
rewrite that instance's `.env`. The recurring trigger was
`server/src/__tests__/worktree-config.test.ts`: an ambient config path
from the developer shell survived into a test whose fixture worktree
name was `PAP-884-ai-commits-component`, explaining the stale name
repeatedly written to the live file.
**Expected behavior**
Worktree repair and worktree runtime-port persistence must mutate a
target only when that target is independently provisioned and persisted
as a worktree. Ambient environment flags alone must never authorize
writes to the default instance or a normal repository-local `.paperclip`
config.
**Steps to reproduce on unpatched `master`**
1. Export `PAPERCLIP_CONFIG` pointing to a default instance config and
set `PAPERCLIP_IN_WORKTREE=true`.
2. Run `server/src/__tests__/worktree-config.test.ts` from that shell.
3. Observe that the default instance `.env` is rewritten with the test
fixture's worktree marker and name.
**Environment**
- Version: `master` at `e4e12bfb8`
- Deployment/install: local source checkout with pnpm
- Adapter: not adapter-specific; core server config
- Database/access context: not applicable
- OS: Linux
**Privacy**
- [x] All paths and values in this description are generic and contain
no credentials or personally identifying data.
## What Changed
- Reject config targets unless their parent directory is the
worktree-specific `.paperclip` layout.
- Require the target's own persisted `.env` to declare
`PAPERCLIP_IN_WORKTREE=true` before repair or runtime-port persistence
can mutate it.
- Scrub ambient `PAPERCLIP_*` variables before every worktree-config
test so developer-machine exports cannot escape test isolation.
- Add regressions for default-instance config poisoning, runtime-port
persistence, and unattested repository-local `.paperclip` targets.
- Preserve valid provisioned worktree behavior by adding persisted
worktree markers to the existing positive fixtures.
## Verification
- `NODE_ENV=test pnpm --filter @paperclipai/server exec vitest run
src/__tests__/worktree-config.test.ts` — 12 tests passed.
- Branch is based directly on current `origin/master`; only two server
files changed.
- No `pnpm-lock.yaml`, workflow, migration, UI, or generated asset
changes.
## Risks
- Low risk: the new guard intentionally refuses repair for targets that
lack provisioning evidence.
- A manually assembled worktree that sets only ambient flags but never
writes its worktree marker will no longer be auto-repaired; the
supported provisioning path already writes that marker.
- No schema, API, migration, or user-facing command changes.
> This is a focused correctness fix and does not overlap with planned
core work in `ROADMAP.md`.
## Model Used
- Implementation and root-cause investigation: Anthropic Claude through
the `claude_local`/Claude Code runtime, reported by the producing agent
as “Claude Fable 5”; the runtime did not expose a more specific provider
model ID or context-window value. Capabilities used: extended reasoning,
shell tool use, code editing, and test execution.
- PR preparation and verification: OpenAI Codex CLI runtime; the harness
did not expose the exact underlying model ID or context-window value.
Capabilities used: repository inspection, shell tool use, Git/GitHub
operations, and test execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used with all model details exposed by
the runtimes
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
#3071 above
- [x] I have described the issue in-PR following the bug report template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket ID
- [x] I have run the focused tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have assessed documentation impact; no documentation change is
required for this internal guard
- [x] I have considered and documented risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is an open-source agentic AI management platform that
ships anonymous usage telemetry to understand product health and
adoption.
> - The telemetry system has a generated contract
(`packages/shared/src/telemetry/generated/paperclip-telemetry.ts`) that
types every first-party event the product emits.
> - When a product change needs a new first-party event that is not yet
in the generated contract, contributors had no public workflow
explaining how to propose an event or later promote it into the contract
once accepted.
> - The gap leads to confusion at call sites: contributors either skip
tracking entirely or emit untyped events that bypass the privacy and
governance safeguards built into the contract.
> - This PR fills that gap by adding `doc/TELEMETRY_WORKFLOW.md` — a
public contributor guide that covers the full propose → promote
lifecycle: the `@ts-expect-error -- proposed-telemetry(...)` marker, the
canonical multi-line `track()` shape, the TS2578 expiry signal, and the
look-up-by-event-name promotion step.
> - `packages/shared/src/telemetry/README.md` gains a cross-reference so
readers of the data-contract doc can find the workflow guide without
searching.
> - `README.md` gains a one-line pointer in the telemetry section so the
workflow is discoverable from the project entry point.
## Linked Issues or Issue Description
No pre-existing public GitHub issue covers this doc gap. Inline
description:
**Problem:** Contributors adding product telemetry for events not yet in
the generated contract have no documented workflow. The
`@ts-expect-error -- proposed-telemetry(...)` pattern exists in the
codebase but is undocumented, leading to inconsistent usage and missing
adoption signals.
**Solution:** A new public guide (`doc/TELEMETRY_WORKFLOW.md`) documents
the marker format, the recommended multi-line `track()` shape that
preserves the TS2578 expiry signal, the dimension rules, and the
promotion checklist. Cross-references are added to
`packages/shared/src/telemetry/README.md` and the top-level `README.md`.
No related open PRs found.
## What Changed
- **New file `doc/TELEMETRY_WORKFLOW.md`** — public contributor guide
for the propose/promote lifecycle: marker format, canonical multi-line
example, TS2578 single-line trap, dimension rules, and step-by-step
promotion checklist.
- **`packages/shared/src/telemetry/README.md`** — added one-line
cross-reference pointing at `doc/TELEMETRY_WORKFLOW.md` for proposed
events not yet in the generated contract.
- **`README.md`** — added one-line pointer in the telemetry section so
the new workflow guide is reachable from the top-level project entry.
## Verification
This is a docs-only change. Verification steps:
1. Open `doc/TELEMETRY_WORKFLOW.md` and confirm it contains:
- The `^[a-z0-9][a-z0-9._:-]{1,63}$` event-name grammar.
- The exact `// @ts-expect-error -- proposed-telemetry(<issue>):
<rationale>` marker.
- The multi-line `client.track()` copy-paste example (directive on line
before event-name string).
- The explanation of the TS2578 single-line trap and the
look-up-by-event-name step.
- The promotion checklist in the "Promote An Event" section.
2. Confirm `packages/shared/src/telemetry/README.md` cross-references
`doc/TELEMETRY_WORKFLOW.md`.
3. Confirm the `README.md` telemetry section links to
`doc/TELEMETRY_WORKFLOW.md`.
## Risks
Low risk — docs-only change. No runtime behavior, schema, or existing
telemetry emission is affected. The risk is that the guidance could
diverge from the actual enforcement in the codebase over time; mitigated
by linking to the generated contract and keeping the doc in the same
repo.
## Model Used
Claude — `claude-sonnet-4-6` (Anthropic Claude Sonnet 4.6). Tool use
enabled. Extended context. Produced via the Paperclip agentic workflow
with `Co-authored-by: Paperclip <noreply@paperclip.ing>`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - It emits telemetry events to understand product usage — registered
event names are gated by a generated `PAPERCLIP_EVENTS` registry so the
client only enqueues known, schema-approved events
> - When product teams want to instrument a new behaviour, they must
first register the event name — but schema registration is a
commit-and-release cycle, which creates friction in fast-moving product
iterations
> - A proposal lane is needed: let developers mark a `track()` call with
a typed `@ts-expect-error` proposal marker so the event name can be
reviewed and tracked in CI before the schema is formally registered
> - The existing client had no guard against unregistered event names,
so any call with an out-of-registry name (or a prototype-inherited key)
would silently enter the queue, state, and network flush path
> - This PR adds an `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard
at the entry point of `track()` to swallow unregistered calls before any
side effects, adds `scripts/extract-proposed-events.mjs` to scan source
for proposal markers and emit a v2 JSON manifest with provenance and
rationale, and documents the complete proposal workflow
> - The benefit is that new instrumentation can be proposed and reviewed
in code without touching the registered schema, and tooling can surface
missing rationale before events graduate to stable
## Linked Issues or Issue Description
No existing GitHub issue covers this change. This PR introduces a new
feature.
**Feature motivation:** Paperclip's telemetry schema is intentionally
stable — registered event names are code-generated and gated. Product
engineers who want to instrument a new behaviour today must land a
schema change first, creating a two-step process that slows iteration. A
proposal lane lets developers write the instrumentation call ahead of
schema registration, protected by a compile-time `@ts-expect-error`
marker that an extractor script can surface for review. This PR
implements both the client-side safety gate and the extraction tooling.
Refs: #9518 (closed predecessor — docs-only; this PR supersedes it with
the full implementation)
## What Changed
- Added `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard at the top of
`TelemetryClient.track()`: unregistered event names (including
prototype-inherited keys) are now swallowed before any state, queue, or
network operation
- Added `scripts/extract-proposed-events.mjs`: scans TypeScript source
for `@ts-expect-error -- proposed-telemetry(<issue>): <rationale>`
markers; emits a v2 JSON manifest per proposed event including name,
rationale, provenance (repo-relative file + line), and a
`rationale_missing` flag for CI enforcement
- Added `scripts/extract-proposed-events.test.mjs`: test suite covering
marker parsing, multi-line markers, path validation, out-of-repo
rejection, and the v2 schema output contract
- Added `doc/TELEMETRY_WORKFLOW.md`: documents the proposal workflow,
the canonical multi-line marker example, rationale requirements, and how
to graduate a proposed event to stable schema
- Updated `packages/shared/src/telemetry/README.md`: added "Proposed
Events" section to the Telemetry Data Contract per the contributing
guide requirement for telemetry changes
## Verification
Run all of the following from the repo root:
```sh
# Extractor unit tests
node --test scripts/extract-proposed-events.test.mjs
# Telemetry client + types tests
pnpm exec vitest run --config vitest.config.ts \
src/telemetry/client.test.ts src/telemetry/client-types.test.ts \
--reporter=verbose
# (run from packages/shared)
# Type-check
pnpm --filter @paperclipai/shared typecheck
# Smoke-run the extractor in local-test mode
node scripts/extract-proposed-events.mjs --ref local-test
```
All four commands pass locally.
## Risks
- **Silent drop on unregistered events:** The `Object.hasOwn` guard
fails closed — any event name not in `PAPERCLIP_EVENTS` is silently
dropped. If the generated registry is missing an event that was
previously tracked, those calls will be silently lost. Mitigation: the
extractor script surfaces proposed events that need registration; the
TypeScript type system already enforces `TelemetryEventName ⊆
PAPERCLIP_EVENTS` at compile time.
- **Extractor is read-only:** `extract-proposed-events.mjs` reads source
and emits JSON; it does not modify any files. No runtime or schema risk.
- Overall risk: **low**. The guard is additive and defensive; the
extractor and docs are additive only.
## Model Used
- Provider: Anthropic
- Model ID: `claude-sonnet-4-6`
- Context window: 200 K tokens
- Capabilities: tool use, extended context, code generation
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The release subsystem uses GitHub Actions dry-run previews so
maintainers can validate stable and canary release behavior before
publishing.
> - Stable release publishing needs a release notes gate so real
`latest` publishes never happen without authored notes.
> - The same gate was also blocking stable dry-run previews, which made
release QA fail before publish-sensitive work could be previewed.
> - This pull request narrows the notes requirement to real non-dry-run
stable publishes.
> - The benefit is that stable dry-run dispatches can preview the
release without same-day notes, while real stable publishes remain
protected.
## Linked Issues or Issue Description
No public GitHub issue exists for this release-QA blocker, so the bug is
described inline below.
### What happened?
Stable dry-run release previews fail when the same-day release notes
file is absent. The release script runs the stable notes-file gate
before release preview work even when `--dry-run` is set.
### Expected behavior
`./scripts/release.sh stable --dry-run` should preview the stable
release without requiring `releases/vYYYY.MDD.P.md`. Real non-dry-run
stable publishes must still fail before build/publish work starts when
the notes file is missing.
### Steps to reproduce
1. Check out current `master` before this fix.
2. Ensure the computed same-day stable release notes file does not exist
under `releases/`.
3. Run `./scripts/release.sh stable --skip-verify --dry-run`.
4. Observe that the script exits with `stable release notes file is
required` instead of reaching the release preview plan.
### Paperclip version or commit
Reproduced on `master` at `9a1d4b7983dfd50e8eb40ee9770e44999d405f60`.
### Deployment mode
Built from source / GitHub Actions release workflow.
## What Changed
- Narrowed the stable release notes gate to `channel=stable` and
`dry_run=false`.
- Updated the release script usage note to say the notes file is
required for non-dry-run stable releases.
- Added a targeted Node test covering dry-run allowed behavior and
non-dry-run blocked behavior.
- Stubbed release fixture registry-version checks so the test isolates
the notes gate without hitting npm.
## Verification
- `node --test scripts/__tests__/release-dry-run-notes.test.mjs` passed
with 2/2 subtests.
- `bash -n scripts/release.sh` exited 0.
## Risks
Low risk. The behavior change only relaxes the notes-file gate for
stable dry-runs. The new test verifies real non-dry-run stable publish
still fails before build/publish work starts when notes are missing.
## Model Used
OpenAI Codex, GPT-5-based coding agent, with shell/tool execution in a
local repository workspace.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane people use to manage
AI-agent companies
> - Operators need to identify the exact build running from the
persistent account menu
> - Formal releases already have a concise public version, but source
builds include a long derived version string
> - The derived version identifies a commit but does not expose the
source branch or a direct path to inspect the code
> - Server Git metadata is auth-sensitive, so the UI must also refresh
it when the current session changes
> - This pull request shows linked branch and commit metadata for source
builds while preserving `v<version>` for formal releases
> - The benefit is faster build diagnosis with correct metadata across
sign-in and sign-out transitions
## Linked Issues or Issue Description
### Pre-submission checklist
- [x] I searched existing open and closed issues and found no duplicate
implementing this exact account-menu behavior.
- [x] The behavior reproduces on `master`.
- [x] The behavior originates in Paperclip's core UI, not an adapter,
provider, or local configuration.
### What happened?
Source builds displayed the full derived version, such as
`2026.626.0+58.git.518fc71ce`, without linking the operator to the
corresponding source branch or commit.
### Expected behavior
Source builds should show the concise branch and short commit SHA with
links to GitHub, while formal releases should continue showing their
public version. Auth transitions should refresh the health metadata that
supplies those Git details.
### Steps to reproduce
1. Run Paperclip from a commit after a release tag.
2. Open the account menu.
3. Inspect the build label beneath the user identity.
4. Sign in or out and reopen the menu.
### Paperclip version or commit
Any source build whose server version uses the
`<version>+<count>.git.<sha>[.dirty]` format.
### Deployment mode
Local dev (`pnpm dev`) or authenticated deployments.
### Installation method
Built from source.
## What Changed
- Detect source-derived version strings and render the source branch
plus seven-character commit SHA in `SidebarAccountMenu`.
- Link source branches and commits to the canonical
`paperclipai/paperclip` GitHub repository.
- Extend server Git metadata with the full SHA and expose it through
health/OpenAPI contracts.
- Refresh auth-sensitive health metadata after sign-in and every
sign-out entry point.
- Preserve the existing `v<version>` label for formal releases and add
focused regression coverage.
## Verification
- `pnpm --filter @paperclipai/ui exec vitest run src/pages/Auth.test.tsx
src/components/SidebarAccountMenu.test.tsx
src/components/SidebarServerInfo.test.tsx`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/health.test.ts src/__tests__/server-info.test.ts`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm check:token-gates`
- `git diff --check public/master...HEAD`
## Risks
- Low risk: formal release rendering retains the existing fallback
behavior when the source-version pattern does not match.
- Source links assume the build came from the canonical public
repository; fork-only branches or commits may not resolve there.
- Health metadata is invalidated after auth transitions, adding one
bounded refetch so the displayed Git details match the new session.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex using GPT-5.4 with medium reasoning, repository/tool
access, shell execution, and code editing; context-window size was not
exposed by the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane people use to manage
AI-agent companies
> - Budgets and spend telemetry are control-plane safety features, not
just reporting
> - Local Codex and Claude adapters can execute through either ACP or
their native CLI engines
> - The ACP lane records usage and reported cost, but CLI JSON output
often reports tokens without a price
> - The CLI lane was either losing per-run usage semantics or coercing
missing cost to zero, making real usage indistinguishable from a
genuinely free run
> - This pull request preserves CLI usage as per-run totals and records
token-bearing runs without a reported price as explicitly unpriced
ledger events
> - The benefit is accurate usage accounting and a visible pricing gap
instead of silently misleading zero-cost telemetry
## Linked Issues or Issue Description
Refs #9471
Refs #9230
**Bug description**
A `codex_local` run using the CLI engine can emit a final
`turn.completed` event with millions of input tokens and tens of
thousands of output tokens while the agent's spend ledger remains
indistinguishable from a true zero-usage, zero-cost run. Claude CLI
output has the same missing-price edge case.
**Expected behavior**
Token-bearing CLI runs should persist their usage. If the adapter
reports a price, the ledger should record it as reported; if the CLI
reports usage but no price, the ledger should explicitly mark the event
as unpriced rather than silently treating missing price data as a
reported `$0` cost.
**Reproduction shape**
1. Configure `codex_local` with `engine: cli`.
2. Run a task that produces a `turn.completed` usage payload.
3. Observe token usage in the run stream.
4. Before this change, missing price data is represented as ordinary
zero-cost spend and the CLI usage basis is not consistently propagated.
## What Changed
- Mark Codex and Claude native CLI usage totals as `per_run` and
propagate that basis through success and failure results.
- Stop coercing missing Claude CLI cost to `0`.
- Add `cost_status` to cost events with `reported` and `unpriced`
values, including an idempotent migration and shared validation/types.
- Persist token-bearing runs without a reported price as `unpriced`
ledger events while retaining zero cents until an authoritative price
exists.
- Add parser, execute-path, heartbeat-accounting, and cost-service
regression coverage for both local CLI adapters.
- Document the cost-status invariant and CLI accounting behavior.
## Verification
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/parse.test.ts
packages/adapters/claude-local/src/server/parse.test.ts
server/src/__tests__/codex-local-execute.test.ts
server/src/__tests__/claude-local-execute.test.ts
server/src/__tests__/heartbeat-cost-accounting.test.ts
server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests
passed.
- `pnpm --filter @paperclipai/shared typecheck`
- `pnpm --filter @paperclipai/db typecheck` — includes migration
numbering and safety checks.
- `pnpm --filter @paperclipai/adapter-codex-local typecheck`
- `pnpm --filter @paperclipai/adapter-claude-local typecheck`
- `pnpm --filter @paperclipai/server typecheck`
## Risks
- Existing cost rows default to `reported`, preserving current
interpretation; only new token-bearing events with absent cost are
marked `unpriced`.
- This change does not invent model pricing. Budget hard stops still
cannot charge an unknown amount, but operators and evals can now
distinguish missing pricing from a genuinely reported zero cost.
- Consumers that enumerate cost-event fields should tolerate the
additive `costStatus` field.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use
and code execution; default reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Its PR CI runs the general-server vitest lane pinned to
`maxWorkers=1` and sharded across 3 runners (introduced in #8360)
> - Suites were assigned to shards round-robin by sorted file index, so
shard test time was unbalanced: a recent PR run split 73s / 153s / 115s,
and the heaviest shard made "General tests (server 2/3)" the slowest
check in the whole workflow at 314s wall
> - The slowest shard sets the lane's wall time, so unbalanced
partitions waste the other two runners and stretch the PR critical path
> - This pull request replaces the round-robin assignment with a
deterministic longest-processing-time partition weighted by a checked-in
per-suite duration manifest
> - The benefit is near-even shard weights (projected 113s / 113s / 113s
with the current manifest), taking roughly 40s off the PR critical path
with no reduction in coverage
## Linked Issues or Issue Description
- Refs #8360 (introduced the 3-way general-server sharding this PR
rebalances)
- No public issue exists. Problem: the general-server test lane's
round-robin shard assignment ignores per-suite duration, so one shard
can carry multiple 30s+ suites while another finishes in half the time;
the slowest shard alone determines the check's wall time.
## What Changed
- `scripts/general-server-shard.mjs` (new): manifest loader and
deterministic LPT (longest-processing-time) partitioner; suites missing
from the manifest get the median recorded weight, and a missing or
malformed manifest degrades to uniform weights so the lane never fails
on stale data
- `scripts/general-server-shard-durations.json` (new): per-suite
duration manifest sampled from a real PR run (240 suites); the
`$comment` field documents how to regenerate it
- `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and
`--dry-run`) now use the balanced partition instead of index round-robin
- `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests
covering skew-balance vs round-robin, determinism, median fallback for
unlisted suites, malformed-manifest degradation, manifest coverage of
the current suite set, and real-partition balance
- `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`:
hardened the `afterEach` sweep — post-run bookkeeping (run-event
records, follow-up wake scheduling) can still insert rows briefly after
a run reaches a terminal status, and a late insert landing between the
`agent_wakeup_requests` and `agents` deletes failed teardown with a
foreign-key violation on the first CI attempt of this PR; the sweep now
retries so a late background write cannot take down the shard
- `release-verify.yml` shares the same runner script and inherits the
balancing with no workflow change
## Verification
- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9
pass (run against current master)
- `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts`
— 6/6 pass against embedded Postgres with the hardened teardown
- `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2
pass
- `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag
shows every suite assigned exactly once across the 3 shards, with
projected weights ~113s each
## Risks
- Low risk: partition changes which runner executes which suite, not
what runs; a completeness test asserts every suite is assigned to
exactly one shard
- The duration manifest will drift as suites are added/changed; unlisted
suites get the median weight and a coverage test flags when the manifest
covers less than half the suite set, so drift degrades balance
gracefully rather than breaking the lane
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking
enabled, agentic tool use (file edits, shell, test execution) via Claude
Code
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Operators tune instance behavior through Settings → Experiments,
where experimental features are toggled on and off
> - The task graph liveness auto-recovery experiment shows a
confirmation dialog (preview of what would be recovered) before it is
enabled
> - After confirming with "Enable only" or "Enable and run", the
dialog's Radix overlay and the `pointer-events: none` body lock were
left behind, dimming the page and blocking all interaction until a
refresh
> - The dialog was unconditionally mounted and only closed inside the
mutation's `onSuccess`, so the overlay teardown depended on the mutation
outcome and could race or never happen
> - This pull request closes the dialog before the mutation fires in
both confirm flows, clears the pending preview alongside the open flag,
and mounts the dialog conditionally so its overlay fully unmounts
> - The benefit is that enabling an experiment behaves like every other
settings change: the dialog goes away, the page stays interactive, and
errors surface in the page-level error banner instead of a dead UI
## Linked Issues or Issue Description
No public GitHub issue exists for this bug; description follows the bug
report template. Refs #4587 (the PR that introduced the configurable
liveness auto-recovery controls this dialog belongs to).
**What happened?** In Settings → Experiments, toggling on "Task graph
liveness auto-recovery" and confirming via "Enable only" left the whole
UI dimmed and unclickable. The dialog content disappeared, but the modal
overlay and the `pointer-events: none` lock on `<body>` remained until a
full page refresh.
**Expected behavior:** Confirming (or dismissing) the auto-recovery
dialog should close it completely and return the page to a fully
interactive state, with the toggle reflecting the new setting.
**Steps to reproduce:**
1. Open Settings → Experiments.
2. Toggle on "Task graph liveness auto-recovery"; the confirmation
dialog with the recovery preview appears.
3. Click "Enable only".
4. The dialog content disappears but the page stays dimmed and nothing
is clickable; refreshing restores the UI and shows the setting was
applied.
**Paperclip version or commit:** master @ 634ae12 · **Deployment mode:**
local instance · **Area:** UI only
(`ui/src/pages/InstanceExperimentalSettings.tsx`).
## What Changed
- Added a `closeRecoveryPreview()` helper that resets both
`previewDialogOpen` and `pendingPreview` together, and used it
everywhere the dialog closes (confirm flows, run-mutation success, and
user dismissal).
- "Enable only" and "Enable and run" now close the dialog *before*
firing the mutation, so overlay teardown no longer depends on the
mutation outcome; mutation errors roll back the optimistic toggle and
surface in the existing page-level error banner.
- The `RecoveryPreviewDialog` is now conditionally mounted
(`previewDialogOpen ? <RecoveryPreviewDialog … /> : null`), guaranteeing
the Radix overlay and body pointer-events lock are fully removed when
closed.
- Added a regression test that walks the real flow — toggle on → preview
dialog appears → "Enable only" — and asserts the update payload, that
the dialog text and `[data-slot="dialog-overlay"]` element are gone, and
that the toggle reads enabled.
Credit: the implementation commit was authored by Cody — thanks! This PR
packages that fix for upstream review.
## Verification
- `pnpm vitest run src/pages/InstanceExperimentalSettings.test.tsx` in
`ui/` — 17/17 tests pass, including the new regression test.
- Independently re-verified beyond the committed assertions: with a
temporary assertion (not committed), confirmed
`document.body.style.pointerEvents` is `none` while the dialog is open
and released after "Enable only" — the actual "can't interact with
anything" symptom, not just overlay DOM removal.
- Manual check: open Settings → Experiments, toggle the auto-recovery
feature, click "Enable only" — the dialog closes, the page stays
interactive, and the toggle shows enabled without a refresh.
## Risks
- Low risk: change is confined to one page component's dialog lifecycle;
no API, schema, or shared-package changes.
- Behavioral shift: the dialog now closes immediately on confirm instead
of staying open with a pending spinner until the mutation resolves.
Errors are still surfaced via the page-level error banner, and the
optimistic toggle rolls back on failure.
## Model Used
- Implementation commit authored by the AI coding agent "Cody"
(Anthropic Claude-based agent). Review, independent verification, and PR
preparation by Claude (Anthropic), model ID `claude-fable-5`, extended
thinking with tool use (code execution, test runs).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Cody <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The adapter layer (hermes-local, process adapters) delegates agent
execution to child processes via `runChildProcess()`
> - `runChildProcess()` accepts an `onSpawn` callback to report child
PID and process group info, but the hermes and process adapters were not
forwarding `ctx.onSpawn` to this call
> - Without PID persistence, the orphan reaper cannot distinguish live
runs from abandoned processes, causing false-positive reaps and 5-minute
timeout errors for active runs
> - This pull request adds `onSpawn: ctx.onSpawn` to both adapter call
sites and declares the option in the `runChildProcess` wrapper type
> - The benefit is that the orphan reaper can now correctly track live
child processes, eliminating false-positive reaps
## Linked Issues or Issue Description
Fixes#8723
Fixes false-positive orphan reaps in hermes-local and process adapters
by forwarding the `onSpawn` callback to `runChildProcess()`. All other
adapters (claude-local, codex-local, cursor-local, gemini-local,
grok-local, opencode-local, pi-local) already forward `ctx.onSpawn` —
these two were the only ones missing it.
## What Changed
- `server/src/adapters/utils.ts`: Added `onSpawn?` to the
`runChildProcess()` options type so callers can forward the callback
- `server/src/adapters/process/execute.ts`: Forward `ctx.onSpawn` to
`runChildProcess()`
- `packages/adapters/hermes/src/server/execute.ts`: Forward
`ctx.onSpawn` to `runChildProcess()`
## Verification
- `pnpm -r typecheck` passes across all packages
- Confirmed all other adapters already forward `ctx.onSpawn` (12 grep
matches across 9 adapter files)
- The 3-line diff is additive only — no existing behavior is changed,
only a previously-ignored callback is now forwarded
## Risks
Low risk. This is a 3-line additive change. The `onSpawn` parameter is
optional (`?`) so existing callers are unaffected. The callback is
already well-established across all other adapters.
## Model Used
Hermes Agent (by Nous Research) — xiaomi/mimo-v2.5-pro via OpenRouter,
with tool use (file editing, git, GitHub API).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run tests locally and they pass (typecheck passes)
- [x] I have added or updated tests where applicable (N/A — type-level
fix only, no behavioral change)
- [x] I have updated relevant documentation to reflect my changes (N/A —
internal fix)
- [x] I have considered and documented any risks above
---------
Co-authored-by: Zephyr <zephyr@motoyuki.dev>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI leans on a shared design-token + component system so
surfaces stay visually consistent as they grow
> - Two small inconsistencies had crept in: several distinct blues were
used to signal "live/running" agent state across the sidebar, task
header, and chat thread; and in the Inbox an unread task's mark-read dot
was pushing that row's status icon and title one column right of read
rows
> - Both read as "not quite aligned" in daily use and undercut the
polish of the lists work that just landed
> - This pull request consolidates the live/running blues onto one
shared recipe and stops the unread dot from indenting the row
> - The benefit is one consistent "live" blue everywhere and Inbox rows
that line up whether read or unread
## Linked Issues or Issue Description
No public GitHub issue exists for this work; describing inline per the
bug-report template.
- **Problem**: (1) the same concept — an agent actively working —
rendered in three visibly different blues: the sidebar `N live` dot, the
task-detail "Live" badge, and the chat-thread "RUNNING" badge each used
a different token/recipe. (2) In the Inbox, unread rows carry a leading
mark-read dot that occupies the chevron column, but a per-row spacer was
still rendering in that same column — so on unread rows the status icon
+ title were shifted one column (~24px) further right than read rows.
Most visible when grouped by workspace.
- **Steps to reproduce**: open the Inbox with a mix of read and unread
tasks (group by workspace). The unread rows' status icons sit further
right than the read rows'. Separately, compare the blue of the sidebar
`N live` dot, a task's "Live" header badge, and a chat "RUNNING" badge —
they don't match.
- **Expected behavior**: unread and read rows align on the same status
column, with the unread dot centered on the workspace group chevron; and
all three "live/running" affordances share one blue.
## What Changed
- Added a shared `liveBlueBadge` recipe in `ui/src/lib/status-colors.ts`
and pointed the task-detail **Live** badge (`IssueDetail.tsx`) and the
chat-thread **RUNNING** badge (`IssueChatThread.tsx`) at it; removed the
now-redundant `brandChipBadge` usage from the chat thread and a stray
`🔵` breadcrumb prefix.
- Changed the sidebar **`N live`** dot (`SidebarNavItem.tsx`) to the
same `blue-600 / dark:blue-400` as its adjacent label text.
- **Inbox** (`Inbox.tsx`): skip the per-row leading spacer when the
unread mark-read dot is present, so the dot alone fills the chevron
column. Unread rows' status icon + title now sit in the same column as
read rows, and the dot centers on the workspace group chevron.
- **Test** (`Inbox.test.tsx`): added a regression test asserting an
unread leaf row renders the mark-read dot and drops the spacer, while a
read row keeps the spacer.
## Verification
- `pnpm typecheck` — clean (all packages)
- `pnpm check:token-gates` — 3/3 CLEAN
- `cd ui && pnpm vitest run src/pages/Inbox.test.tsx` — 14/14 (includes
the new regression test)
- Full Storybook visual suite (514 stories, both themes) — green locally
(CI cannot run this suite yet — the baseline-manifest archive is
unpublished, a pre-existing condition from #9134)
- Manual (workspace-grouped Inbox, 2× dark): measured the unread badge
center at the same x as the workspace chevron (276 = 276) and the
unread-row status icon at the same x as read-row status icons (292 =
292). Before/after screenshots in a PR comment below.
## Risks
Low risk — presentation only. No data, routing, or state changes. The
blue consolidation is a token/class swap; the Inbox change removes a
redundant spacer element on unread rows only (read rows and
non-grouped/mobile views are unaffected). The unread-row behavior is
covered by the new unit test.
## Model Used
Claude (Anthropic), Opus 4.8 — model id `claude-opus-4-8`; extended
thinking + tool use, driving local verification (typecheck, token gates,
vitest, Playwright visual suite + pixel measurements).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server package ships runtime asset trees used by built-in agents
and onboarding templates.
> - A prior server build omitted those source asset trees from `dist`,
allowing a built artifact to differ from runtime expectations.
> - The copy step is now present, but the existing build-gap gate only
checked TypeScript coverage for packages whose build skips `tsc`.
> - This pull request extends that standing gate so source asset files
under the server runtime asset trees must exist at the matching `dist`
paths after build.
> - The benefit is that future server asset additions fail loudly in CI
instead of silently shipping an incomplete `dist`.
## Linked Issues or Issue Description
### What happened?
After a server build, runtime asset files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**` could be
missing from `dist/` with no build failure. The existing build-gap gate
only checked TypeScript coverage for packages that skip `tsc`; it did
not verify that non-TypeScript source assets were copied to `dist`. A
server build that forgot the `cp -R` step, or that added a new asset
tree without updating the copy command, would produce an incomplete
`dist` without any CI signal.
### Expected behavior
After `pnpm --filter @paperclipai/server build`, every
non-TypeScript/non-JavaScript source file under `server/src/built-ins/`
and `server/src/onboarding-assets/` must exist at the matching path
under `server/dist/`. If any file is missing, the build-gap gate must
exit non-zero with a diagnostic listing the missing files and the
command to fix them.
### Steps to reproduce
1. Remove a copied runtime asset: `rm
server/dist/built-ins/agents/reflection-coach/AGENTS.md`
2. Run the guard: `node scripts/run-typecheck-build-gaps.mjs
--runtime-assets-only`
3. Before this fix: the command exits 0 and the missing file goes
undetected.
### Paperclip version or commit
Reproduced on `master` at `c36f1a4af` (`@paperclipai/server` 0.3.1).
### Deployment mode
Not deployment-specific — the build-gap check runs in CI on any
checkout.
## What Changed
- Extended `scripts/run-typecheck-build-gaps.mjs` with a source-derived
server runtime asset parity check for non-`.ts`/non-`.js` files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**`.
- Added a guard-only mode, `--runtime-assets-only`, for focused
pass/fail verification after a server build.
- Wired `pnpm run typecheck:build-gaps` to prepare plugin SDK build
deps, build the server package, then run the existing build-gap gate
plus the new asset check.
## Verification
Pass path:
```text
$ pnpm --filter @paperclipai/plugin-sdk ensure-build-deps
> @paperclipai/plugin-sdk@1.0.0 ensure-build-deps .../packages/plugins/sdk
> node ../../../scripts/ensure-plugin-build-deps.mjs
$ pnpm --filter @paperclipai/server build
> @paperclipai/server@0.3.1 build .../server
> tsc && mkdir -p dist/onboarding-assets dist/built-ins && cp -R src/onboarding-assets/. dist/onboarding-assets/ && cp -R src/built-ins/. dist/built-ins/
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
Regression simulation (guard catches the missing file):
```text
$ rm server/dist/built-ins/agents/reflection-coach/AGENTS.md
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] Missing server runtime asset(s) in dist:
- source: server/src/built-ins/agents/reflection-coach/AGENTS.md
expected dist: server/dist/built-ins/agents/reflection-coach/AGENTS.md
Run pnpm --filter @paperclipai/server build and ensure source runtime asset trees are copied into dist.
```
Standing gate (full end-to-end):
```text
$ pnpm run typecheck:build-gaps
[typecheck:build-gaps] typechecking 4 workspace(s): paperclipai, @paperclipai/plugin-authoring-smoke-example, @paperclipai/plugin-llm-wiki, @paperclipai/ui
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
## Risks
Low risk. The check only reads source and dist files during the
build-gap gate. The main tradeoff is that the gate now builds
`@paperclipai/server` so a clean checkout has generated `dist` content
to validate.
## Model Used
OpenAI Codex, GPT-5 based coding agent with repository tool use and
shell execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane people use to coordinate
AI-agent companies
> - Human operators use the Decisions attention queue to review and
resolve work that needs them
> - Decision rows were composed as a fixed content column plus a
right-side controls column
> - At phone widths, timestamps, actions, menus, and evidence thumbnails
compressed the decision headline until it was barely readable
> - This pull request makes each row respond to its own container width
and stacks metadata, content, evidence, and actions on narrow surfaces
while preserving the dense desktop layout
> - The benefit is a useful, thumb-reachable Decisions workflow on
phones and narrow side panels without regressing wide-screen density or
scrolling performance
## Linked Issues or Issue Description
No public GitHub issue exactly matches this bug, so it is described here
using the bug-report fields.
**What happened**
Decision rows used a fixed two-column layout. On narrow screens, the
right-hand timestamp, overflow menu, decision buttons, and optional
thumbnails squeezed the headline into a truncated sliver.
**Expected behavior**
Decision headlines should remain readable on mobile, supporting context
should flow below the headline, and primary actions should remain easy
to tap. Wide rows should retain the compact desktop presentation.
**Steps to reproduce**
1. Open the Decisions / What needs me surface with populated attention
items.
2. Reduce the row container to a phone-width layout (approximately
390px).
3. Observe rows with multiple actions or evidence thumbnails.
**Paperclip version / deployment mode**
Current `master`, board UI in local or hosted deployments.
**Related public work found during dedup search**
- Refs: #9311 — original What needs me attention queue work.
- Refs: #9468 — recent Decisions scrolling performance work preserved by
this change.
## What Changed
- Reworked `AttentionQueueRow` into a container-query-driven vertical
stack on narrow surfaces, with the existing compact layout restored at
wide row widths.
- Made decision titles wrap to two lines, moved project/evidence context
below the headline, and promoted actions to full-width mobile tap
targets.
- Preserved upstream row memoization and `content-visibility` scrolling
optimizations while rebasing onto current `master`.
- Added three 390px Storybook scenarios covering populated rows,
type/detail variants, and snoozed/dismissed curtains.
- Updated the focused row test to assert the new thumbnail/context
alignment.
## Verification
- `pnpm exec vitest run ui/src/components/AttentionQueueRow.test.tsx` —
1 file passed, 16 tests passed.
- `pnpm check:token-gates` — all token gates clean.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/ui build-storybook` — completed
successfully.
- `git diff --check public/master...HEAD` — passed.
## Risks
- Low risk: the behavior is isolated to the Decisions row presentation
and its Storybook coverage.
- Container-query breakpoints could need future visual tuning for
unusual embedded widths, but the wide layout remains available at the
row-level breakpoint.
- The mobile layout increases row height by design in exchange for
readable content and usable actions.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex CLI coding agent. The exact model ID and context-window
size are not exposed to this runtime; reasoning, repository editing,
shell execution, and test execution capabilities were enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Operators use the Decisions page to review an uncapped attention
feed across active, snoozed, and dismissed items
> - Large feeds mounted every row eagerly, and routine interactions
re-rendered the full queue
> - That made initial paint and scrolling progressively slower as
decision history accumulated
> - This pull request bounds rendering, stabilizes row props, and lets
off-screen rows skip layout and paint work
> - The benefit is a responsive Decisions page even for companies with
large attention histories
## Linked Issues or Issue Description
### What happened?
Opening `/decisions` for a company with a large attention history
eagerly mounted every visible-feed row. Expanding, selecting,
dismissing, snoozing, or restoring an item could also re-render the
entire queue.
### Expected behavior
The page should render a bounded initial window, progressively reveal
more rows near the scroll boundary, and avoid re-rendering unaffected
rows during interactions.
### Steps to reproduce
1. Populate a company with hundreds of attention items.
2. Open `/decisions`.
3. Scroll and interact with individual rows.
4. Observe increasing initial render, layout, paint, and interaction
cost on the previous implementation.
### Paperclip version or commit
Reproduced on `master` before this PR.
### Deployment and installation
Local development, built from source. This is a core UI issue, not
adapter- or database-specific.
### Additional context
Searched open public issues and PRs; no duplicate was found.
## What Changed
- Added a pure `planAttentionRenderRows` helper that allocates one
render budget across active groups and open snoozed/dismissed curtains
in document order.
- Render 50 rows initially and add 100 more when the Decisions page
approaches the scroll boundary.
- Memoized `AttentionQueueRow`, stabilized parent callbacks and inbox
dismissal actions, and passed row items through a shared expand
callback.
- Added `content-visibility: auto` and intrinsic containment so
accumulated off-screen rows avoid unnecessary layout and paint work.
- Added render-plan coverage and a regression test proving identical row
props do not re-render after a parent update.
## Verification
- `pnpm -C ui typecheck`
- `pnpm -C ui exec vitest run src/lib/attention.test.ts
src/components/AttentionQueueRow.test.tsx
src/components/Sidebar.test.tsx src/pages/Inbox.test.tsx` — 96 tests
passed
- `pnpm check:token-gates` — all gates clean
## Risks
- Low risk: the change is UI-only and does not alter API or database
contracts.
- The main behavioral risk is incorrect row-budget accounting across
collapsed groups or open curtains; the pure planner has focused tests
for ordering, truncation, and collapsed/closed sections.
- Progressive rendering means rows beyond the current budget are
intentionally absent until scrolling nears the boundary, matching the
existing Issues list pattern.
> This is a targeted performance fix and does not overlap planned core
feature work in `ROADMAP.md`.
## Model Used
- Anthropic Claude Fable 5 assisted with the implementation using
repository tools and code execution.
- OpenAI Codex `gpt-5.6-sol` prepared and verified the PR with high
reasoning effort, repository tools, shell execution, and
GitHub/Paperclip API access. The runtime did not expose a context-window
size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
documentation changes were required for this UI-only behavior)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip coordinates AI-agent work through repeated heartbeat runs.
> - Adapter prompts combine a default heartbeat template with scoped
wake context.
> - Fresh heartbeats received the same execution contract from both
layers, wasting prompt tokens and obscuring which layer owns the
contract.
> - Resume deltas and template-less adapters do not share that
composition path, so removing the wake-payload copy unconditionally
would drop required guidance.
> - Empty comment batches also emitted instructions and metadata that
only matter when comments exist.
> - This pull request makes execution-contract inclusion explicit by
prompt path, preserves OpenClaw gateway behavior, and suppresses no-op
comment boilerplate.
> - The benefit is one contract per heartbeat path and roughly 300 fewer
prompt tokens on a fresh zero-comment wake.
## Linked Issues or Issue Description
- Fixes#9221
- Refs #9200
- Refs #7634
## What Changed
- Stop emitting the execution-contract paragraph from fresh scoped wake
payloads because the default heartbeat template already contains the
full contract.
- Keep the contract in resume deltas, and add `includeExecutionContract`
for adapters that do not render the default heartbeat template.
- Opt `openclaw-gateway` into wake-payload contract rendering so
template-less gateway runs retain the guidance.
- Omit comment-batch acknowledgement/fetch guidance and empty `pending
comments` / `latest comment id` metadata when a fresh wake has no
pending comments.
- Add regression and acceptance coverage proving composed fresh prompts
contain `Execution contract` exactly once while resume and template-less
paths retain it.
Measured effect: the fresh zero-comment wake block drops from 1,840 to
855 characters (about 300 tokens saved per fresh heartbeat; about 220 on
comment wakes), and the composed fresh prompt contains `Execution
contract` once instead of twice.
## Verification
- `npx vitest run packages/adapter-utils/src/server-utils.test.ts` — 63
passed
- `npx vitest run server/src/__tests__/codex-local-execute.test.ts` — 13
passed
- `npx vitest run
server/src/__tests__/heartbeat-comment-wake-batching.test.ts
server/src/__tests__/openclaw-gateway-adapter.test.ts
server/src/__tests__/low-trust-red-team-routes.test.ts` — 27 passed
- `pnpm --filter @paperclipai/adapter-utils typecheck` — passed
- `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` —
passed
## Risks
- Low risk: prompt text and adapter composition only; no database or API
migration.
- The main compatibility risk is a template-less adapter losing the
contract. The explicit option and OpenClaw gateway regression coverage
protect the known template-less path.
- External adapters that call `renderPaperclipWakePrompt` directly can
opt into `includeExecutionContract: true` when they do not render the
default template.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5.4, reasoning mode with tool use and code
execution; context-window size is not exposed by the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane people use to coordinate
AI agents and their work
> - Heartbeat admission decides when an agent should start another
adapter session for an issue
> - After process-loss recovery, assignment pollers and reconcilers can
repeatedly request another wake while the issue remains `in_progress`
> - When the preceding runs succeeded without issue-visible progress,
those event-free wakes provide no new information but still pay the full
cost of an adapter session
> - Existing liveness evidence is too broad for this case because
workspace tool calls can make a run look active without moving the issue
> - This pull request adds an issue-scoped admission throttle for
consecutive no-progress re-wakes while preserving every wake that
carries new information or recovery intent
> - The benefit is bounded recovery cost without delaying comments,
operator actions, failures, or other meaningful events
## Linked Issues or Issue Description
No public GitHub issue exists for this bug.
**What happened?**
After a process died, external wake drivers could re-wake the same agent
for the same `in_progress` issue every few seconds. Each succeeded run
that produced no issue-visible progress could be followed by another
full adapter session despite no new issue input. In the observed
recovery smoke, one recovery consumed 25 sessions and 2.4× the
direct-run cost.
**Expected behavior**
Repeated event-free re-wakes should back off after consecutive
successful runs produce no issue-visible progress. Any new information,
explicit operator intent, or failed-run recovery should continue
immediately.
**Steps to reproduce**
1. Start an issue heartbeat and simulate process loss while the issue
remains `in_progress`.
2. Allow assignment/reconciliation drivers to request repeated
event-free wakes for the same agent and issue.
3. Complete each follow-up run successfully without adding a comment,
issue mutation, document, work product, interaction, or continuation.
4. Observe repeated adapter sessions starting every few seconds without
new issue input.
**Environment**
- Version: reproduced on `master` before this change
- Deployment: local development, built from source
- Adapter scope: core bug; not adapter-specific
- Database: reproduced and tested with embedded Postgres
## What Changed
- Add a pure issue re-wake throttle that detects consecutive succeeded
runs without issue-visible progress and applies a 120-second exponential
cooldown capped at 30 minutes.
- Gate event-free `enqueueWakeup` requests and return the explicit skip
reason `issue_rewake_throttled` while the cooldown is active.
- Always bypass throttling for comment wakes, new issue activity,
explicit resumes, `forceFreshSession`, event-shaped reasons, and
post-failure recovery.
- Add focused pure unit coverage and database-backed heartbeat admission
coverage for throttle and bypass behavior.
## Verification
- `cd server && pnpm vitest run
src/__tests__/issue-rewake-throttle.test.ts` — 12 passed.
- `cd server && pnpm vitest run
src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6 passed with
embedded Postgres.
- `cd server && pnpm run typecheck` — passed.
- Neighbor suites previously verified:
`heartbeat-dependency-scheduling`, `heartbeat-process-recovery`,
`run-continuations`, `heartbeat-issue-liveness-escalation`,
`recovery-stale-issue-lock-sweep`, and `heartbeat-comment-wake-batching`
— 131 tests passed.
## Risks
- A progress classifier that is too narrow could defer a legitimate
event-free poll; the cooldown is bounded and new issue activity bypasses
it immediately.
- A progress classifier that is too broad could allow the original
heartbeat storm; tests intentionally distinguish issue-visible mutations
from workspace-only activity.
- Low compatibility risk: no schema, API contract, or migration changes.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex coding agent. The runtime does not expose the exact
underlying model ID or context-window size; reasoning, terminal tool
use, code inspection, GitHub CLI access, and test execution were
enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My public PR branch name describes the change and contains no
internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source control plane for running and governing
AI-agent companies.
> - Adapter executions feed token usage, billing identity, and run cost
into the control plane's spend telemetry.
> - The default ACP execution lane for local Claude and Codex adapters
did not propagate per-turn usage or cost, so paid runs could be recorded
with zero spend and no tokens.
> - Claude CLI result events could also undercount output tokens by
reading only the main-loop usage block instead of the complete per-model
ledger.
> - The shared executor needs to distinguish per-run usage from
session-cumulative usage so the server does not apply the wrong delta
heuristic.
> - This pull request captures ACP usage and cumulative-cost deltas,
resolves adapter billing identity, uses Claude's complete model-usage
ledger, and preserves per-run usage in server normalization.
> - The benefit is accurate token and cost accounting across the default
paid Claude and Codex execution paths.
## Linked Issues or Issue Description
### What happened?
Paid `claude_local` and `codex_local` runs using the default ACP engine
can complete successfully while the control plane records zero or null
cost and missing token usage. Claude CLI result parsing can additionally
undercount output tokens when subagent or sidechain usage is present.
### Steps to reproduce
1. Run a paid Claude or Codex local adapter through the ACP engine.
2. Complete a turn that reports usage and cumulative cost through ACP
status/events.
3. Inspect the execution result and normalized run telemetry.
### Expected behavior
The execution result contains per-turn token usage, a per-run USD cost
delta, and the correct billing identity. Server normalization records
those per-run values without applying a session-cumulative delta a
second time.
### Actual behavior before this change
ACP execution results returned no usage and `costUsd: null` with unknown
billing. The server therefore recorded zero spend and no tokens for paid
runs. Claude CLI parsing could use an incomplete usage block.
## What Changed
- Capture ACP usage from runtime status and `usage_update` events,
reporting it as `usageBasis: per_run`.
- Convert agent-reported cumulative ACP cost into a per-turn delta,
including counter-reset and no-report safeguards.
- Add a shared billing-identity resolver and map Claude and Codex
authentication/provider modes to control-plane billing types.
- Prefer Claude result-event `modelUsage` totals so subagent and
sidechain tokens are included.
- Skip the server's session-cumulative usage delta when an adapter
explicitly reports per-run usage.
- Add regression coverage for usage capture, event fallback, cost
resets, stale reports, billing identities, model-usage totals, and
server spend normalization.
## Verification
- `pnpm exec vitest run
packages/adapter-utils/src/acpx-engine/execute.test.ts
packages/adapters/claude-local/src/server/parse.test.ts
packages/adapters/claude-local/src/server/acp.test.ts
packages/adapters/codex-local/src/server/acp.test.ts
server/src/__tests__/costs-service.test.ts
server/src/__tests__/monthly-spend-service.test.ts` — 6 files, 126 tests
passed.
- `pnpm --filter @paperclipai/adapter-utils typecheck` — passed.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` — passed.
- `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- A broader Claude-local suite has a pre-existing rate-limit
classification failure in `test.probe.test.ts`; it also fails on clean
`master` and is unrelated to this change.
## Risks
- Cost reporting depends on the agent's cumulative counter semantics;
reset handling falls back to the post-turn amount and is covered by
regression tests.
- Incorrect billing-mode inference could misclassify spend;
provider/auth mappings mirror each adapter's existing CLI behavior and
have focused tests.
- The new `usageBasis` contract changes server normalization only when
adapters explicitly opt into `per_run`; existing adapters retain prior
behavior.
- No database migration, workflow, lockfile, or UI changes are included.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Implementation commit: Anthropic Claude Fable 5, tool-enabled coding
workflow (exact context window and runtime configuration were not
recorded in the commit metadata).
- PR preparation and verification: OpenAI Codex, tool-enabled coding
agent (runtime model ID and context window are not exposed to this
session).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI-agent
companies and their recurring work
> - Scheduled routines provide native cron-driven execution for
recurring agent tasks
> - Watcher-style routines currently dispatch a model run even when the
control plane has been quiet since their last useful run
> - Existing pause, catch-up, and concurrency policies do not
distinguish external work from a routine's own bookkeeping
> - This pull request adds a generic activity gate that checks
company-scoped activity provenance before scheduled dispatch
> - The benefit is backward-compatible zero-token quiet skips while real
human, agent, or delegated-child activity still wakes the routine
## Linked Issues or Issue Description
- Refs #8534
## What Changed
- Added `activity_gate_policy` and `activity_gate_scope` routine columns
with backward-compatible `always` / `company` defaults.
- Added a company-bounded `evaluateActivityGate()` predicate that uses
the last dispatched run as its open window, excludes the routine's own
execution runs and scheduler bookkeeping, ignores pure-read actions, and
supports company/project scope.
- Integrated the predicate into scheduled ticks after pause/worktree
eligibility checks; quiet ticks create visible skipped run-history rows
with reason `no_external_activity` and gate-window diagnostics without
advancing the activity window.
- Kept webhook, manual, and API dispatch paths ungated; catch-up
schedules evaluate the gate once per scheduler tick.
- Added migration-default, provenance predicate, project-scope,
quiet-window, scheduler, and webhook-bypass coverage.
## Verification
- `pnpm --filter @paperclipai/db typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm exec vitest run server/src/__tests__/routines-service.test.ts` —
51 tests passed
- Embedded Postgres `EXPLAIN` for the company-scope gate scan:
```text
Limit (cost=24.56..24.58 rows=1 width=24)
-> Incremental Sort (cost=24.56..24.60 rows=2 width=24)
Sort Key: activity.created_at, activity.id
Presorted Key: activity.created_at
-> Nested Loop Anti Join (cost=0.44..24.55 rows=1 width=24)
Join Filter: (own_run.id = activity.run_id)
-> Index Scan using activity_log_company_created_idx on activity_log activity (cost=0.15..8.19 rows=1 width=40)
Index Cond: ((company_id = '00000000-0000-0000-0000-000000000001'::uuid) AND (created_at > (now() - '01:00:00'::interval)) AND (created_at <= now()))
```
## Risks
- The migration adds two non-null text columns, but constant defaults
preserve all existing routine behavior and avoid a backfill step.
- Project scope resolves activity through issue/run/routine provenance;
tests cover in-project and cross-project issue activity, while every
top-level and correlated query remains company-bounded.
- This is the scheduler/schema foundation. Public API validation and
documentation for configuring the new fields are intentionally handled
in the next scoped follow-up.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex using `gpt-5.4` with medium reasoning, repository/tool
access, terminal code execution, and test execution. The runtime did not
expose a context-window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR extends the
existing Scheduled Routines roadmap item
- [x] I have searched GitHub for duplicate or related PRs and linked the
related efficiency request above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
user-facing configuration is exposed in this scoped foundation PR)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Execution workspaces can run managed services that must report
reliable lifecycle and readiness state
> - Service startup previously waited for readiness before committing
the starting row, making concurrent control actions see stale state
> - Fixed service ports also needed clearer configuration and ownership
diagnostics to avoid cross-workspace collisions
> - This pull request persists startup state before readiness, validates
port ownership, and exposes configurable service ports in the workspace
UI
> - The benefit is dependable service controls and actionable
diagnostics when workspace runtimes start slowly or compete for ports
## Linked Issues or Issue Description
### What happened?
Slow-starting workspace services could remain invisible to concurrent
stop/restart controls until readiness completed, and fixed-port
conflicts lacked enough ownership context for safe repair.
### Expected behavior
A starting service is persisted immediately, control operations can
observe it, configured ports are editable, and conflicts identify the
owning process/workspace.
### Steps to reproduce
1. Configure a workspace service that delays binding its HTTP port.
2. Start the service and immediately request another control action.
3. Observe stale persisted state before this change.
4. Configure two workspaces for the same fixed port and observe limited
conflict diagnostics.
### Paperclip version or commit
`origin/master` at `02e2dd271`
### Deployment mode
Local dev; built from source; not adapter-specific; database-backed
workspace runtime state.
## What Changed
- Commit the `starting` runtime-service row before waiting for readiness
and transition it after the probe completes.
- Add port-owner inspection and cross-workspace conflict details to
local service supervision.
- Preserve configurable runtime service ports through workspace
configuration updates.
- Surface service-port editing and validation in the execution workspace
details UI.
- Add server and UI regression coverage for slow readiness, concurrent
controls, port persistence, and conflict diagnostics.
## Verification
- `vitest --project @paperclipai/server
src/__tests__/workspace-runtime.test.ts
src/__tests__/execution-workspaces-service.test.ts` — 118 tests passed.
- `vitest --project @paperclipai/ui
src/pages/ExecutionWorkspaceDetail.service-ports.test.ts` — 4 tests
passed.
- `node scripts/check-token-gates.mjs` — all token gates clean.
## Risks
- Moderate risk: changes touch workspace service lifecycle persistence
and local process/port inspection.
- No schema migration is required; tests exercise slow readiness,
concurrent control, persisted ports, and cross-workspace conflicts.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and
code execution; context-window size was not exposed by the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Isolated development workspaces need a repeatable run, verify, and
repair procedure
> - Port conflicts can be caused by another live Paperclip run that
keeps respawning and reclaiming a configured port
> - Restarting the target service alone does not resolve that class of
conflict
> - This pull request teaches the workspace-repair skill to identify the
owner, guard the master checkout, and stop the conflicting run before
repair
> - The benefit is safer recovery guidance that addresses the actual
port owner instead of creating restart loops
## Linked Issues or Issue Description
### What happened?
The workspace repair procedure could recommend restarting a managed
service while a separate live run still owned and reclaimed the
configured port.
### Expected behavior
The procedure identifies the owning process/run, protects the live
master checkout, and stops the conflicting owner before restarting the
intended service.
### Steps to reproduce
1. Start two managed workspace runs configured for the same fixed port.
2. Restart only the target workspace service.
3. Observe the sibling run reclaiming the port and the repair failing to
hold.
### Paperclip version or commit
`origin/master` at `02e2dd271`
### Deployment mode
Local dev; built from source; not adapter-specific; not
database-related.
## What Changed
- Expand the port-conflict diagnosis to distinguish dead owners from
live respawning runs.
- Add master-checkout safety checks before killing or restarting
processes.
- Document owner-first recovery and explicit verification of final port
ownership.
- Tighten the success checklist so a repaired workspace must prove
health and correct ownership.
## Verification
- Reviewed the rendered Markdown diff and command sequence for
consistent owner-first recovery.
- No executable code changes are included in this documentation-only PR.
## Risks
- Low risk: documentation and agent procedure only.
- Process termination guidance remains intentionally guarded by owner
identification and master-worktree checks.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and
code execution; context-window size was not exposed by the runtime. The
original change also credits Claude Fable 5 in the commit trailer.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The board coordinates repeated API polling across tabs to reduce
redundant requests
> - The shared polling coordinator retained cached result and
publication entries after the last subscriber left
> - Dynamic polling keys could therefore grow those maps for the
lifetime of the page
> - This pull request evicts inactive keys while preserving useful
short-lived handoff state and request deduplication
> - The benefit is bounded client memory without regressing cross-tab
polling behavior
## Linked Issues or Issue Description
### What happened?
Shared polling cached result/publication entries indefinitely after a
polling key no longer had subscribers.
### Expected behavior
Inactive keys are eventually removed, while recently published values
remain available long enough for normal subscriber handoff.
### Steps to reproduce
1. Create and unsubscribe many distinct shared polling keys in one page
lifetime.
2. Inspect the coordinator's cached results and publication timestamps.
3. Observe that the old maps retain every historical key.
### Paperclip version or commit
`origin/master` at `02e2dd271`
### Deployment mode
Local dev; built from source; not adapter-specific; not
database-related.
## What Changed
- Track inactive polling keys and schedule bounded cache eviction.
- Preserve cached data while a key is active or inside its retention
window.
- Cancel stale cleanup timers when polling resumes and clear coordinator
caches during disposal.
- Add focused fake-timer coverage for retention, resubscription, and
disposal behavior.
## Verification
- `vitest --project @paperclipai/ui src/lib/cross-tab-poll.test.ts` — 11
tests passed.
## Risks
- Low-to-moderate risk: eviction timing affects client polling
coordination.
- Tests cover the retention boundary, resumed subscriptions, and
coordinator cleanup to reduce regression risk.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and
code execution; context-window size was not exposed by the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Company skills expose their source metadata in a narrow details
sidebar
> - Long filesystem paths and repository locators were truncated, hiding
the part operators often need to distinguish sources
> - The sidebar can preserve the complete value by wrapping at arbitrary
path boundaries instead of ellipsizing it
> - This pull request renders full source paths and repository labels
without widening the layout
> - The benefit is that operators can inspect and copy the actual skill
source from the UI
## Linked Issues or Issue Description
### Pre-submission checklist
- [x] I have searched existing open and closed issues and this is not a
duplicate.
- [x] I can reproduce this on `master`.
- [x] I have confirmed the behavior originates in Paperclip itself, not
an agent adapter, API provider, or local configuration.
### What happened?
Long company-skill source paths and repository locators were truncated
in the skill details sidebar.
### Expected behavior
The complete source value remains visible and wraps within the available
sidebar width.
### Steps to reproduce
1. Open a company skill whose source path is longer than the details
sidebar.
2. View the Source field.
3. Observe that the old UI replaces the middle or end of the value with
an ellipsis.
### Paperclip version or commit
`origin/master` at `02e2dd271`.
### Deployment mode
Local dev (`pnpm dev`).
### Installation method
Built from source (`pnpm dev` / `pnpm build`).
### Agent adapter(s) involved
None; this is a company-skills UI layout issue.
### Logs, configuration, or screenshots
Not applicable; the behavior is directly visible in the Source field.
### Additional context
The narrow sidebar should remain width-constrained. Wrapping
intentionally trades vertical space for full source inspectability.
## What Changed
- Replace source-path truncation with width-constrained arbitrary
wrapping.
- Apply the same wrapping behavior to linked repository/source labels.
- Add a regression test proving the full long path is rendered without
ellipsis.
## Verification
- `vitest --project @paperclipai/ui src/pages/CompanySkills.test.tsx` —
11 tests passed.
- `node scripts/check-token-gates.mjs` — all token gates clean.
## Risks
- Low risk: the change is limited to text layout in the company skill
details view.
- Very long unbroken values may make the Source section taller,
intentionally trading vertical space for inspectability.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and
code execution; context-window size was not exposed by the runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Operators scan task state constantly, so the task **status**
vocabulary (backlog / todo / in progress / in review / done / blocked /
cancelled) has to read instantly
> - Those statuses render through one shared component, `StatusGlyph`,
whose icons were hand-rolled SVG geometry lifted from an internal spec
> - Hand-rolled glyphs are harder to reason about, drift from the rest
of the UI (which uses Lucide everywhere else), and mix fill/stroke
styles across statuses
> - This pull request swaps the hand-rolled geometry for named Lucide
icons — one clean, consistent icon family — with no change to colours,
sizing, or accessibility
> - The benefit is a status icon set that is consistent with the rest of
the app's iconography, trivially adjustable (change a mapping, not SVG
path math), and simpler to maintain
## Linked Issues or Issue Description
No existing public GitHub issue. Describing the change in-PR
(feature/polish):
**Problem / motivation.** The task status icons in `StatusGlyph` were
bespoke inline SVGs (a half-filled disc for *in progress*, a filled disc
+ knockout check for *done*, ring+bar for *blocked*, ring+slash for
*cancelled*, etc.). The rest of the UI uses [Lucide](https://lucide.dev)
icons, so the status set was the odd one out — and its mixed fill/stroke
shapes were harder to scan and to tweak.
**Proposed solution.** Map each status to a Lucide icon and render that
instead:
| Status | Lucide icon |
| --- | --- |
| backlog | `circle-dashed` |
| todo | `circle` |
| in_progress | `rotate-cw` |
| in_review | `circle-dot` |
| done | `circle-check` |
| blocked | `circle-minus` |
| cancelled | `ban` |
| in_queue (covered-blocked) | `circle-minus`, recoloured blue |
Colours (the `--status-task-icon-*` tokens), the `sm/md/lg` size scale,
`currentColor` recolouring, and the `role="img"` / `aria-label`
behaviour are all unchanged — only the shapes change.
**Alternatives considered.** Keeping the bespoke geometry (rejected:
inconsistent with the app and harder to maintain).
**Related PRs** (linked for reviewer context, not dependencies):
- Refs #8580 — the merged PR that established the current hand-rolled
status glyphs this PR restyles.
- Refs #8838 — open PR forwarding Radix trigger props through
`StatusGlyph`; touches the same component (no overlap with this change).
- Refs #1760 — open proposal to redesign the *cancelled* status icon
specifically; this PR moves cancelled to Lucide `ban`.
## What Changed
- `ui/src/components/StatusGlyph.tsx`: replaced the per-status
hand-rolled SVG `glyphBody()` geometry with a `status → Lucide icon` map
(`circle-dashed`, `circle`, `rotate-cw`, `circle-dot`, `circle-check`,
`circle-minus`, `ban`). Kept the token-driven colour wiring, size scale,
`currentColor` recolouring, a11y label handling, and the `in_queue` =
blocked-icon-recoloured-blue behaviour.
- `ui/src/components/StatusGlyph.test.tsx`: updated to lock the new icon
mapping (per-status Lucide class, size scale, colour var, `in_queue`,
a11y) instead of the old geometry.
Net: two files, +74 / −138 (the component got smaller). Because every
status surface (list, board, detail header, status picker,
sub-task/blocked-by pills, chips) routes through `StatusGlyph`, this
single-component edit covers them all.
## Verification
- `pnpm check:token-gates` → **3/3 clean** (no hardcoded
colour/spacing/font values introduced).
- `pnpm typecheck` → clean across all packages.
- `cd ui && pnpm vitest run` → **2509/2509 passing**, including the
updated `StatusGlyph` test.
- Manual: ran the worktree dev server and confirmed the new icons render
everywhere (task list, task detail, related-task chips, and the status
picker showing all seven).
**Storybook visual-regression note:** this is an intentional visual
change, so the status-icon stories will diff against the published
baseline. The baseline snapshots need to be regenerated and republished
by a maintainer (`pnpm test:storybook-visual:update` from a trusted
environment) as part of accepting this change — the visual-regression CI
check is expected to be red until then. No baseline is published in the
environment this PR was authored in, so that step is left to a
maintainer.
## Risks
- **Low risk / cosmetic.** No logic, data, or API changes — only the
rendered icon shapes. Colours, sizes, and accessibility labels are
unchanged.
- The most noticeable shifts are *in progress* (half-disc → rotating
arrow), *done* (solid disc+check → outline circle+check), and
*cancelled* (ring+slash → ban). These are deliberate.
- The only CI check expected to fail is the Storybook visual-regression
job, pending a maintainer baseline update (see Verification).
## Model Used
Claude Opus 4.8 (`claude-opus-4-8`), run in Claude Code with extended
thinking and tool use (file edits, local test runs, browser-driven
visual verification).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server startup path reports the product version from package
metadata and, in source checkouts, Git metadata.
> - Packaged installs can run from `node_modules`, where Git metadata is
normally unavailable and that absence is expected.
> - The fallback path was still attempting Git metadata probing in
packaged contexts, which could print scary diagnostic noise during
onboarding even though the package version fallback was working.
> - This pull request makes the packaged path skip Git probing only when
the package does not look like a source checkout, and keeps fallback
diagnostics opt-in.
> - The benefit is a quieter first-run experience without weakening
source-checkout version detection or debug diagnostics.
## Linked Issues or Issue Description
No public GitHub issue exists.
### Bug Report
#### Pre-submission checklist
- [x] I have searched existing open and closed issues and this is not a
duplicate.
- [x] I am on the latest released version of Paperclip or can reproduce
on `master`.
- [x] I have confirmed the error originates in Paperclip itself, not in
an agent adapter, API provider, or local configuration.
#### What happened?
When Paperclip starts from a packaged install, server version resolution
can fall back from Git metadata to package metadata. That expected
fallback path could emit scary Git diagnostic noise during onboarding
even though startup could continue normally.
#### Expected behavior
Packaged Paperclip startup should use package metadata quietly when Git
metadata is unavailable. Source checkouts should still use Git-derived
versions, and operators who explicitly opt into version-resolution
diagnostics should still receive useful Git failure details.
#### Steps to reproduce
1. Run Paperclip from a packaged install where the server package is
under `node_modules` and does not include package-local Git metadata.
2. Start the server in an environment where `git describe` cannot
resolve repository metadata for that package.
3. Observe that version fallback can produce Git diagnostic noise during
startup even though the package version fallback is expected.
#### Paperclip version or commit
Reproduced against the pre-fix server version resolution behavior on
`master`-derived builds.
#### Deployment mode
Self-hosted server / packaged local install.
#### Installation method
npm / pnpm package install.
#### Agent adapter(s) involved
Not adapter-specific; this is core server startup/version behavior.
#### Database mode
Not database-related.
#### Access context
Unclear / not applicable.
#### Relevant logs or output
Git fallback diagnostics from `git describe` could appear during
packaged startup. The exact path and Git output depend on the operator
environment.
#### Additional context
The fix keeps diagnostics available behind
`PAPERCLIP_DEBUG_VERSION_RESOLUTION=1` and preserves source-checkout Git
version detection, including source paths that happen to contain a
`node_modules` segment.
#### Privacy checklist
- [x] I have reviewed all pasted output for PII and redacted where
necessary.
## What Changed
- Skip Git metadata probing for packaged installs under `node_modules`
only when no package-local Git metadata is present.
- Preserve Git-derived version detection for source or linked workspace
checkouts, even when their path contains a `node_modules` segment.
- Keep fallback diagnostics behind the existing debug/diagnostic opt-in
path.
- Include useful Git failure details such as stderr/stdout/stack/cause
when diagnostics are enabled.
- Add version tests covering packaged fallback behavior, source-checkout
detection, richer diagnostics, and quiet default output.
## Verification
- `pnpm vitest run server/src/__tests__/version.test.ts` passed after
the Greptile follow-up changes.
- `pnpm --filter @paperclipai/server typecheck` passed.
- `git diff --check` passed.
- Greptile completed with confidence score 5/5 and no blocking issues on
the latest reviewed commit.
## Risks
Low risk. The change is scoped to version fallback behavior.
Source-checkout Git version detection remains covered, while packaged
`node_modules` contexts intentionally rely on package metadata instead
of Git probing.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex, GPT-5-class coding agent with shell/tool use in the
Paperclip workspace.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Instance settings include an Environments section where operators
configure execution environments, each with an environment-variables
editor for run-time bindings
> - The editor showed a bare "Unsaved changes" banner that never said
which variables changed, sometimes appeared the moment a saved config
was opened (a lossy round-trip through the editor's emit rules made
clean values look dirty), and the environment form let you navigate away
without any confirmation, silently dropping the draft
> - Operators could not tell what was unsaved, distrusted the phantom
banner, and lost half-finished environment edits to a stray click — the
agent configuration page already confirms before discarding, so
environments behaved inconsistently
> - This pull request lists the new/edited/removed variable names under
the banner, normalizes both sides of the dirty comparison so saved
values no longer look dirty on open, and confirms before cancel, in-app
navigation, or tab unload while the form has unsaved changes
> - The benefit is that the banner is trustworthy and specific, and
unsaved environment edits can no longer be lost without an explicit
confirmation
## Linked Issues or Issue Description
Related (not fixed by this PR): #8930 introduced the current
environment-variables editor and its unsaved-changes banner; #9386 moved
environment create/edit from a modal to routed pages, which this PR's
navigation guard builds on.
No existing public issue for the defects themselves; described per the
bug report template:
**What happened?**
The environment-variables editor in Environments settings showed a bare
"Unsaved changes" banner with no indication of which variables changed.
For some saved configurations (names with surrounding whitespace,
incomplete secret references, duplicate names differing only by
whitespace) the banner appeared immediately on opening the edit form,
before any user input. Navigating away from the environment form —
cancel, an in-app link, or closing the tab — silently discarded the
draft with no confirmation.
**Expected behavior**
The banner should say which variables are new, edited, or removed; a
freshly opened saved configuration should show no banner; and leaving
the form with unsaved changes should require an explicit confirmation,
consistent with the agent configuration page.
**Steps to reproduce**
1. Open Settings → Instance settings → Environments and edit an
environment whose saved config round-trips lossily (e.g. an env var name
stored with trailing whitespace) — the "Unsaved changes" banner appears
with no user edits.
2. Add or edit a variable — the banner gives no hint of what is unsaved.
3. With a dirty draft, click any in-app link or Cancel — the draft is
dropped with no confirmation.
**Deployment mode**
Self-hosted (local development instance), reproducible on `master`.
## What Changed
- The unsaved-changes banner in `EnvironmentVariablesEditor` now renders
a change summary line — `New: … · Edited: … · Removed: …` — showing up
to three names per group with a `+N more` overflow and the full list in
a `title` tooltip. A rename shows as one addition plus one removal.
- Dirty detection normalizes both the committed value and the draft
through the same rules the editor uses when emitting values (trimmed
names, incomplete secret refs dropped, last-writer-wins on trimmed
duplicates), so a saved config that round-trips lossily no longer shows
a phantom banner on first open.
- The editor exposes an `onDirtyChange` callback and warns via
`beforeunload` while its local draft is dirty.
- The environment create/edit page (`CompanyEnvironments`) tracks a
payload-level baseline fingerprint of the form as initialized and treats
the page as having unsaved changes when the current form differs from it
or the editor draft is dirty. While dirty it confirms ("Discard unsaved
environment changes?") on Cancel, intercepts same-origin in-app link
clicks, and warns on tab unload.
## Verification
- `node_modules/.bin/vitest run
ui/src/pages/CompanyEnvironments.test.tsx
ui/src/components/environment-variables-editor/EnvironmentVariablesEditor.test.tsx`
— 48 tests pass, including new coverage for: the change-summary banner
text, no phantom banner for lossy round-trip values, beforeunload only
while dirty, cancel confirmation on the edit page, and unload/link-click
warnings after edits are staged into the form.
- `tsc -b` in `ui/` passes.
- Manual: edit an environment, add/edit/remove variables, observe the
summary line; click Cancel or an in-app link and observe the
confirmation; save and observe navigation proceeds without prompting.
## Risks
- Low risk, UI-only. The click interceptor is scoped to same-origin
anchor navigation while the environment form page has unsaved changes
and is removed on cleanup; modified-key/middle-button clicks and
external links are left alone.
- The dirty-normalization intentionally ignores differences the editor
could never persist (incomplete secret refs, untrimmed duplicate names);
those were previously reported as unsaved changes that could not be
saved away.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code, with
extended thinking and tool use (code editing, test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent adapters (Claude, Codex, Gemini) default to the ACP engine
lane, which needs a live bidirectional stdio session with the agent
process
> - Sandbox execution targets only exposed one-shot command execution,
so every ACP-capable adapter refused remote targets and fell back to the
CLI lane with a "supports only the local Paperclip host" warning
> - Running agents in sandboxes is a core deployment mode, and losing
ACP there means losing streaming updates, structured events, and
default-lane parity with local runs
> - This pull request adds a provider-agnostic process-session bridge
that relays the ACP stdio session into the sandbox over the existing
sandbox runner contract, and updates the adapters to use it
> - The benefit is that the default ACP lane now behaves the same on the
local host and in any sandbox provider, with CLI fallback reserved for
targets that genuinely cannot host a bidirectional session
## Linked Issues or Issue Description
No existing public issue covers this; inline description following the
feature request template:
**Problem or motivation**
Configuring an ACP-capable adapter (e.g. Claude) with a sandbox
environment made every run fall back to the CLI lane with the warning
"Claude ACP currently supports only the local Paperclip host, but this
run targets a remote environment." The ACP engine only knew how to spawn
a local subprocess, while sandbox providers only expose one-shot command
execution — so there was no way to hold the bidirectional stdio session
ACP requires.
**Proposed solution**
Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable
proxy script connects to a token-authenticated loopback TCP server,
which relays JSON-framed stdin/stdout/stderr events to and from a small
relay script executed inside the sandbox via the provider's ordinary
runner. Claude/Codex/Gemini adapters now treat sandbox targets with a
runner as ACP-capable, resolve agent commands against the remote target,
and fall back to CLI only when the sandbox exposes no bidirectional
path. The sandbox callback bridge injects a run-scoped API endpoint and
bridge token so the agent inside the sandbox can reach Paperclip
(including work-product handoffs) without ever receiving the host run
JWT.
**Alternatives considered**
A provider-specific lane was prototyped first: Daytona minting SSH
access metadata at lease time, converted into an SSH execution target.
It was dropped because it only worked for providers able to advertise
SSH, added per-provider surface area, and left every other sandbox
provider on the CLI fallback. The merged design rides the one-shot
runner contract all providers already implement; a regression test pins
that sandbox targets stay on the bridge lane even when lease metadata
advertises SSH access.
**Roadmap alignment**
Directly advances the "Cloud / Sandbox agents" roadmap item — agents
running in remote and sandboxed environments keep the same control-plane
behavior as local ones. No overlap with other planned core work.
## What Changed
- `packages/adapter-utils/src/execution-target.ts`: new
`startAdapterExecutionTargetProcessSessionBridge()` plus helpers —
writes a token-authenticated local proxy script (spawnable by ACPX) and
a remote relay script synced into the sandbox, with a loopback TCP
server streaming JSON-framed stdio between them; events emitted before
the ACP client attaches are buffered so none are lost.
- `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine
can execute against remote sandbox targets through the bridge instead of
requiring a local subprocess, including remote cwd/env shaping.
- `packages/adapter-utils/src/sandbox-callback-bridge.ts`:
sandbox-scoped API bridging extended to allow work-product handoffs; the
sandbox payload env carries a bridge token, never the host run JWT.
- `packages/adapters/claude-local`, `codex-local`, `gemini-local`
(`src/server/acp.ts`): default-lane selection no longer rejects all
remote targets; command resolution is remote-aware
(`ensureAdapterExecutionTargetCommandResolvable`,
`resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped
to sandboxes that expose only one-shot execution.
- `server/src/__tests__/environment-execution-target.test.ts`: pins that
sandbox targets resolve to the bridge lane, including when lease
metadata advertises SSH access.
- Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP
engine's remote transport is sandbox-only, so default-lane selection
falls back for those targets across all three adapters, and tests
covering CLI-specific remote behavior pin `engine: "cli"` explicitly.
- The bridge authenticates loopback connections before they can own the
session or receive buffered output (token required, idle unauthenticated
peers dropped), and remote event writes are serialized so the exit event
always lands after stdout/stderr have drained.
- Daytona plugin: formatting-only residue from the earlier iteration; no
functional change.
## Verification
- `vitest run` over the touched suites —
`packages/adapter-utils/src/acpx-engine/execute.test.ts`,
`packages/adapter-utils/src/execution-target-sandbox.test.ts`,
`packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three
adapter `acp.test.ts` files, and
`server/src/__tests__/environment-execution-target.test.ts` — 102 tests
pass.
- End to end: with a Claude agent configured on a Daytona sandbox
environment, the primary-model test now selects the default ACP lane (no
fallback warning), and the full round trip (wake → sandbox execution →
API bridge → comment post) was exercised twice from inside a live
sandbox.
## Risks
- Behavioral shift: adapters that previously always fell back to CLI on
sandbox targets now default to ACP there; `engine=cli` still pins the
CLI lane explicitly.
- The bridge relays stdio as JSON lines over loopback TCP guarded by a
per-session random token; the remote relay runs inside the sandbox under
the provider's runner. Providers with slow one-shot execution will see
higher session startup latency — the CLI fallback remains for genuinely
incapable targets.
- No schema or migration changes.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking
enabled, agentic tool use via the Claude Agent SDK harness;
implementation iterated with local Vitest verification.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
shipped docs describe the old local-only ACP limitation)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Cody <noreply@paperclip.ing>
Co-authored-by: Cody <cody@paperclip.local>