> Follow-up to #11006 (merged): rebased onto master and ready for
review.
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release subsystem now publishes canary (every master push),
nightly (scheduled, smoke-gated, added in #11006), and stable (manual)
> - There is still no human-approved release-candidate lane between
nightly and stable, and nothing enforces that a stable actually soaked
anywhere before shipping
> - Betas need a real approval gate, and stables need a soak policy that
is data, not prose
> - This pull request adds the beta channel: a manual promotion of a
chosen nightly behind the `npm-beta` environment gate, re-smoked after
publish, plus a stable preflight that enforces a 3-day beta soak with a
written-justification bypass
> - The benefit is a complete canary → nightly → beta → stable train
where every stable shipped as a beta first, and emergencies leave a
written trace
## Linked Issues or Issue Description
**Subsystem affected**
Release automation: `scripts/release.sh`, `scripts/release-lib.sh`,
`.github/workflows/release.yml`, `.github/workflows/docker.yml`,
`.github/workflows/release-smoke.yml`.
**Problem or motivation**
After #11006 the project has canary and nightly prerelease lanes, but no
release-candidate lane. Stable promotion has no enforced soak: any ref
can ship as stable directly. There is no approval boundary for a
broader-audience prerelease, and no structured way to record why an
emergency release skipped validation.
**Proposed solution**
Add a `beta` channel: a manual dispatch that promotes a chosen nightly's
source commit, publishes behind the `npm-beta` GitHub environment
(required reviewers are the gate), re-smokes the published beta, and
tags `beta/vX`. Enforce in the stable path that the source commit
shipped as a beta at least 3 days earlier (measured from the beta's npm
publish time), with a `skip_soak_justification` input as the recorded
emergency bypass.
**Alternatives considered**
Codifying the soak policy in docs only. Rejected: an unenforced policy
decays; the preflight makes the policy executable while the
justification input keeps the emergency path usable and auditable.
## What Changed
- `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel —
requires HEAD to carry a `nightly/v*` tag, publishes the package set as
`YYYY.MDD.P-beta.N` under dist-tag `beta`, tags
`beta/vYYYY.MDD.P-beta.N`
- `.github/workflows/release.yml`:
- `channel: beta` dispatch path: `select_beta` resolves the newest (or
an explicit `source_version`) nightly and fails loudly on selection
problems; `publish_beta` runs behind the `npm-beta` environment, pushes
the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release
smoke suite against the exact published beta version
- stable path: new `preflight_stable` job enforces the 3-day beta soak
from the beta's npm publish time; `skip_soak_justification` bypasses
with the reason echoed into the job summary; dry runs report without
blocking
- `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both
images, with exact version stamping
- `.github/workflows/release-smoke.yml`: `beta` added to the dispatch
choice list
- Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate,
and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta`
environment setup, including the warning to create the environment
before the first beta dispatch (GitHub auto-creates unprotected
environments on first reference)
- Tests: beta version-counting coverage in
`scripts/release-registry-versions.test.mjs`; beta identity and
nightly-tag guard coverage in
`scripts/__tests__/release-dry-run-notes.test.mjs`
## Verification
- `node --test` on the two touched suites: 17 pass, including the 3 new
beta tests
- `bash -n` on both shell scripts and YAML parse of all three workflows
- After merge, in order: create the `npm-beta` environment, dispatch
`channel: beta` with `dry_run: true` to preview, then a real promotion
of a published nightly through the approval gate, then a stable dry-run
against a young beta to see the soak gate report
## Risks
- If the `npm-beta` environment does not exist when the first beta
dispatch runs, GitHub creates it with no protection rules and the beta
publishes without approval. Mitigated by documentation and by creating
the environment before merge (operator step)
- Until the first beta exists, every stable dispatch requires
`skip_soak_justification`. This is deliberate — the first beta ships
immediately after this merges — but it is a behavior change to the
stable dispatch
- The soak clock reads the beta's npm publish time from the registry; a
registry outage makes the preflight fall back to requiring justification
(fail-closed)
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with
extended thinking and full tool use (repository exploration, local test
execution, live registry and git verification). All code, tests, and
docs in this PR were model-authored under human direction.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending — will confirm before
merge)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending — will confirm before merge)
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip connects AI agents to local and remote runtimes.
> - The Codex local adapter needs a safe device-login flow.
> - A future sandbox integration needs strict prompt validation, secret
protection, cleanup, and private credential storage.
> - This pull request adds tested building blocks for that flow.
> - The result gives a later Daytona integration a clear security
boundary.
## Linked Issues or Issue Description
No public issue covers this change.
**Problem or motivation**
The Codex local adapter has no safe, reusable flow to prove device login
inside an isolated sandbox.
**Proposed solution**
Add parser, runner, credential export, and proof helpers. Validate the
prompt, protect login data, store credentials in a private run-scoped
home, and dispose all sandbox resources.
**Alternatives considered**
Do not connect a production Daytona driver in this change. Use an
injected sandbox driver and focused tests first. This keeps the security
controls testable before live provider integration.
**Roadmap alignment**
The change extends the Codex local adapter. It does not add a core
Paperclip route or duplicate a planned core feature.
**Additional context**
The flow keeps the login URL, code, and token out of logs, results, and
errors. The proof home uses a company-scoped root and a run-scoped
private directory.
## What Changed
- Add a pure parser for the exact Codex device-login URL and one-time
code shape.
- Add a sandbox runner with prompt handling, timeout, cancellation, and
disposal.
- Add a credential export step with company scoping, path checks,
payload checks, private modes, locking, and cleanup.
- Add redacted device-login fixtures and focused tests for parsing,
secret redaction, runner outcomes, credential export, and cleanup.
## Verification
- Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`.
- Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc
--noEmit`.
- Review tests for strict URL and code validation, timeout,
cancellation, disposal, secret redaction, path safety, payload safety,
file modes, and cleanup.
## Risks
- This change provides building blocks, not a live Daytona proof.
- A later integration must connect the runner to a concrete sandbox
driver.
- Credential export depends on existing Codex authentication cache
helpers.
- Incorrect path or payload assumptions can reject valid credentials.
## Model Used
Codex, GPT-5, tool use, code execution, and repository review. The
Paperclip runtime controls the exact context window and reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release subsystem's nightly lane gates every nightly publish on
the release smoke suite, which drives real onboarding in a browser
against the published artifact
> - With the harness fixed (#11187, #11189), the gate reached the
Playwright suite for the first time in CI — and the spec still walks the
old onboarding wizard, so it fails at "Create your first agent" on every
current build
> - The wizard was redesigned to a mission-first five-step flow, and the
spec rotted silently because the suite never ran in CI before
> - This pull request rewrites the spec to drive the current wizard end
to end
> - The benefit is a smoke gate that actually tests today's product,
verified against a real published canary
## Linked Issues or Issue Description
**Subsystem affected**
Release smoke testing:
`tests/release-smoke/docker-auth-onboarding.spec.ts`.
**Problem or motivation**
Nightly run 31431273139 failed in the smoke Playwright suite: the spec
expects the old wizard step "Create your first agent", but current
builds show the redesigned mission-first flow (front door → company →
mission → team lead → connect model → review). The page snapshot in the
run artifact shows the "Define your mission" step where the spec
expected the agent step. Both retries failed identically — this is
deterministic spec drift, not flake.
**Proposed solution**
Rewrite the spec for the current flow: fill the company name, define the
mission directly (confirming creates the company), name the team lead,
hire it through the adapter step — the adapter environment probe reports
unhealthy in the CLI-less smoke container by design and must not block
the hire — then launch to the dashboard. Assert the company, the
ceo-role agent, and the company goal through the API. The first-task and
assignment-run assertions are removed together with the wizard flow that
created them.
## What Changed
- `tests/release-smoke/docker-auth-onboarding.spec.ts`: rewritten for
the mission-first wizard; sign-in and wizard-opening helpers and the
company-name step are unchanged
## Verification
- Full local run against the real nightly candidate: launched the smoke
container for `paperclipai@2026.810.0-canary.3` via
`scripts/docker-onboard-smoke.sh` (with the #11189 bind fix), then ran
`pnpm run test:release-smoke` against it — 1 passed (4.5s)
- After merge: dispatch `release.yml` with `channel: nightly` to run the
full gate in CI
## Risks
- Low. Test-only change. The spec now asserts less about first-task
creation because the wizard no longer creates a first task; if a
first-run trigger returns to onboarding, the spec should grow that
assertion back
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with
extended thinking and full tool use (CI artifact forensics, UI source
tracing, local Docker + Playwright reproduction and verification). All
changes model-authored under human direction.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending — will confirm before
merge)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending — will confirm before merge)
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release subsystem's nightly lane gates every nightly on the
release smoke suite, which boots the published artifact in a Docker
container and drives real onboarding
> - The gate kept failing even after the readiness budget fix (#11187),
and the new container-log dump revealed the server was healthy but
listening on 127.0.0.1 inside the container, unreachable through
Docker's port mapping
> - `onboard --yes` without an explicit `--bind` prefers trusted-local
quickstart defaults: it writes a loopback bind into the instance config
and ignores the deployment env vars the harness passes, and that config
outranks `HOST` at runtime
> - This pull request pins the smoke container to the `lan` bind preset
and adds a wiring test for it
> - The benefit is a working nightly gate, verified end to end against a
real published canary
## Linked Issues or Issue Description
**Subsystem affected**
Release smoke testing: `docker/Dockerfile.onboard-smoke`,
`scripts/__tests__/release-verify-workflow.test.mjs`.
**Problem or motivation**
Nightly run 31428558684 failed in smoke with the server unreachable at
the mapped port for the full 420 second budget. The container logs
(captured thanks to #11187) show a fully booted server with `Bind
loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the
deployment env vars, but `onboard --yes` without `--bind` deliberately
prefers trusted-local defaults, writes `bind: loopback` into the
instance config, and the config outranks `HOST` at runtime. A loopback
listener inside a container is invisible to the port mapping, so the
health check can never pass. This behavior predates the current stable,
so the harness was silently broken against every recent version — it
only surfaced now because the nightly lane is the suite's first CI
consumer.
**Proposed solution**
Pass `--bind lan` in the smoke container command (the flag is supported
by `latest` and canary alike; it selects the all-interfaces preset and
keeps the env-driven authenticated deployment), and pin the flag with a
wiring test so it cannot regress silently.
## What Changed
- `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard
--yes --bind lan --data-dir ...`, with a comment explaining why the flag
is load-bearing
- `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test
asserting the smoke Dockerfile pins a non-loopback bind preset
## Verification
- Full local harness run against the real nightly candidate
`2026.810.0-canary.1`: container healthy, bind banner shows `lan
(0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap
invite accepted, board session verified), `/api/health` returns
`bootstrapStatus: ready`
- `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4
pass
- After merge: dispatch `release.yml` with `channel: nightly` to run the
gate end to end in CI
## Risks
- Low. The change only affects the smoke container. `--bind lan` inside
a container exposes the port to the container network only; reachability
from outside still goes through Docker's explicit port mapping
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with
extended thinking and full tool use (CI log forensics, upstream source
tracing, local Docker reproduction and verification). All changes
model-authored under human direction.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending — will confirm before
merge)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending — will confirm before merge)
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release subsystem's nightly lane (#11006) gates every nightly
publish on the release smoke suite, which boots the published artifact
in a Docker container
> - The suite's first CI execution failed at the health readiness check:
the harness hard-codes a 90 second budget, but a CI container
cold-installs paperclipai from npm and initializes embedded postgres
with no warm caches
> - When the timeout expired with the container still running, the
harness printed no container logs, so the failure gave no diagnostics
> - This pull request makes the readiness budget configurable, raises it
for CI, and dumps container logs on timeout
> - The benefit is that the nightly gate measures the artifact, not the
runner's cold caches, and a red smoke run is diagnosable from its logs
## Linked Issues or Issue Description
**Subsystem affected**
Release smoke testing: `scripts/docker-onboard-smoke.sh`,
`.github/workflows/release-smoke.yml`.
**Problem or motivation**
Run 31426044332 (first forced nightly after #11006) failed in
`smoke_nightly` with `server did not become ready at
http://localhost:3232/api/health` after exactly 90 seconds. The
harness's readiness window is hard-coded to 90 attempts at 1 second.
Locally that works because the npm cache is warm; in CI the container
downloads the full package set and embedded postgres first. The timeout
path also printed no container logs when the container was still
running, so there was no way to see how far boot had progressed.
**Proposed solution**
Make the readiness budget an environment variable
(`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use),
set it to 420 in the CI workflow, and dump the last 150 container log
lines when the readiness check times out on a still-running container.
## What Changed
- `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env
var (default 90) replaces the hard-coded readiness budget; timeout with
a still-running container now prints the tail of `docker logs`
- `.github/workflows/release-smoke.yml`: sets
`SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs
## Verification
- `bash -n` on the harness and YAML parse of the workflow
- The real proof is the next `channel: nightly` dispatch of
`release.yml`, which re-runs this suite in CI with the new budget
## Risks
- Low. The local default is unchanged; CI runs simply wait longer before
declaring failure, and a genuinely broken artifact still fails (with
logs now)
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with
extended thinking and full tool use. Diagnosis from CI run logs; patch
model-authored under human direction.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending — will confirm before
merge)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending — will confirm before merge)
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Company import/export lets an operator move a full company package
between instances, with the Import page uploading the package as one
compressed `.zip`
> - The server caps that upload at 128 MB, and real company packages
with attachments now exceed it — imports fail at the preview step
> - The failure message tells the user to use the CLI folder import, but
that path posts inline JSON capped at 64 MB, so the advice is a dead end
for exactly these packages
> - This pull request raises the zip upload cap to a 1 GB default, makes
it operator-configurable through an environment variable, scales the
decompression-bomb guards from the cap in effect, and replaces the
misleading hint
> - The benefit is that large real-world company packages import
successfully, and operators with unusual needs can tune the cap without
a code change
## Linked Issues or Issue Description
**What happened?**
A company import fails at the preview step with `Preview failed: Import
package exceeds 134217728 bytes`. The package is a valid Paperclip
export. Its compressed size is larger than the 128 MB server cap (one
reported package is 257 MB). The error panel suggests the CLI folder
import, but that path sends the package as one inline JSON body capped
at 64 MB, so it also fails.
**Expected behavior**
A valid company package of realistic size imports successfully through
the Import page. If a package is too large, the error must state the
limit clearly and suggest a step that can work.
**Steps to reproduce**
1. Export a company with enough attachments to make the compressed
package larger than 128 MB.
2. Open the Import page and upload the `.zip`.
3. Click "Preview import".
4. The preview fails with `Import package exceeds 134217728 bytes`.
**Deployment mode**
Reported from a managed deployment; the limit applies to all deployment
modes.
## What Changed
- Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default
(`server/src/http/body-limits.ts`).
- Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid
or non-positive values fall back to the default.
- Scale the zip decompression-bomb guard from the configured cap at the
import route: the aggregate inflated ceiling is 4x the cap. The
per-entry ceiling stays at 512 MB because V8's string length limit
applies to an entry regardless (`server/src/routes/companies.ts`,
`packages/shared/src/portability-zip.ts`).
- Report the 422 limit error in MB instead of raw bytes.
- Replace the "use the CLI folder import for very large packages" hint
on preview failure with advice that works: re-export the package without
large attachments (`ui/src/pages/CompanyImport.tsx`).
- Update the stale comment in `ui/src/lib/import-preflight.ts` that made
the same CLI claim.
- Add tests for the new default, the env override, and the
invalid-override fallback.
## Verification
- `pnpm vitest run server/src/__tests__/body-limits.test.ts
packages/shared/src/portability-zip.test.ts
server/src/__tests__/company-portability-routes.test.ts
server/src/__tests__/company-portability.test.ts
server/src/__tests__/company-portability-import-batching.test.ts` — all
pass.
- `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes,
including the updated failure-panel copy assertion.
- `pnpm typecheck` — clean across the workspace.
- Manual: upload a `.zip` larger than the configured cap; the preview
fails with `Import package exceeds the 1024 MB upload limit` and the new
hint. A package between 128 MB and 1 GB now previews and imports.
## Risks
- Peak per-import memory rises with the cap: the upload is buffered in
memory and unzipped in one pass. A 1 GB compressed package can use
several GB transiently. Imports are instance-admin actions, so the
exposure is a deliberate operator action, not anonymous traffic.
Operators on small hosts can lower the cap with
`PAPERCLIP_IMPORT_ZIP_MAX_BYTES`.
- The aggregate bomb guard moves from a fixed 512 MB to 4x the
configured cap. It still bounds expansion far below what a decompression
bomb needs.
- No migration and no API shape change. The 422 message text changes; no
code matches on the old text.
## Model Used
- Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended
thinking and tool use (file edits, local test runs, live-instance
inspection).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release subsystem publishes the `paperclipai` npm package set
and the Docker images on two lanes: canary on every master push, and
stable on manual promotion
> - There is no middle ground between those lanes. Users must track
every merge or wait weeks for a stable. Docker `:latest` also tracks
master, so Docker users have no stable image at all
> - A calm prerelease lane needs to exist, and it must never ship a
build that failed its checks
> - This pull request adds the nightly channel: a scheduled job that
selects the newest master commit with a green canary publish, runs the
full release smoke suite against that exact published canary, and only
then republishes it as the nightly. It also separates Docker tags by
lane, so `:latest` finally means stable
> - The benefit is that users can follow prereleases at a nightly
cadence with a smoke-tested guarantee, and Docker users get real
`:canary`, `:nightly`, and stable image tags
## Linked Issues or Issue Description
**Subsystem affected**
Release automation: `scripts/release.sh`, `scripts/release-lib.sh`,
`.github/workflows/release.yml`, `.github/workflows/docker.yml`,
`.github/workflows/release-smoke.yml`.
**Problem or motivation**
The project publishes only `canary` (every master push) and `latest`
(manual stable). Users who want prereleases without per-merge churn have
no option. Docker has a second problem: master builds overwrite
`:latest`, and CI-published stables never produced Docker images,
because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger
in `docker.yml`. No stable-versioned image exists in ghcr today.
**Proposed solution**
Add a `nightly` channel. A scheduled job selects the newest
canary-tagged master commit, smoke-tests that exact published canary,
and republishes the same commit as `YYYY.MDD.P-nightly.N` under the
`nightly` dist-tag. Separate Docker tags by lane (`:canary` for master,
`:nightly` for nightly tags, `:latest` plus version tags for stable tags
only), and have the release jobs dispatch `docker.yml` at the new tag so
lane images actually build.
**Alternatives considered**
Moving the `nightly` dist-tag to the existing canary version without a
republish. Rejected: the version string would say `canary` while the
user is on nightly, which breaks at-a-glance lane identification in bug
reports and `--version` output.
## What Changed
- `scripts/release-lib.sh`: channel-parameterized
`next_prerelease_version` and `prerelease_tag_name` helpers (canary
helpers delegate to them), a `require_channel_tag_at_head` guard, and
the no-provenance retry for Sigstore transparency-log duplicates now
covers the `nightly` dist-tag as well as `canary`
- `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry
a `canary/v*` tag, publishes the full public package set as
`YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source
commit `nightly/vYYYY.MDD.P-nightly.N`
- `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) —
select candidate, smoke it via `release-smoke.yml`, publish on green
under the existing `npm-canary` environment, push the tag, dispatch
`docker.yml`. New `channel` dispatch input (default `stable`, so
existing stable dispatches are unchanged) with `nightly_source_version`
and `dry_run` support for forced runs. The stable path now also
dispatches `docker.yml` at the new `v*` tag
- `.github/workflows/docker.yml`: lane tag mapping for both image jobs —
master pushes publish `:canary` and no longer move `:latest`;
`nightly/v*` tags publish `:nightly`; only stable `v*` tags publish
`:latest` and the versioned tags. New `workflow_dispatch` trigger for
the release-job dispatches. Build-version stamping uses the exact
nightly version on nightly tag builds
- `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch
choice list
- `doc/CHANNELS.md` (new): user-facing guide to the channels
- `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping
table, and a nightly failure playbook
- `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses
`npm-canary` and needs no npm trusted-publisher changes
- Tests: channel-parameterized version helper coverage in
`scripts/release-registry-versions.test.mjs`, and nightly flow coverage
(publish identity, notes not required, canary-tag guard) in
`scripts/__tests__/release-dry-run-notes.test.mjs`
## Verification
- `node --test` on the release script suites: 68 pass, including 6 new
tests. The only failure, `acpx-patch-packaging.test.mjs`, needs
installed `node_modules` and fails identically on a pristine checkout of
master in the same environment
- `bash -n` on both shell scripts and YAML parse of all three workflows
- Live fail-path check: `./scripts/release.sh nightly --print-version`
from a master tip with no canary tag fails with `HEAD has no canary/v*
tag`
- Live success-path check: the same command from the
`canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0`
- Live selection check: the candidate-selection shell logic run against
the real repository selects the commit of `canary/v2026.806.0-canary.7`,
which matches the current npm `canary` dist-tag exactly
- After merge: dispatch `release.yml` with `channel: nightly` and
`dry_run: true` to preview, then a real forced run to validate end to
end before the first scheduled run
## Risks
- Docker `:latest` changes meaning from "latest master build" to "latest
stable release". This is deliberate and will be announced. Users who
want the old behavior pull `:canary`. Until the first stable release
after this change, `:latest` stays at its current (master-built) image
- The nightly is a rebuild of the same source commit, not the
byte-identical canary artifact that was smoked. The lockfile pins
dependencies, and the publish path's registry-visibility and
clean-prefix install gates still run on the nightly artifacts
- All npm publishing must stay inside `release.yml` because npm trusted
publishing pins that workflow file per package. The nightly jobs were
added to `release.yml` for exactly that reason; this constraint is now
documented in `RELEASING.md`
- The stable-lane Docker dispatch fails gracefully (a warning with
manual instructions) when the source ref predates `docker.yml`'s
`workflow_dispatch` trigger
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with
extended thinking and full tool use (repository exploration, local test
execution, live registry and git verification). All code, tests, and
docs in this PR were model-authored under human direction.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending — will confirm before
merge)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending — will confirm before merge)
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip orchestrates AI agents and relies on issue checkout as the
core task-claiming primitive
> - The issue checkout route is the HTTP boundary that translates
service and database outcomes into agent-usable API responses
> - Routine-linked issues are protected by the partial unique index
`issues_open_routine_execution_uq`, which covers only rows whose
`execution_run_id` is set
> - `svc.checkout` sets `execution_run_id`, so a concurrent claim moves
the row into that index and can raise a 23505 mid-request
> - Unhandled, that surfaces as a 500 and crashes the agent run instead
of being a recoverable conflict
> - Drizzle wraps driver failures in its own `Failed query: ...` error,
so the Postgres error carrying `code` and the constraint name is
reachable only through `cause`
> - This pull request translates that violation into a 409 at the
checkout route, detecting it through the cause chain the way
`isReviewPathRecoveryIdempotencyConflict` already does
> - The benefit is that agents handle routine execution contention
through the normal heartbeat conflict path instead of failing on an
internal server error
## Linked Issues or Issue Description
Fixes#3660
Related pull requests found while searching for duplicates:
- #3699 — an earlier attempt at this same route-level fix, closed
unmerged. Same shape, and its check has the flat-error bug described
under Verification.
- #3633 — related work on postgres.js `constraint_name` handling in
conflict detection.
- #5662 — covers the adoption path (`assertCheckoutOwner`) that this
pull request does not.
## What Changed
- Added `server/src/db-errors.ts` with `isUniqueViolation(error,
constraintName?)`, which walks the `cause` chain (depth-capped) and
accepts the postgres.js `constraint_name`, the node-postgres
`constraint`, or the driver message as evidence of SQLSTATE 23505.
- Wrapped `svc.checkout()` in `POST /issues/:id/checkout` with a narrow
try/catch that uses that helper to return **409 Conflict** for
`issues_open_routine_execution_uq`, and rethrows every other error
unchanged.
- Added `server/src/__tests__/db-errors.test.ts` covering the wrapped
and bare error shapes, both constraint field names, the message
fallback, non-matching constraints, non-unique-violation codes, and a
self-referential cause chain.
## Verification
- The new unit test includes the wrapped case `{ cause: { code: "23505",
constraint_name: ... } }` that a flat `error.code` check fails, so it is
a real regression guard rather than a restatement of the implementation.
- The wrapped shape is what this codebase observes in practice:
`server/src/__tests__/plugin-tenant-isolation.test.ts` asserts
`cause?.code === "23505"` against embedded Postgres,
`packages/db/src/pipelines-schema.test.ts` asserts that constraint
failures throw `Failed query`, and
`server/src/services/recovery/review-path-recovery.ts` walks the same
chain.
- CI (verify, e2e, policy) exercises this change against current master
through the pull request merge ref.
- Not verified locally: no monorepo install or typecheck was run in this
environment.
## Risks
- Low. One route gains a catch that matches a single constraint and
rethrows all other errors, so no unrelated failure can be swallowed.
- The 409 body `{ error: ... }` matches the other 409 responses this
route already returns.
- Scope limit: this covers the checkout route only. The adoption path
reached through `assertCheckoutOwner` (heartbeat, plugins, and pipelines
routes) can still surface the same violation as a 500; #5662 targets
that path.
- `isUniqueViolation` is new and intentionally generic. Existing flat
23505 checks elsewhere in the server are left untouched by this pull
request.
## Model Used
- Original change: OpenAI Codex, GPT-5-class tool-using coding agent in
the Codex CLI environment; exact backend model revision is not exposed
in that runtime.
- Follow-up revision (cause-chain detection plus tests): Anthropic
Claude Opus 5 (`claude-opus-5`), tool-using coding agent with extended
thinking and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
## What was done
Replaced the strict `!nonEmpty(process.env.PORT)` guard in
`maybePersistWorktreeRuntimePorts` with a new `isPortPinnedByRuntimeEnv`
helper function. This function checks if `process.env.PORT` is set, but
only suppresses persisting the port to configuration if the ambient
`PORT` matches the newly allocated `selectedPort`.
## Why it matters
Fixes issue #1849. Previously, if an ambient `PORT` environment variable
was exported globally (like inheriting from the shell running the parent
workspace), worktrees would silently fail to write their
collision-avoiding ports (e.g. 3103 instead of 3100) back to their
respective local `config.json` files. This resulted in orphaned
sub-worktrees and lost port tracking on reboot. With this fix, worktrees
correctly persist their assigned ports even while nested under an
inherited environment variables stack, while continuing to respect
manual, explicit pinning.
## How to verify
1. Export a port in the shell explicitly: `export PORT=3100`.
2. Launch a sub-worktree instance which receives an auto-assigned free
port (e.g., `3103`).
3. View the underlying `config.json` for that worktree inside
`.paperclip/worktrees/`.
4. The config file should correctly contain `{"server": {"port": 3103}}`
rather than dropping the write operation.
## Risks
None expected. The `Number()` and `Number.isInteger()` checks handle
parsing edge cases cleanly, defaulting robustly to preventing writes if
`process.env.PORT` is somehow malformed (e.g., set to a non-integer),
ensuring absolute safety during misconfigurations.
Co-authored-by: manavshrivastavagit <manavshrivastava@users.noreply.github.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The agent detail page has a Costs section. That section renders a
data table of per-run spend.
> - The table has a header row, but its `<th>` elements carry no `scope`
attribute.
> - A screen reader uses `scope="col"` to bind each data cell to its
column header. Without it, the reader announces a number without telling
the user which column it belongs to.
> - A table of costs is exactly the case where that hurts. Every cell is
a bare figure.
> - This pull request adds `scope="col"` to the five header cells in
that table.
> - The benefit is that assistive technology announces the cost table
correctly. The change is markup only, so sighted users see no
difference.
## Linked Issues or Issue Description
No public GitHub issue covers this. The problem is described in-PR,
following the enhancement template.
**What existing behavior does this improve?**
The Costs table rendered by `CostsSection` in
`ui/src/pages/AgentDetail.tsx`.
**Subsystem affected**
ui/ — React + Vite board UI
**Current behavior**
The table renders five header cells: Date, Run, Input, Output, and Cost.
None of them set `scope`. A screen reader must guess the header-to-cell
relationship, so a user hears a value with no column name attached to
it.
**Proposed behavior**
Each header cell sets `scope="col"`. A screen reader then announces the
column name together with each cell, so a cost figure is read as part of
the Cost column.
**Reason and benefit**
`scope` is the standard way to associate header cells with data cells in
an HTML table. The attribute has no visual effect, so the fix carries no
design cost and makes the table usable with a screen reader.
**Breaking changes**
None. `scope` is a presentational-neutral HTML attribute. No component
API, no styling, and no test changes.
**Related pull requests**
- #2215 proposed the same attribute for the Routines table. It is
closed, because that table no longer exists on master.
- #1524 and #1522 applied `scope="col"` to other tables. Both are
closed.
## What Changed
- Added `scope="col"` to the five `<th>` elements in the `CostsSection`
table in `ui/src/pages/AgentDetail.tsx`.
- Rebased the branch onto current master.
- Dropped the original `ui/src/pages/Routines.tsx` hunks. Master rebuilt
the Routines page around folder-grouped rows, so the table those hunks
targeted no longer exists.
- Dropped the original `HintIcon` opacity change. It altered a visible
colour, which is out of scope for a markup-only accessibility fix.
## Verification
- Run `pnpm --filter @paperclipai/ui typecheck` and `pnpm --filter
@paperclipai/ui build`. This is a markup-only change, so a clean
type-check and build is the relevant automated signal.
- Open an agent detail page and go to the Costs section. Inspect the
header row. Each `<th>` now carries `scope="col"`.
- Navigate the same table with a screen reader, cell by cell. Each cell
is announced with its column name.
- Compare the rendered page before and after. It is unchanged, because
`scope` has no styling effect.
## Risks
Low risk. The change adds one standard HTML attribute to five header
cells in a single table. It introduces no code path, changes no
component API, and has no visual effect. The worst case is that the
attribute is redundant for a reader that already infers the column,
which is harmless.
## Model Used
Anthropic Claude Opus 5, exact model ID `claude-opus-5`. It ran with
extended thinking and repository read/write tools, inside a
maintainer-operated triage agent. The model rebased the branch, dropped
the two out-of-scope hunks, and wrote this description. The original
change was authored by @bluzername, and the model used for that work is
not recorded here.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` references)
- [x] I have considered and documented any risks above
- [x] My branch name describes the change and contains no internal
ticket id
- [ ] I have run tests locally and they pass — not run. This is a
markup-only change and the package has no test covering this table.
- [ ] I have added or updated tests where applicable — no test added,
which is why this PR is titled `refactor:`.
- [ ] I have updated relevant documentation to reflect my changes — no
documentation describes this markup.
- [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work — not checked by the maintainer who rebased this.
- [ ] All Paperclip CI gates are green — CI re-runs on this push.
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
Greptile re-reviews on this push.
---------
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip posts system comments when automatic run recovery cannot
continue
> - These comments currently mix the main event with recovery
identifiers and routing details
> - The task chat shell also renders these comments as large raw text
blocks
> - Operators need a short explanation first and inspectable evidence on
demand
> - This pull request emits structured recovery notices and renders them
as compact humanized rows
> - The benefit is a quieter task thread that keeps the full recovery
evidence available
## Linked Issues or Issue Description
Related prior extraction source: #11070. This pull request replaces only
its structured recovery notice slice with a focused branch based on
current master.
**What existing behavior does this improve?**
Paperclip recovery escalations and the experimental task chat
system-comment renderer.
**Current behavior**
Recovery escalation comments put action identifiers, owner details, run
details, and failure codes into the visible markdown body. The task chat
shell renders the complete system comment as a large text block.
**Proposed behavior**
The server emits a short system notice with typed metadata sections. The
task chat shell classifies known recovery families and renders one
compact row. An operator can expand the row to inspect the full body and
metadata.
**Reason and benefit**
The main thread stays readable during repeated recovery activity. Typed
links and evidence remain available without exposing raw failure text in
the default view.
**Breaking changes**
The visible recovery comment body is shorter. Recovery action
deduplication now reads the structured metadata and still recognizes
legacy body markers. No API schema or database migration changes.
## What Changed
- Emit stranded recovery escalations with `system_notice` presentation
and typed recovery, owner, run, and failure-code metadata.
- Share bounded metadata row builders across recovery notice producers
and preserve legacy deduplication compatibility.
- Humanize known recovery notice families and render compact expandable
task-chat rows.
- Route system-authored comments ahead of derived agent authorship so
recovery notices do not appear as agent bubbles.
- Add focused server and UI regression coverage.
## Verification
- `pnpm check:token-gates` — 3/3 clean.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/shared exec vitest run
src/validators/issue.test.ts` — 32 tests passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/services/recovery/stranded-notice.test.ts
src/__tests__/issue-recovery-actions.test.ts` — 57 tests passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/services/recovery/successful-run-handoff.test.ts
src/services/recovery/stranded-notice.test.ts` — 39 tests passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/heartbeat-process-recovery.test.ts -t 'escalates an
exhausted failed successful-run handoff without using generic
continuation recovery first|escalates an exhausted successful handoff
run that still leaves no disposition'` — 2 tests passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/heartbeat-process-recovery.test.ts -t 'blocks assigned
todo work after the one automatic dispatch recovery was already used'` —
passed.
- `pnpm --filter @paperclipai/ui exec vitest run
src/lib/system-notice-humanizer.test.ts
src/components/task-chat/TaskChatSystemNotice.test.tsx
src/components/task-chat/task-chat-adapter.test.ts` — 15 tests passed.
- Storybook visual baselines were not updated because this chat-shell
path has no affected snapshot baseline. Focused rendering tests and
token gates cover this change.
## Risks
- Consumers that parse recovery action identifiers from comment markdown
must move to structured metadata. Server deduplication remains backward
compatible with legacy comments.
- The humanizer uses stable recovery-family phrases. Unknown notices use
a generic truncated first-sentence fallback.
- The UI changes only the experimental task chat presentation. The
stored comment body and expanded metadata remain available.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex with GPT-5, reasoning mode, repository tools, shell
execution, and GitHub integration. The runtime did not expose a
context-window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - People send task instructions through the board task chat.
> - A page refresh or task switch can discard an unfinished message in
the redesigned composer.
> - The existing task chat already supplies a task-specific draft key.
> - The redesigned composer must use that key without changing
attachment or send behavior.
> - This pull request restores, saves, and clears text drafts in the
redesigned composer.
> - The benefit is that users can return to unfinished task messages
without losing their text.
## Linked Issues or Issue Description
Related prior work: #11070. This pull request extracts only the final
composer draft behavior from that larger draft.
**Subsystem affected**
ui/ — React + Vite board UI.
**Problem or motivation**
The redesigned task chat composer does not use the draft key that the
task thread already provides. A refresh, navigation, or unmount can lose
an unfinished message.
**Proposed solution**
Persist text drafts by task key in local storage. Restore a draft when
the composer mounts. Save changes after a short delay and flush pending
text during unload or unmount. Clear the draft only after a successful
send.
**Alternatives considered**
The composer could save on every keystroke. A short delay avoids
unnecessary synchronous storage writes. The feature could also stay in
the larger predecessor PR, but a focused PR is easier to review and
verify.
**Roadmap alignment**
This is a focused usability improvement for the task conversation
surface. It does not add or duplicate a roadmap capability.
## What Changed
- Added safe draft storage helpers for load, save, and clear operations.
- Connected the task-specific draft key to the redesigned task chat
composer.
- Preserved drafts across debounce windows, unmounts, page unloads,
failed sends, and React Strict Mode probes.
- Cleared drafts after successful sends without changing current
attachment safeguards.
- Added focused composer and thread integration tests.
## Verification
- `pnpm check:token-gates`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/task-chat/TaskChatComposer.test.tsx
src/components/TaskChatThread.test.tsx`
## Risks
- Local storage can be unavailable or full. The helpers catch storage
errors and keep the composer usable.
- Only text is persisted. Attachments, work mode, and assignee
selections remain session state.
- Draft keys remain task-scoped, so text does not cross task boundaries.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex with model `gpt-5`. The context-window size is not exposed
in this environment. The model used agentic reasoning, tool use, code
execution, and test execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Developers can run Paperclip from linked Git worktrees.
> - The server development watcher scans paths near the active checkout.
> - A main checkout can contain many complete sibling worktrees under
`.paperclip/worktrees`.
> - Scanning those sibling checkouts can stall the watcher before it
starts the server.
> - This pull request excludes the shared worktree directory from the
development watcher.
> - The benefit is that development startup stays responsive as the
number of worktrees grows.
## Linked Issues or Issue Description
**What happened?**
The server development watcher traversed sibling checkouts under
`.paperclip/worktrees`. Large worktree collections could make `pnpm dev`
stall before the watcher started the server process.
**Expected behavior**
The watcher must observe only source paths that can reload the active
checkout. It must ignore sibling worktrees in both a main checkout and a
linked worktree.
**Steps to reproduce**
1. Create several linked worktrees under `.paperclip/worktrees`.
2. Add normal dependency and build output trees to those worktrees.
3. Run `pnpm dev` from the main checkout or one linked worktree.
4. Observe the watcher scan sibling worktrees before it starts the
server.
**Paperclip version or commit**
Reproduced on `master` before this change.
**Deployment mode**
Local development with `pnpm dev`.
## What Changed
- Detect whether the active server root is inside the managed
linked-worktree directory.
- Ignore the shared `.paperclip/worktrees` root from both main and
linked checkouts.
- Add regression coverage for the resolved ignore path and its globstar
form.
## Verification
- `./node_modules/.bin/vitest run
server/src/__tests__/dev-watch-ignore.test.ts --reporter=verbose`
- `pnpm --filter @paperclipai/server typecheck`
## Risks
- Low risk. The change affects only local development watch exclusions.
- A non-standard checkout that copies the same `.paperclip/worktrees`
directory layout will receive the same exclusion.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5, context window not disclosed, with reasoning,
tool use, and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Humans oversee those agents in teams, so each team needs its own
spend controls
> - The BudgetPolicyCard component shows how much of a budget is
consumed
> - The utilization bar in that card is a styled div with no ARIA role,
value, or label
> - A screen reader user therefore cannot hear how much budget is used
> - This pull request adds role="progressbar" and the matching ARIA
value attributes
> - The benefit is that assistive technology announces budget
utilization the same way sighted users see it
## Linked Issues or Issue Description
No existing issue. Related pull requests: #1869 and #1878 add the same
progressbar
semantics to ProviderQuotaCard and QuotaBar. They touch different files.
The problem follows the bug report template:
**What happened?**
The budget utilization bar in `BudgetPolicyCard` renders as a plain
`div`. It has no
`role`, no `aria-valuenow`, and no accessible name. A screen reader
announces nothing
for it. The user can read the "Remaining" amount, but not the
utilization percentage.
**Expected behavior**
The bar is announced as a progress bar. It reports the current
utilization percentage,
with its minimum and maximum.
**Steps to reproduce**
1. Open a project or agent page that shows the budget card.
2. Start VoiceOver (Cmd+F5 on macOS).
3. Move the cursor to the budget utilization bar.
4. VoiceOver announces nothing.
**Paperclip version or commit**
master, `ui/src/components/BudgetPolicyCard.tsx`.
**Deployment mode**
Local dev (pnpm dev).
## What Changed
- Added `role="progressbar"` to the inner bar element.
- Added `aria-valuenow` with the rounded utilization percentage.
- Added `aria-valuemin={0}` and `aria-valuemax={100}`.
- Added `aria-label` in the form `Budget utilization: 73% used`.
`aria-valuenow` and `aria-label` use the same `progress` value. That
value is already
capped at 100 by `Math.min(100, summary.utilizationPercent)`, so the
reported value
stays inside the min/max range when a scope is over budget.
An earlier revision also changed the budget amount `Input` to
`type="number"`. That
change is removed. It changed input behavior and was not an
accessibility fix.
See "Risks".
## Verification
1. Open a page that shows the budget card.
2. Start VoiceOver (Cmd+F5 on macOS) and move to the utilization bar.
3. VoiceOver announces "Budget utilization: X% used, progress
indicator".
4. Inspect the element. `aria-valuenow` equals the displayed percentage,
and it stays
at 100 when utilization is above 100%.
The change adds attributes only. There is no visual change.
## Risks
Low risk. The change adds ARIA attributes to one element in one
component. It changes
no logic, no styling, and no layout.
The `type="number"` change is removed on purpose. With `type="number"`,
a browser
reports an empty string for text it cannot parse. `parseDollarInput("")`
then returns
`0` instead of `null`, so the existing "Enter a valid non-negative
dollar amount."
error never appears and unparseable input silently becomes a $0.00
budget. The
`inputMode="decimal"` input keeps that validation path working.
## Model Used
Original implementation by @bluzername. The author did not state a
model.
Scope reduction, branch update, and this description: Claude Opus 5 (1M
context,
extended thinking, tool use), prepared under maintainer review.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change
- [x] I have considered and documented any risks above
- [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
## Thinking Path
> - Paperclip runs AI agents through adapter execution services.
> - The adapter runtime records spans for each stage of agent startup.
> - The managed runtime packs workspace tarballs before it uploads them.
> - These pack operations had no host span, so `stage.sync` omitted pack
time.
> - This pull request adds one `pack` span around both tarball builds.
> - The span nests under the active `stage.sync` step and improves trace
detail.
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The managed runtime workspace sync builds a git-history tarball and a
workspace-overlay tarball before upload.
**Subsystem affected**
The change affects `packages/adapter-utils`, which provides adapter
execution and managed runtime support.
**Current behavior**
The host builds both tarballs without an OpenTelemetry span. The
`stage.sync` trace therefore omits the host pack duration.
**Proposed behavior**
The host wraps both tarball builds in one `pack` span. The executor
parents this span under the active startup step.
**Reason and benefit**
The trace shows the time that the host spends packing workspace data.
Operators can use the existing runtime span tree to find sync delays.
**Breaking changes**
None. The default span runner remains a no-op runner, and the existing
control flow remains unchanged.
## What Changed
- Add an optional `runtimeSpan` runner to the managed runtime
preparation path.
- Create one host `pack` span around the two workspace tarball builds.
- Parent the `pack` span under the active startup step.
- Add unit coverage for span emission and span nesting.
## Verification
- Run `tsc --noEmit` for `@paperclipai/adapter-utils`.
- Run the `@paperclipai/adapter-utils` Vitest suite.
- Confirm the suite reports 448 passed tests and 4 skipped tests.
- Confirm the trace test records `pack` under `stage.sync`.
## Risks
The change adds optional tracing only. The default no-op runner
preserves behavior when tracing is not configured.
## Model Used
OpenAI Codex, GPT-5, tool use and code execution. The model reviewed the
handoff and opened this pull request. The implementation author supplied
the commit.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.
Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
This reverts commit 11e56654f8.
#10786 ported the onboarding flow from a standalone prototype and
repointed `/onboarding` at the new `CloudOnboardingFlow`, deleting the
existing `OnboardingWizard.tsx` in the process. The ported flow is not
ready to be the shipping onboarding experience: it landed as a single
large port rather than an incremental migration, it pulled `motion`,
`three` and `@types/three` onto the UI dependency list for prototype
visuals, and it deleted the wizard that four in-flight pull requests
(#9900, #9501, #8982 and one more) were building on — those went
CONFLICTING the moment the file disappeared.
Rather than keep the half-migrated state on master while that is sorted
out, back the port out whole and re-land it incrementally. This restores
`OnboardingWizard.tsx` and the previous versions of the four e2e specs,
drops the `onboarding-preview.html` Vite entry, the DesignGuide
onboarding section and the `data-viz-misc` storybook story, and removes
the three prototype dependencies from `ui/package.json`.
This is an exact mechanical inverse of the squash commit — 41 files,
+2089/-3647, no hand edits. Reverting this commit restores all 41 files
byte for byte, so the port is recoverable in full when it is ready.
`pnpm-lock.yaml` is deliberately not touched. #10786 never updated it;
bot commit 4683f26c9 (#11036) added the `motion`/`three` entries
afterwards, so the lockfile is now ahead of the manifest. CI owns
lockfile updates (`.github/workflows/pr.yml`) and the policy job
regenerates it from the changed manifest.
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app that manages AI agents for work
> - Sandbox providers let agents run in remote and isolated environments
> - Daytona session commands need a path that sends agent output to the
host without host polling
> - Host polling adds delay and repeats provider output work
> - This pull request adds typed execute.log notifications and a log
sink for incremental output
> - This pull request adds an optional ACP session stream with
final-result replay protection
> - The benefit is lower output delay while the default flags keep
current behavior unchanged
## Linked Issues or Issue Description
**Subsystem affected**
Cross-cutting. The change spans the plugin SDK, Daytona provider,
adapter utilities, and server execution services.
**Problem or motivation**
The Daytona ACP bridge polls a host output file while an agent command
runs. This adds delay and can repeat work. The host also needs a safe
route for provider output chunks.
**Proposed solution**
Add a typed `execute.log` notification with host-issued invocation
correlation. Add an ordered log sink to the environment execute path.
Add an optional ACP session-log path that parses newline-delimited JSON
frames and removes the host output poll for that path.
**Alternatives considered**
Keep the output-file poll as the only path. This keeps the current
behavior but does not provide timely output. The new path stays behind
flags, so the existing path remains the default fallback.
**Roadmap alignment**
This change supports the shipped Cloud / Sandbox agents milestone in
`ROADMAP.md`, including Daytona support.
## What Changed
- Add the typed `execute.log` worker-to-host notification and
company-scoped host route.
- Add ordered `stdout` and `stderr` chunk delivery before the final
execute result.
- Add the Daytona session log sink and the optional ACP streamed session
path.
- Add monotonic frame handling so live and final output reach the host
once.
- Keep `useLogStream` and `streamAgentSessionOutput` off by default.
- Add unit and integration coverage for the notification, execution
target, runtime, and Daytona paths.
## Verification
- Run adapter-utils tests: 445 tests pass locally.
- Run server environment tests: 73 tests pass locally.
- Run Daytona plugin tests: 131 tests pass locally.
- Run TypeScript checks for shared, adapter-utils, and server.
- Review the pull request checks after GitHub completes them.
- All required GitHub checks pass on the current head.
## Risks
The new paths change output delivery only when a feature flag enables
them. The final execute result remains available for parsing and
fallback. The main risk is a provider stream or frame-order error; the
final-result parser limits that risk.
## Model Used
OpenAI Codex, GPT-5, tool use and code execution. The runtime did not
supply a context-window value.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
No operator documentation change applies because both new flags remain
disabled by default.
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps subsystem connects company tools through governed provider
connections.
> - A company can need more than one account for the same provider.
> - The current database constraint and Apps flow assume one named
connection per company.
> - New quarantined actions also need an explicit review decision before
activation.
> - This pull request supports multiple provider connections and
complete action review decisions.
> - The benefit is safer access control and a clear multi-account Apps
workflow.
## Linked Issues or Issue Description
Refs: #11040
**Subsystem affected**
Cross-cutting. This change affects the Apps UI, the tool access API, the
shared request contract, and the database schema.
**Problem or motivation**
The connection name constraint prevents a company from keeping more than
one connection for a provider. The Apps UI also reuses an existing OAuth
connection when a user asks to connect another account. Action review
can enable selected entries without recording a decision for every
quarantined action.
**Proposed solution**
Remove the company and connection name uniqueness constraint. Let users
open, count, edit, and create multiple provider connections. Require the
finish request to cover every quarantined action exactly once before the
server activates reviewed entries.
**Alternatives considered**
The UI could generate unique internal names and keep the database
constraint. This would preserve a one-connection assumption in the data
model and would make display names part of identity. The server could
also infer review decisions from enabled actions. This would not
distinguish a reviewed disabled action from an action that the user did
not review.
**Roadmap alignment**
This change extends the completed MCP Tool Gateway and Apps milestone.
It also supports the Connected Apps roadmap item. It follows the
navigation and connection management work in #11040.
## What Changed
- Remove the company-scoped connection name uniqueness index with an
ordered and idempotent migration.
- Add a reviewed action list to the finish-app contract and reject
incomplete or duplicate review decisions.
- Activate reviewed entries and keep unreviewed quarantined entries
blocked.
- Enable a completed connection and preserve the company and connection
scope in all updates.
- Show provider connection counts and open the provider setup page from
Browse.
- Let users edit existing connections or connect another account without
reusing an active OAuth connection.
- Update focused server and UI coverage for multiple connections and
action review.
## Verification
- Ran the focused Apps UI suite. All 116 tests passed in 11 files.
- Ran the focused server and CLI suite. All 276 tests passed in 3 files.
- Ran `pnpm --filter @paperclipai/db check:migrations`. The migration
safety check passed.
- Ran `pnpm -r typecheck`. All projects passed.
- Ran `pnpm build`. All projects built successfully.
- Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One
worktree-safety assertion failed because the execution workspace reloads
its worktree marker. The same test passed with an isolated non-worktree
marker.
- Ran `pnpm check:token-gates`. It reports 12 existing violations in the
unchanged `PaperclipOrbit3D.tsx` file from the target branch.
- Started the six affected Playwright specifications. Chromium could not
start because the host does not provide `libatk-1.0.so.0`. The GitHub
e2e jobs will verify these specifications.
- GitHub Actions passed every final-head CI gate, including all three
e2e shards and the aggregate `e2e` and `verify` jobs.
- Greptile reviewed final commit `9af9200426` at 5/5 with zero review
threads.
## Risks
- Removing the name uniqueness index permits duplicate display names.
Stable connection IDs and UIDs remain unique within a company.
- The finish-app endpoint accepts the new review field as optional for
backward compatibility. When clients send it, the server requires a
complete decision for all quarantined actions.
- Multiple OAuth connections depend on the explicit new-connection route
flag. Focused tests cover active and draft connection reuse.
- The migration is ordered after migration 0210. Its `DROP INDEX IF
EXISTS` statement is safe to repeat.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The
agent used repository tools, code execution, test execution, and agentic
reasoning. The Codex runtime manages the context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The Skills Store lets a company find and install reusable agent
procedures
> - The MCP integration preparation procedure existed outside the app
catalog
> - Paperclip users could not find or install that procedure from the
product
> - This pull request adds the procedure as an optional software
development skill
> - The skill keeps research, human approval, and connector delivery as
separate gates
> - The benefit is a repeatable and governed path from vendor research
to one connector pull request
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The app-shipped Skills Store can install optional skills, but it does
not include the MCP integration preparation workflow from
`paperclip-content`.
**Subsystem affected**
`packages/skills-catalog`.
**Current behavior**
An agent must know where the external workflow lives. The agent cannot
find or install it from the Paperclip skills catalog.
**Proposed behavior**
The optional catalog includes `prepare-mcp-integration`. The installed
skill directs agents through cited research, a research-only content
pull request, an exact-revision human gate, and one Paperclip connector
pull request per approved connection.
**Reason and benefit**
This change makes the existing integration and connector playbooks
available as one installable Paperclip workflow. It also prevents agents
from starting connector code before the research gate is approved.
**Breaking changes**
None. The skill is optional and markdown-only.
Related source work: paperclipai/paperclip-content#13.
## What Changed
- Add the optional `prepare-mcp-integration` catalog skill under
software development
- Add Paperclip catalog metadata for roles, requirements, tags, and
trust classification
- Add a worked Notion MCP research-gate example
- Regenerate the checked-in catalog manifest
- Add the new key to the shipped optional skill test
## Verification
- `pnpm --filter @paperclipai/skills-catalog build:manifest`
- `pnpm --filter @paperclipai/skills-catalog validate`
- `pnpm --filter @paperclipai/skills-catalog test`
- Confirm the generated catalog contains
`paperclipai/optional/software-development/prepare-mcp-integration`
- Confirm the trust level is `markdown_only` and compatibility is
`compatible`
## Risks
- Low risk. This change adds one optional markdown-only catalog entry.
- The workflow can become stale if the two upstream playbooks change.
The skill requires agents to read the current playbooks before each
phase and to update upstream rules when reusable specifications change.
## Model Used
OpenAI Codex with GPT-5.4, reasoning mode, shell tool use, and code
editing.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip skills guide agents through repeatable work.
> - The plan-to-task skill guides agents when they create issue graphs
from plans.
> - The prior guidance could encourage an issue for every concrete
deliverable.
> - That guidance can create unnecessary subtasks for work that one
owner can complete end to end.
> - This pull request defines the boundaries that justify a separate
task.
> - It also adds a merge-back pass that removes task splits without a
qualifying boundary.
> - The benefit is a smaller issue graph with clear ownership,
dependencies, and review gates.
## Linked Issues or Issue Description
**Issue type**
Unclear or confusing documentation.
**Where is the issue?**
`skills/paperclip-converting-plans-to-tasks/SKILL.md`
**What's wrong?**
The skill says that each concrete deliverable must become an issue. This
can make agents split one end-to-end job into tasks for each step, file,
component, or phase. The result is more coordination work without a real
execution boundary.
**Suggested fix**
Tell agents to start with one end-to-end task. Permit separate tasks
only for ownership, parallel work, dependencies, independent review or
approval, or substantial follow-up work. Require a merge-back pass
before agents create the issue graph.
## What Changed
- Added a rule to use the fewest tasks that can complete and verify the
work.
- Defined the boundaries that qualify work for a separate issue.
- Added a merge-back pass for proposed subtasks without a qualifying
reason.
- Updated dependency, parallel work, verification, and checklist
guidance to enforce the smaller graph.
## Verification
- Ran `pnpm exec vitest run packages/shared/src/frontmatter.test.ts`.
- Result: 1 test file passed and 21 tests passed.
- Ran `git diff --check public-gh/master...HEAD`.
- Confirmed that the PR changes one skill file and has no lockfile,
workflow, image, or migration changes.
## Risks
- Low risk. This change updates skill guidance only.
- Agents might combine work too aggressively. The qualifying boundaries
preserve separate ownership, parallel execution, dependencies, review,
approval, and substantial follow-up work.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI GPT-5 through Codex. The exact deployment ID and context window
are not exposed. The agent used reasoning, repository tools, command
execution, and GitHub tools.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps UI manages app discovery and app connections.
> - The managed worktree runtime starts agent work in repository
worktrees.
> - The Apps routes do not match the main discovery flow, and the
connections view lacks a delete action.
> - Legacy managed worktrees can also start before their pending seed
operation runs.
> - This pull request makes app discovery the main Apps route and makes
connection management explicit.
> - It also seeds legacy managed worktrees before runtime startup and
makes the CLI read the repository-local config.
> - The benefit is a clearer Apps workflow and a safer managed-worktree
startup path.
## Linked Issues or Issue Description
**What existing behavior does this improve?**
This improves the Apps navigation, app connection management, managed
git-worktree startup, and CLI worktree selection.
**Subsystem affected**
Cross-cutting. The change affects `ui/`, `server/`, `cli/`, and
development documentation.
**Current behavior**
The `/apps` route opens the connections list while discovery uses a
nested route. The connections list has no delete action. Some legacy
managed worktrees can start runtime work before their pending seed
operation runs. The CLI can also read an ambient Paperclip config
instead of the repository-local config.
**Proposed behavior**
The `/apps` route opens Browse, and `/apps/connections` opens the
connection list. Users can delete a connection after confirmation.
Runtime startup seeds legacy managed worktrees when required. The CLI
resolves the current worktree from the repository-local
`.paperclip/config.json` file.
**Reason and benefit**
Users can discover apps from the canonical Apps route and can manage
existing connections from a dedicated route. Legacy worktrees receive
their required repository content before agent runtime starts. CLI
worktree selection stays scoped to the current repository.
**Breaking changes**
The `/apps` and `/apps/browse` route behavior changes. Old Browse links
redirect to `/apps`. The change does not modify an API schema or
database schema.
## What Changed
- Make Browse the canonical `/apps` page and move the connection list to
`/apps/connections`.
- Align Apps navigation, redirects, attention links, empty states, and
connection actions with the new routes.
- Add connection deletion with confirmation and clear failure feedback.
- Seed legacy managed git worktrees before runtime startup when their
seed status is pending.
- Read the CLI worktree selection from the repository-local Paperclip
config.
- Update focused UI, server, CLI, and development documentation
coverage.
## Verification
- Ran 202 focused UI, server, and CLI tests. All tests passed.
- Ran `pnpm -r typecheck`. All projects passed.
- Ran `pnpm build`. All projects built successfully.
- Ran `pnpm test:run`. The server and UI stages passed 7,168 tests. The
CLI stage found one environment-sensitive secrets test because this
workspace injects static AWS credentials. The isolated CLI file passed
all 8 tests after those injected variables were unset.
- Ran `pnpm check:token-gates`. It reports 12 existing color-token
violations in the unchanged `PaperclipOrbit3D.tsx` file from the target
branch. This pull request does not modify that file.
- Ran focused regression coverage for repository-root CLI config
resolution and connection deletion state. All tests and affected package
typechecks passed.
- Collected all 27 tests in the six changed Playwright specifications
successfully.
- GitHub Actions passed every latest-head CI gate, including all three
e2e shards and the aggregate `e2e` and `verify` jobs.
- Greptile reviewed the final commit at 5/5 with zero unresolved
threads.
## Risks
- Existing bookmarks for `/apps/browse` redirect to `/apps`.
- Connection deletion changes visible connection state and requires user
confirmation.
- The legacy seed path runs only for managed git worktrees with pending
seed state. Tests cover the startup condition.
- The rebase preserves the target branch's direct OAuth policy for the
Notion connection flow.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex with the GPT-5 model family assisted this change. The agent
used reasoning, repository tools, code execution, and test execution.
The runtime does not expose the exact model snapshot or context-window
size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Each run executes inside a persisted **execution workspace** (a row
in `execution_workspaces`) that is either freshly created or
**restored/reused** across runs of the same issue
> - Before adapter launch, a guard rejects a restored workspace whose
`projectWorkspaceId` is null while the issue resolves a concrete project
workspace (`persisted_workspace_missing_project_workspace_id`) — a
safety check against binding a run to a workspace with no
project-workspace link
> - The reuse/**restore** path updated the existing row (cwd, branch,
status, metadata…) but **never set `projectWorkspaceId`**, so a row
persisted with a null value stayed null on every restore
> - Result: for an issue that resolves a project workspace, the guard
fires, `reuse_existing` re-selects and re-binds the *same* stale null
row on the next attempt, and the run crash-loops forever with no
self-heal
> - This pull request backfills `projectWorkspaceId` during restore
(prefer the existing binding, fall back to the resolved one) so the row
heals on first reuse and the guard stops firing
> - The benefit is that reused workspaces created before their project
had a primary project workspace self-repair on next use instead of
crash-looping, while genuine mismatches are still surfaced by the guard
## Linked Issues or Issue Description
No public GitHub issue exists — describing the bug inline per the bug
report template (`.github/ISSUE_TEMPLATE/bug_report.yml`):
### What happened?
In `heartbeatService`, the execution-workspace reuse/restore branch
calls
`executionWorkspacesSvc.update(reusableExistingExecutionWorkspace.id, {
… })` without a `projectWorkspaceId` field. Only the sibling CREATE
branch sets `projectWorkspaceId`. So an execution workspace that was
persisted with a null `projectWorkspaceId` (e.g. created before its
project had a primary project workspace) is never backfilled on restore.
When such a workspace is later reused for a run whose issue resolves a
concrete project workspace, the pre-launch guard throws
`persisted_workspace_missing_project_workspace_id`, the run fails, and
`reuse_existing` re-binds the identical stale row on the next attempt —
an unbounded crash-loop with no self-heal.
### Expected behavior
On restore, the reused workspace's `projectWorkspaceId` is backfilled
from the resolved project workspace when it is currently null, so the
guard passes and the run launches. An existing non-null binding is never
overwritten (a genuine mismatch is still surfaced by the separate
`project_workspace_mismatch` guard).
### Steps to reproduce
1. Have an `execution_workspaces` row with `project_workspace_id = NULL`
that is eligible for reuse.
2. Give its project a primary project workspace (so the issue now
resolves a concrete `projectWorkspaceId`).
3. Dispatch a run for an issue in that project that reuses the
workspace. The restore `update()` leaves `project_workspace_id` null,
the launch guard throws
`persisted_workspace_missing_project_workspace_id`, and every subsequent
reuse re-binds the same null row and fails identically.
### Paperclip version or commit
`master` (branched from `14f20be92`); reproduced on a live self-hosted
instance.
### Deployment mode
Self-hosted, embedded Postgres, local adapters.
## What Changed
- New exported pure helper
`reconcileReusedExecutionWorkspaceProjectWorkspaceId(existing,
resolved)` in `server/src/services/heartbeat.ts`, returning `existing ??
resolved ?? null`. It prefers an existing binding (never nulls out a
good value or silently rebinds a genuine mismatch — the guard still
surfaces those), backfills a null binding from the resolved value, and
stays null when neither is present.
- Wire the helper into the reuse/restore
`executionWorkspacesSvc.update(...)` call so the restored row's
`projectWorkspaceId` is set to
`reconcileReusedExecutionWorkspaceProjectWorkspaceId(reusableExistingExecutionWorkspace.projectWorkspaceId,
resolvedProjectWorkspaceId)`. The CREATE branch already set
`projectWorkspaceId`; this brings the restore branch to parity.
## Verification
- Added 3-case unit coverage in
`server/src/__tests__/heartbeat-workspace-session.test.ts` for the
helper: (a) backfills a null existing binding from the resolved value,
(b) never overwrites an existing binding even when a resolved value is
present, (c) returns null when both existing and resolved are absent
(null and undefined inputs).
- Confirmed the `update()` patch type accepts the field:
`executionWorkspacesSvc.update` takes `Partial<typeof
executionWorkspaces.$inferInsert>`, and `projectWorkspaceId` is a column
on that table; both
`reusableExistingExecutionWorkspace.projectWorkspaceId` and
`resolvedProjectWorkspaceId` are `string | null`, matching the helper's
`string | null | undefined` params / `string | null` return.
- Live-instance exposure check (embedded Postgres): 354
`execution_workspaces` rows carry a null `project_workspace_id`; all of
them belong to projects with **no** project workspace, so
`expectedProjectWorkspaceId` currently resolves null and the guard does
not fire today. The fix is durable heal-on-reuse protection for the
moment any such project gains a primary project workspace (or a null row
is reused for an issue that resolves one).
- CI (full pnpm workspace install) runs the authoritative test +
typecheck for this change on this PR.
## Risks
- Low risk; scoped to the execution-workspace restore path, no schema or
API change.
- The helper only ever *adds* a `projectWorkspaceId` where the row had
none; it never overwrites an existing binding, so it cannot mask a real
`project_workspace_mismatch` (that guard still runs after).
- Complementary to (not overlapping with) #10130, which escalates a
terminal `workspace_validation_failed` run to `blocked` from the
recovery side; this PR prevents the guard from firing on reuse in the
first place. Neither depends on the other.
## Model Used
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use /
code execution (repo edit, embedded-Postgres exposure query, unit-logic
verification).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched open PRs touching heartbeat / execution-workspace /
reuse; only #10130 is related, and it is complementary)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (no
user-facing docs affected)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (in progress)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents governed access to external tools through catalog
connectors.
> - The connector playbook is the repeatable template for adding such
connectors.
> - Notion just shipped as the first MCP-direct connector with RFC 7591
dynamic client registration (#11009).
> - The playbook had no guidance for MCP-direct connections, OAuth
discovery, or DCR.
> - This pull request documents that path and encodes three mandatory
standards into the connector template.
> - It also adds a Notion dry-run appendix recorded from the live probe
and the shipped implementation.
> - The benefit is that the next MCP-direct connector follows a
recorded, verified path.
> [!IMPORTANT]
> **Depends on #11009.** This PR documents the Notion MCP connector that
ships in #11009 (RFC 7591 DCR, discovery-first endpoints,
`redirectConstraints` enforcement, refresh-rotation hardening).
Reviewing this doc against master before #11009 merges will show the
documented behavior as "not implemented" — that is expected. Draft until
#11009 lands, then re-review.
## What Changed
`doc/connections/CONNECTOR-PLAYBOOK.md` only (+364 lines, no code):
- New **"MCP-Direct Connections"** section: RFC 9728/8414 endpoint
discovery chain, an **RFC 7591 dynamic client registration** subsection
(public client, PKCE S256, env-client precedence, persist-and-reuse),
and a **redirect-URI constraints** subsection (`https-or-loopback-http`,
fail-fast wizard error).
- Three mandatory documentation standards encoded into the connector
template itself: a service-involvement statement (DCR providers need
neither Paperclip ID nor Paperclip Connect — instance-local per the
PAP-14828 spec §10 item 8.4; cloud and self-hosted use the same path), a
required **Connection Flow** section (sequence diagram + exact
authorize/token/registration/callback endpoints), and a required
**Administrator Setup** section.
- A **Notion dry-run appendix** mirroring the Linear appendix, recorded
from the live PAP-16649 probe: verified request sequence, redirect-URI
probe results table, sequence diagram (mermaid), shipped manifest
sketch, representative tool risk classes, admin setup (nothing to
register), governance defaults, and validation hooks.
## Verification
- Docs-only change; no code paths affected. `git diff --stat` shows
exactly one file.
- Every endpoint, error code, and constraint in the appendix was checked
against the shipped implementation on the #11009 head:
`server/src/services/tool-access.ts` (`assertOAuthRedirectConstraints`,
DCR registration metadata, refresh serialization),
`server/src/routes/tool-access.ts` (`POST
/api/tools/oauth/:connectionId/start`, `GET /api/tools/oauth/callback`),
and `packages/shared/src/app-definitions/notion.json`
(`redirectConstraints: "https-or-loopback-http"`).
- The request log mirrors the live probe record from PAP-16649
(2026-08-06/07), not vendor docs alone.
- Mermaid source renders cleanly (rendered PNG attached to PAP-16653).
## Risks
- Low: documentation only. Main risk is doc/implementation drift if
#11009 changes before merging — mitigated by the dependency note above
and re-review after #11009 lands.
- The template changes add mandatory sections for future connector
proposals; existing proposals are not retroactively invalidated.
## Model Used
Claude Fable 5 (claude-fable-5).
## Linked Issues or Issue Description
PAP-16653 (parent PAP-16637 P5). Companion catalog package landed in
paperclip-content (`integrations/catalog/platforms/notion/areas/mcp/`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task list and the chat views show a Live badge and a Working
shimmer for an issue that has an active run.
> - A finished task kept the Live badge and the Working shimmer after
the run ended and the sandbox stopped.
> - The user interface reads run liveness from the
`heartbeat_runs.status` row. The run finalizer writes the terminal
status in a step that is separate from the agent `status=done` update.
When the sandbox or the run process stops between the two steps,
`heartbeat_runs.status` stays `running` forever.
> - A run row that stays `running` makes a finished task look
perpetually Live, and the user interface has no guard for an issue that
already reached a terminal status.
> - This pull request closes the invariant "environment lease released
implies the run is terminal" on the server, and adds a user interface
guard that suppresses live state for a terminal issue.
> - The benefit is that a finished task stops showing Live and Working,
both at the source (the run row) and at the surface (the badge and the
shimmer).
## Linked Issues or Issue Description
**Bug description**
- A completed task kept the Live badge and the Working shimmer after its
run ended and the sandbox was torn down.
**Steps to reproduce**
- Run an agent task to completion. Let the sandbox tear down while the
run finalizer is between the `status=done` update and the terminal
run-status write.
- Open the task list or the chat view for the finished task.
**Expected behavior**
- A finished task shows no Live badge and no Working shimmer.
**Actual behavior (before this change)**
- The finished task showed the Live badge and the Working shimmer
because its `heartbeat_runs.status` row stayed `running`.
This pull request supersedes the two separate pull requests #10954
(frontend) and #10955 (backend). It carries all of their changes for the
same race.
## What Changed
Server:
- Run teardown terminalizes a still-running or still-queued run before
it releases the environment lease. It writes `succeeded` when the issue
already reached `done`, `cancelled` when the issue is `cancelled`, and
`interrupted` otherwise. It never overwrites a status that another path
already made terminal.
- The recovery stale-lock sweep terminalizes an orphaned running run to
`interrupted` after it confirms the process and the sandbox are both
gone. It requires recorded process metadata, so it never terminalizes a
live run, a queued run, or a scheduled retry.
- Each terminal transition writes a run event.
- The stale-lock sweep continues and clears the lock when the audit
write fails. It logs the failure loudly.
- New server tests cover both invariants.
User interface:
- A shared guard suppresses the Live badge and the Working shimmer when
the issue status is terminal.
- The guard keeps non-terminal `queued` and `running` issues live.
- The guard prefers the newest issue live-status snapshot.
- New user interface tests cover the guard and the snapshot preference.
## Verification
Server:
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && tsc
--noEmit` in `server/` — 0 errors.
- `pnpm --filter server test
heartbeat-run-lease-release-terminalization.test.ts
recovery-stale-issue-lock-sweep.test.ts` — 12 tests pass.
User interface:
- `pnpm --filter @paperclipai/ui typecheck` — 0 errors.
- `pnpm exec vitest run ui/src/lib/liveIssueIds.test.ts
ui/src/lib/issue-chat-messages.test.ts` — 40 tests pass.
## Risks
- Low risk. The server change only forces a still-live run row to a
terminal status when the lease releases or when the recovery sweep
confirms the process is dead. It never overwrites an existing terminal
status, and it guards the recovery path with process metadata to avoid
terminalizing a live run.
- The user interface change is additive. The guard only suppresses live
state for a terminal issue and keeps queued and running issues live.
- No database migration. No change to any external endpoint.
## Model Used
- Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use, and
code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.
Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - First-run onboarding is the subsystem that turns a brand-new install
into a working company: it creates the company, its goal, a lead agent,
and that agent's first task
> - The existing `OnboardingWizard` carried all of that wiring
correctly, but its UI had drifted from the current design direction, and
a separate design prototype (`paperclip-onboard`) existed as a
standalone visual mock with no backend
> - Porting the prototype's *logic* would have thrown away working,
well-tested backend orchestration; leaving the two apart meant the
design never shipped
> - Separately, cloud and local (self-hosted) installs need meaningfully
different first runs — local has no sign-in and must let the user pick a
locally-installed CLI adapter — so a single linear wizard could not
serve both
> - This pull request rebuilds the presentational layer from the
prototype on top of the existing backend orchestration, and splits it
into two thin flow containers over a shared core
> - The benefit is that the shipped onboarding matches the intended
design, cloud and local can diverge without duplicating logic, and each
can later ship to a different app version while sharing one set of step
components
## Linked Issues or Issue Description
No existing issue — describing inline (feature request).
**What problem does this solve?**
Onboarding is the first thing a new user sees, and the shipped wizard
had drifted from the current design. In parallel, cloud and local
installs need different first-run paths: local has no hosted sign-in,
and its agent runs on a CLI adapter installed on the user's machine,
which the cloud path never has to ask about. There was no way to express
that difference without either forking the whole wizard or bolting
conditionals onto a single linear flow.
**Proposed solution**
Extract the onboarding step views and shell into a shared core, then
compose two thin flow containers (cloud and local) over it. Keep all
backend orchestration in the existing `useOnboardingFlow` hook so no
working logic is rewritten.
**Alternatives considered**
- *Single flow with a `variant` prop* — most DRY, but the two flows are
intended to ship on different app versions, and a shared file would have
to be split later anyway.
- *Two fully independent copies* — simplest per-flow, but every shared
refinement (spacing, motion, copy) would have to be made twice and would
drift.
## What Changed
- **Shared core** under `ui/src/components/onboarding/`:
`OnboardingScaffold` owns the full-screen shell and the single
`AnimatePresence` step crossfade, so both flows transition identically;
step views (Start / Company / Agent / Task), `FooterNav`, `AgentPreview`
and the motion constants are extracted for reuse.
- **`CloudOnboardingFlow`** — `start → company → agent → task`; mounted
in the real app via `OnboardingWizardVariant`. Behaviour matches the
retired wizard, including `previewMock` and the existing-company ("add
an agent") entry point.
- **`LocalOnboardingFlow`** — skips sign-in and adds an optional email
ask (with a privacy assurance), a local model/adapter step that hires
with `requireEnvProbe: true`, and a "star us on GitHub" interstitial
before completing. **Harness-only for now** — the real app still mounts
the cloud flow.
- **Deleted `OnboardingWizard.tsx`** (1,786 lines); updated its
Storybook stories and the `OnboardingWizardVariant` test to the new
components.
- **Orbiting 3D paperclip backdrop** behind the auth and welcome screens
(`three`), code-split so it only downloads on those screens; honours
`prefers-reduced-motion` and disposes its GL context on unmount.
- **`motion`** added for step transitions and the agent-capsule
choreography.
- Visual values routed through design tokens per `DESIGN.md`; `Stepper`
generalized to take a step total (backward compatible); `/design-guide`
page and the component index updated.
- **Standalone preview harness** (`ui/onboarding-preview.html`) with
`?flow=` and `?step=` for backend-free review, wired as a second Vite
rollup input.
- **Adapter env probe bound to the adapter it ran against.**
`hireLeadAgent` reused `adapterEnvResult` for any adapter, so when a
hire failed and the user picked a *different* local adapter and retried,
the previous adapter's verdict satisfied the `requireEnvProbe` guard
while the hire posted the new adapter's config — hiring it unprobed. The
cache is now keyed on the adapter type plus the exact config posted to
the test endpoint, the config is built once and shared by probe and
hire, a failed probe clears the cache, and `clearAdapterEnvResult()`
(called on adapter change) stops the step displaying a stale verdict.
Cloud is unaffected — it hires with `requireEnvProbe: false`. Reported
by Greptile.
- **E2E specs re-pointed at the new flow.** Four specs still drove the
deleted wizard (`onboarding`, `conference-room-typing-intro`,
`planning-mode-visual-verification`, `nux-phase4-screenshots`) and
failed with `element(s) not found` on `"Name your company"` /
`input[placeholder="Acme Corp"]`. Rather than repeat the new drive
sequence four times, `tests/e2e/onboarding-flow.ts` adds one driver per
step (`startCloudOnboarding`, `completeCompanyStep`,
`completeAgentStep`, `completeTaskStep`, `completeCloudOnboarding`) and
the specs import it, so the next flow change touches a single file. Two
now-dead `**/test-environment` route stubs went with it — the cloud flow
hires with `requireEnvProbe: false`, so that probe never fires.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `npx vitest run` over the onboarding suites
(`OnboardingWizardVariant`, `AgentCapsule`, `onboarding-launch`,
`onboarding-goal`, `onboarding-route`, `onboarding-adapter-config`) — 33
tests pass.
- `pnpm --filter @paperclipai/ui build` — succeeds; the three.js chunk
splits out separately (522 kB raw / 133 kB gzip) rather than entering
the main bundle.
- Both flows driven end-to-end in the preview harness in `previewMock`
(no database writes), plus the cloud flow rendered in the real
authenticated app at `/onboarding` to confirm the mount swap.
- The four re-pointed e2e specs pass locally against the new flow.
- New `ui/src/hooks/useOnboardingFlow.test.tsx` — 4 cases pinning the
adapter-probe cache (switch-adapter retry, cold path, explicit clear,
and the cloud flow's `requireEnvProbe: false`). Verified non-vacuous:
the switch-adapter case fails against the pre-fix code.
- Rebased onto current `master`; `pnpm-lock.yaml` is deliberately
**not** committed — `.github/workflows/pr.yml` regenerates it when a
manifest changes and shares it with downstream jobs as the `pr-lockfile`
artifact.
## Risks
- **Deleting `OnboardingWizard.tsx` is the one change that alters
existing app behaviour.** The cloud flow is intended to be
behaviour-equivalent, and its entry points are covered by the updated
`OnboardingWizardVariant` test, but this is the area to review most
closely.
- **Conflict risk with open PRs that touch the old wizard**: #9900,
#9501, #8982 and #6636 all modify
`ui/src/components/OnboardingWizard.tsx`, which this PR removes.
Whichever lands second will need its change re-applied to the new step
components. Flagging so ordering can be decided deliberately.
- **New dependencies**: `motion` and `three` (+ `@types/three`). `three`
is large, so it is lazily imported and code-split — it does not affect
the main bundle. Both are MIT.
- The **local flow is not reachable in the app** yet (harness/canary
only), so it carries no runtime risk today; wiring it up is a follow-up.
- The auth screens remain **presentational only** — they are not wired
to real auth, unchanged from before this PR.
- **Pre-existing, not introduced here:** `OnboardingWizardVariant`
renders outside `<Routes>` in `App.tsx`, so its `useParams()` never
resolves `:companyPrefix` and `/{prefix}/onboarding` opens the welcome
screen instead of jumping to the agent step. `master` has the identical
structure, so this PR faithfully ports existing behaviour; the working
"add an agent" entry is the launcher card behind the overlay, which is
what the screenshot spec drives. Worth a separate fix.
## Model Used
Claude Opus 5 (`claude-opus-5`) via Claude Code, with extended thinking
and tool use (repo search/edit, local test + build execution, and
browser-driven visual verification of the rendered flows). Portions of
the session also ran on `claude-opus-4-8` and `claude-fable-5`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The agent detail page has a Dashboard tab. The Dashboard tab shows a
"Live Run" section for the agent's current heartbeat.
> - The "Live Run" section has two clickable pieces: the section heading
and the running row. Both pieces linked to the same run detail page.
> - Two controls that go to the same place waste a navigation affordance
and hide the task the agent runs.
> - This pull request splits the two destinations. The heading goes to
the run. The running row goes to the task.
> - The benefit is that a user reaches the run internals from the label
and the work item from the row, in one click each.
## Linked Issues or Issue Description
No public GitHub issue exists for this change. The description follows
the enhancement issue template.
**What existing behavior does this improve?**
The "Live Run" section on the agent detail page, Dashboard tab (the
`LatestRunCard` component in `ui/src/pages/AgentDetail.tsx`).
**Current behavior**
The "Live Run" heading and the running row both link to the run detail
page (`/agents/:agentId/runs/:runId`). The row shows the run code and an
invocation-source chip. There is a separate "View details →" link that
also goes to the run detail page. A user cannot reach the task the run
works on from this section.
**Proposed behavior**
The heading becomes a link to the run detail page and appends the short
run code, shown as `Live Run · <run code>`. The redundant "View details
→" link is removed. The running row links to the task detail page when
the run's context snapshot resolves to a known issue, and the row then
shows the task status glyph, the task slug, and the task title. A pure
timer heartbeat with no resolvable task keeps the previous behavior: run
code plus source chip, linking to the run detail page.
**Reason and benefit**
The heading and the row now go to distinct, intuitive destinations. A
user reaches the run internals from the label and the work item from the
row, each in a single click. The fallback keeps heartbeats with no task
readable and avoids a blank or broken row.
## What Changed
- Made the "Live Run" / "Latest Run" heading a `Link` to the run detail
page and appended the short run code (`run.id.slice(0, 8)`) in a mono
span, formatted `Live Run · <run code>`. Kept the pulsing live dot.
- Removed the redundant "View details →" link.
- Changed the running row `Link` target to the task detail page
(`/issues/:identifier`) when a task resolves, falling back to the run
detail page otherwise.
- Resolved the task from the run context snapshot
(`contextSnapshot.issueId`, falling back to `contextSnapshot.taskId`)
against a `Map` of the agent's assigned issues threaded in from
`AgentOverview`.
- When a task resolves, replaced the run code and source chip in the row
with the task status glyph (`StatusGlyph`), the task slug, and the task
title. Kept the running spinner, the run status badge, and the
timestamp.
## Verification
- `pnpm check:token-gates` → 3/3 gates clean.
- `pnpm --filter ui typecheck` → passes.
- `pnpm --filter ui exec vitest run
src/pages/AgentDetail.progress.test.ts
src/pages/AgentDetail.instructions.test.tsx` → 10/10 pass.
- Manual (needs a reviewer with a browser): open an agent detail page →
Dashboard tab.
- For a live issue-execution run: the heading reads `Live Run · <run
code>` and opens the run detail page; the row shows the task status
icon, slug, and title and opens the task detail page.
- For a pure timer heartbeat with no task: the row falls back to run
code + source chip and opens the run detail page. No blank row.
- Confirm both states in light and dark mode.
## Risks
Low risk. The change is presentational and scoped to one component. The
task lookup is defensive: it reads the context snapshot with a fallback
key and only renders the task row when the issue is present in the
already-loaded assigned-issue set, so an unknown or missing issue
degrades to the previous run-detail behavior rather than breaking.
## Model Used
- Provider: Anthropic (Claude).
- Model: claude-opus-4-8 (Opus 4.8).
- Context window: 200K.
- Reasoning mode: extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox provider plugins let agents run commands in remote
environments.
> - The Daytona provider polls the exit code while it waits for command
logs.
> - Polling delays log delivery and does not support long-lived streamed
commands.
> - This pull request adds an opt-in Daytona log stream with one
reconnect and a poll fallback.
> - The benefit is faster log delivery while the existing default path
stays unchanged.
## Linked Issues or Issue Description
Refs: #10941
**Subsystem affected**
packages/plugins — sandbox provider plugins.
**Problem or motivation**
The Daytona provider polls the command exit code every 50 milliseconds
while it waits for logs. This delays output and does not support a
long-lived streamed command.
**Proposed solution**
Add the `useLogStream` provider option. Stream stdout and stderr from
the Daytona callback log form, read the exit code after the stream ends,
retry the read with bounded backoff, and fall back to the existing poll
path after a disconnect.
**Alternatives considered**
Keep polling for all commands. This keeps the current behavior but does
not provide timely logs or a path for long-lived commands.
**Roadmap alignment**
This supports the roadmap item for cloud and sandbox agents, including
Daytona.
## What Changed
- Add the opt-in `useLogStream` option with a default of `false`.
- Stream stdout and stderr from the Daytona callback log form.
- Drop replayed log prefixes by delivered byte offset after reconnect.
- Retry once after disconnect, then use the existing poll path.
- Read the exit code once after a successful stream and retry when the
code is not ready.
- Add tests for ordered output, exit-code reads, disconnect fallback,
and reconnect replay handling.
## Verification
- Run `node_modules/.bin/vitest run --config
packages/plugins/sandbox-providers/daytona/vitest.config.ts
packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`.
- Confirm that the full Daytona plugin test project passes with 127
tests.
- Confirm that the changed files type-check against Daytona SDK 0.203.0
types.
- Review the PR against parent PR #10941 before it reaches `master`.
## Risks
- The stream path changes behavior only when `useLogStream` is `true`.
- A stream failure can add one reconnect attempt before the existing
poll fallback.
- The stream path has no command deadline because it supports long-lived
commands.
- No endpoint, stored data, telemetry shape, authentication rule, or
result shape changes.
## Model Used
OpenAI Codex, GPT-5, tool use and code execution. The context window and
reasoning mode are not exposed by this run.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The CLI and server both update Paperclip values in `.env` files
> - The server preserved operator content, but the CLI rebuilt the
complete file
> - A CLI rerun could remove comments, custom values, ordering, and
newline style
> - Both paths need one editor with one value encoding and duplicate key
policy
> - The final integration also needs one regression test across the
related setup and sync safety mechanisms
> - This pull request moves the editor to the shared package and adds
cross-cutting rerun-survival coverage
> - The benefit is safe setup and worktree repair reruns that preserve
operator edits
## Linked Issues or Issue Description
**What happened?**
The CLI rebuilt the complete `.env` file when it wrote a managed
Paperclip value. This action removed comments, blank lines, custom keys,
original quoting, and the original newline style.
**Expected behavior**
Paperclip must update only the managed assignments. It must preserve all
unrelated bytes. It must skip the file replacement when all managed
values are current.
**Steps to reproduce**
1. Add comments, custom keys, quoted values, and CRLF newlines to the
Paperclip `.env` file.
2. Run a CLI path that calls the agent JWT secret setup.
3. Observe that the old writer replaces the complete file.
**Paperclip version or commit**
The problem exists on `master` before this pull request.
Related public context: Refs #437.
## What Changed
- Add one shared line-preserving `.env` editor for the CLI and server.
- Define minimal and JSON value encodings in the shared helper.
- Update every stale duplicate of a managed key and preserve current
duplicate encodings.
- Preserve comments, ordering, blank lines, unknown keys, export
prefixes, trailing comments, and newline style.
- Write changed files through a same-directory temporary file and atomic
rename.
- Limit CLI updates to non-empty `PAPERCLIP_*` entries.
- Skip the write when all managed values are current.
- Add shared, CLI, and server regression coverage.
- Refresh the branch after the related config, sandbox, and skill safety
changes landed.
- Add a cross-cutting integration test for config, env-file,
managed-sandbox, and managed-instructions rerun survival.
## Verification
- `pnpm exec vitest run packages/shared/src/env-file.test.ts
packages/shared/src/config-schema.test.ts
cli/src/__tests__/agent-jwt-env.test.ts
cli/src/__tests__/config-store.test.ts
server/src/__tests__/config-file.test.ts
server/src/__tests__/worktree-config.test.ts` passes 39 tests.
- `pnpm exec vitest run
server/src/__tests__/rerun-survival.integration.test.ts` passes 4 tests.
- `pnpm -r typecheck` passes on the previous head. GitHub CI reruns it
on the refreshed head.
- The previous head passed the complete general, serialized, workspace,
and E2E matrix. GitHub CI reruns that matrix on the refreshed head.
- `pnpm build` passes on the previous head. GitHub CI reruns it on the
refreshed head.
## Risks
- Low risk. The production change only affects managed `.env`
assignments.
- Existing managed assignments can keep their original quoting when
their decoded values are current.
- Changed CLI values keep the prior minimal encoding policy. Changed
server values keep the prior JSON encoding policy.
- Duplicate managed assignments now follow one explicit rule: Paperclip
updates each stale occurrence.
- The master refresh had one import-block conflict. The resolution keeps
both the config merge imports and the env-file imports.
- The added integration file is test-only. It has no database, API, or
UI contract effect.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex from the GPT-5 family produced this change with reasoning,
tool use, and code execution. The runtime did not expose the exact model
ID or context window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents can select company skills and synchronize them to adapter
runtimes
> - The skill sync API replaced the complete selection without an
explicit destructive choice
> - Company package import also replaced conflicting skills by default
> - These defaults could remove operator edits during setup and import
reruns
> - This pull request adds explicit assignment merge modes and safe
package conflict handling
> - The benefit is that reruns preserve operator work unless the caller
explicitly requests replacement
## Linked Issues or Issue Description
**What existing behavior does this improve?**
This change improves agent skill synchronization and company package
import.
**Subsystem affected**
This is a cross-cutting change across the shared contracts, server, CLI,
and UI.
**Current behavior**
Agent skill synchronization replaces the full desired skill set from a
modeless request. Package import replaces a conflicting skill when the
caller does not select a conflict mode.
**Proposed behavior**
Agent skill synchronization requires `add`, `remove`, or `replace`.
Package import skips conflicts by default. Each imported skill reports
whether it was created, renamed, replaced, or skipped.
**Reason and benefit**
Setup and import reruns must preserve operator edits by default.
Explicit destructive modes make data loss less likely and make each
outcome inspectable.
**Breaking changes**
Callers of the agent skill sync API must now send `mode`. Callers that
need the former behavior must send `replace`. Package import now uses
`skip` when `onConflict` is absent.
## What Changed
- Added required `add`, `remove`, and `replace` modes to the shared
agent skill sync contract.
- Added actionable `422` validation for missing or invalid modes.
- Updated first-party UI and CLI callers with explicit modes.
- Changed package skill conflict handling to use `skip` by default.
- Kept plugin-owned and built-in stock skill imports on explicit
`replace`.
- Added created, renamed, replaced, and skipped results to company
imports.
- Added regression coverage for merge modes and package conflict
outcomes.
## Verification
- `pnpm check:token-gates`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm test:run:serialized` (128 suites passed)
- `pnpm --filter @paperclipai/skills-catalog test` (20 tests passed)
- Focused agent skill route, company skill service, portability, CLI,
and UI tests passed.
- GitHub CI passed build, typecheck, canary, all general and serialized
test shards, all browser shards, policy, security, and final
verification on commit `2cfbb3e4c5`.
- Greptile reviewed the latest commit at 5/5 with zero unresolved
threads.
## Risks
- This change intentionally rejects modeless agent skill sync requests.
- The safe package default can leave an existing skill unchanged where
the old default overwrote it.
- All first-party callers now select a mode. Regression tests cover each
outcome.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI `gpt-5.6-sol` through Codex. The runtime used agentic
reasoning, tool use, code execution, and repository editing. The runtime
did not expose the context window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip creates a managed sandbox environment for each company
during boot
> - Operators can change environment fields after Paperclip creates the
environment
> - The boot reconciler replaced those changes without checking for
drift
> - This pull request adds stock hashes and transactional drift
reconciliation
> - The benefit is that Paperclip can update untouched stock fields
without losing operator work
## Linked Issues or Issue Description
**What happened?**
The managed sandbox boot reconciler rewrote the stock description,
configuration, metadata, and status on every start. It did not detect
operator changes first. A restart could therefore remove an operator's
changes.
**Expected behavior**
Paperclip must preserve operator changes by default. It must update an
untouched stock environment when Paperclip ships new stock values. It
must perform each row update and stock-hash update atomically.
**Steps to reproduce**
1. Start Paperclip and let it create the managed sandbox environment.
2. Change one Paperclip-owned stock field on that environment.
3. Restart Paperclip.
4. Observe that the previous reconciler replaced the change.
**Paperclip version or commit**
This bug reproduces on `master` before this change.
**Deployment mode**
Local development and self-hosted server boot are affected.
## What Changed
- Add a shared deterministic stock-hash and drift classifier for
built-in resources.
- Track the managed sandbox stock hash with the company-scoped built-in
resource binding.
- Reconcile the environment and its stock metadata in one transaction
with a row lock.
- Preserve operator-modified and unmanaged rows and report their skipped
update state.
- Use archive ownership tokens so provider recovery reactivates only
Paperclip-archived rows and preserves later operator archive decisions.
- Keep operator-owned environment variables and unrelated metadata out
of the stock fingerprint.
- Add activity records for managed environment creation, updates,
skipped drift, tracking initialization, and archive changes.
- Add regression tests for current stock, available stock updates,
operator drift, unmanaged rows, archive and reactivation, user-owned
fields, and concurrent reconciliation.
## Verification
- `pnpm -r typecheck`
- `pnpm build`
- `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run --
--mode serialized` (128 suites passed)
- Repository general server, UI, CLI, shared, skills catalog, database,
adapter, plugin SDK, and plugin creator projects passed. Two
embedded-Postgres tests exceeded the host's five-second default under
the aggregate run and passed in the complete database project with
`--testTimeout=20000`. One timing-sensitive sandbox stream test passed
on its focused retry.
- Focused managed-environment unit and integration coverage passed: 49
tests across the drift classifier, boot report, and environment service
suites.
## Risks
- The main risk is an incorrect ownership boundary in the stock
fingerprint. The fingerprint includes only Paperclip-owned stock fields.
Tests confirm that environment variables and unrelated metadata survive
reconciliation.
- Concurrent reconciliation could otherwise overwrite a late operator
edit. The implementation locks the environment row and updates the row
and hash binding in one transaction. A concurrency test covers this
path.
- There is no schema migration. Existing managed rows initialize
tracking without replacing their current values.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI GPT-5 Codex. The exact serving snapshot and context-window size
were not exposed. The model used reasoning, repository tools, code
execution, and test execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The CLI and server share a JSON configuration contract for local
installations and worktrees.
> - Existing config writes removed extension keys because Zod stripped
unknown object properties.
> - Invalid config files could also be replaced with defaults before an
operator preserved the original bytes.
> - Configuration updates must preserve operator edits and must not
rewrite files when the effective value is unchanged.
> - This pull request adds extension-preserving merges, guarded
invalid-config repair, atomic writes, and focused regression tests.
> - The benefit is safe setup and configuration reruns without data loss
or unnecessary mtime changes.
## Linked Issues or Issue Description
**What happened?**
Known-field updates through the CLI or server removed unknown top-level
and nested config keys. Non-interactive configure and onboard paths
could replace a present but invalid config with defaults.
**Expected behavior**
Writers preserve extension keys, skip semantic no-op writes, and require
explicit interactive confirmation before an invalid config is replaced.
Repair preserves an exact collision-safe backup first.
**Steps to reproduce**
1. Add an unknown top-level key and an unknown nested provider key to
`config.json`.
2. Update a known field through the CLI or worktree config writer.
3. Observe that the extension keys are removed on the base branch.
4. Write invalid JSON and run configure or onboard without an
interactive terminal.
5. Observe that the original file can be replaced without a durable
invalid-file backup on the base branch.
**Paperclip version or commit**
`master` at the pull request base commit.
## What Changed
- Accept unknown properties at each extensible config object boundary
while keeping every known field validated.
- Merge known-field updates into the parsed source config and preserve
only unknown extension data.
- Warn about near-match key names without deleting or changing them.
- Skip writes when the effective config is unchanged, which keeps file
mtimes stable.
- Write config changes through a temporary file, file sync, rename, and
directory sync.
- Distinguish a missing config from an invalid config in configure and
onboard.
- Back up invalid bytes as `config.json.invalid-N` and verify the source
still matches that backup before repair.
- Require interactive repair confirmation and reject non-interactive
replacement with an actionable message.
- Document the config preservation and repair behavior.
## Verification
- `pnpm exec vitest run packages/shared/src/config-schema.test.ts
cli/src/__tests__/config-store.test.ts
cli/src/__tests__/configure-repair.test.ts
cli/src/__tests__/configure.test.ts cli/src/__tests__/onboard.test.ts
server/src/__tests__/config-file.test.ts
server/src/__tests__/worktree-config.test.ts`
- `pnpm -r typecheck`
- `AWS_ACCESS_KEY_ID= AWS_SECRET_ACCESS_KEY= VITEST_MAX_WORKERS=1 pnpm
test:run`
- `pnpm build`
- Confirm all pull request checks are green on the latest commit.
- Confirm Greptile reports 5/5 with no unresolved comments.
## Risks
- Passthrough keeps misspelled keys. Near-match warnings make this
visible without destructive cleanup.
- Merge behavior must distinguish unknown extension keys from optional
known keys. Schema-aware regression tests cover preservation and
known-key deletion.
- Repair must not overwrite bytes that changed after backup. The writer
compares the current source with the selected backup before atomic
replacement.
- The change does not alter database schema, company scoping, or
activity logging.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex, GPT-5 model family. The exact deployment model ID and
context window are not exposed. Agentic reasoning, tool use, and code
execution were enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip uses adapter and sandbox code to start agents and run
sandbox work
> - The current sandbox spans use mixed names and do not group related
run-time work
> - Mixed names make traces harder to read and compare across providers
> - This pull request renames provider spans, adds run-time wrapper
spans, and keeps the host allowlist closed
> - The benefit is clearer traces with the same sandbox behavior and
trust boundary
## Linked Issues or Issue Description
**What existing behavior does this improve?**
This improves OpenTelemetry span names and grouping for sandbox startup,
execution, callback relay, and agent session work.
**Subsystem affected**
Cross-cutting (multiple of the above): adapter utilities, sandbox
providers, shared telemetry documentation, and server instrumentation.
**Current behavior**
Sandbox provider spans use mixed names. Related run-time operations
expose inner `sandbox.exec` spans without a named wrapper span. The host
mapper uses a closed allowlist for provider span names.
**Proposed behavior**
Use descriptive provider-scoped span names. Add wrapper spans for agent
session input, agent session output polling, and callback relay. Keep
the host mapper allowlist closed and map unknown names to `other`.
**Reason and benefit**
Clear names make traces easier to read and reduce ambiguity during
sandbox operation analysis. Wrapper spans show the full operation while
preserving the inner execution spans.
**Breaking changes**
None. This change updates telemetry span names and grouping only. It
does not change sandbox behavior, endpoint behavior, or the host trust
boundary.
**Additional context**
Related prior work:
[#10758](https://github.com/paperclipai/paperclip/pull/10758).
## What Changed
- Rename Daytona provider sync and session spans with descriptive
provider-scoped names.
- Add three run-time wrapper spans for agent session input, output
polling, and callback relay.
- Add a shared span runner that preserves no-op behavior without a real
tracer.
- Keep the host mapper allowlist closed and map unknown names to
`other`.
- Update telemetry documentation and span-name tests.
## Verification
- Focused adapter-utils span tests pass for startup timing, callback
relay, and sandbox execution.
- Focused Daytona plugin span tests pass for renamed leaf spans and
session open or close spans.
- Focused server tests pass for host mapping and instrumentation.
- The stacked diff contains one commit on top of
`feat/daytona-persistent-session-model`.
## Risks
- Span names change for existing telemetry consumers.
- The wrapper spans add trace structure but do not change sandbox
execution.
- The host mapper keeps the existing closed allowlist and `other`
bucket.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI GPT-5 (Codex agent); exact deployment revision and context window
are not exposed in this run; tool use and code execution enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents governed access to external tools.
> - The Apps gallery lists Notion, but the server required manually
configured OAuth credentials.
> - Notion's hosted MCP server supports OAuth discovery and dynamic
client registration.
> - Notion also requires HTTPS or a loopback HTTP redirect URI.
> - This pull request adds a direct Notion MCP OAuth path with PKCE and
reusable dynamic clients.
> - It also adds the current Apps UI states for connect and
reauthorization.
> - The benefit is a secure Notion connection with no manual client
credential setup.
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The Apps gallery, Apps connect route, OAuth token lifecycle, and managed
MCP gateway.
**Subsystem affected**
`server/`, `packages/shared/`, `scripts/`, and `ui/`.
**Current behavior**
The Notion gallery cards are disabled. The server uses the classic
Notion OAuth endpoints and requires operator-supplied client
credentials. It does not register an OAuth client from provider
metadata. Concurrent refreshes can also replay a rotating refresh token.
**Proposed behavior**
Enable the Notion Apps flow. Discover OAuth metadata from
`https://mcp.notion.com/mcp`. Register and reuse a public RFC 7591
client with PKCE. Require HTTPS or loopback HTTP callbacks. Serialize
refreshes, store each rotated refresh token before the new access token
can be used, and show a reconnect state for `invalid_grant`.
**Reason and benefit**
Operators can connect the built-in Notion MCP app without creating or
copying OAuth credentials. Paperclip keeps dynamic clients and rotating
tokens in the company secret store.
**Breaking changes**
None. Explicit environment client credentials still take priority.
Existing Slack and Linear OAuth endpoint hints remain unchanged. Other
OAuth apps remain disabled unless they are allowlisted.
**Additional context**
PR #10910 is a related, broader Connections v3 wizard replacement. This
PR is the focused current Apps flow. The MCP Tool Gateway and Connected
Apps items in `ROADMAP.md` cover this planned capability.
## What Changed
- Classify all 20 reviewed Notion MCP tools with provider-scoped read
and write defaults.
- Require approval for selected Notion mutations, including move,
duplicate, and convert actions that generic verb matching missed.
- Preserve company-scoped connection and catalog resolution for Notion
profiles and policies.
- Add RFC 7591 dynamic client registration with
`token_endpoint_auth_method=none` and mandatory PKCE.
- Store the dynamic client ID on the connection and store any returned
client secret in the company secret store.
- Reuse the registered client for later connects and keep explicit
environment credentials as the first choice.
- Discover protected-resource and authorization-server metadata from the
Notion MCP endpoint.
- Add `redirectConstraints: "https-or-loopback-http"` to the generated
Notion app definition and shared contract.
- Reject non-loopback plain HTTP callbacks before network access with a
TLS setup error.
- Serialize client registration and token refresh operations within the
server process.
- Store a rotated refresh token before publishing the refreshed access
token.
- Treat `invalid_grant` as terminal and move the connection to a clear
reauthorization state.
- Add focused coverage for registration reuse, callback constraints,
refresh rotation, and terminal grants.
- Enable the Notion Apps route and add connect, redirect, success,
error, and reconnect UI states.
- Keep non-allowlisted OAuth apps blocked and cover the UI policy with
regression tests.
## Verification
- The focused Notion policy integration test passed with embedded
PostgreSQL.
- The focused 20-tool classification test passed.
- The server typecheck passed on the governance head.
- `pnpm -r typecheck` passed on the rebased head.
- `pnpm --filter @paperclipai/server typecheck` passed after the
security follow-up.
- `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts
-t \u0027DCR|refresh tokens|invalid_grant|abandoned lease\u0027` passed
10 focused security tests.
- `pnpm build` passed on the rebased head.
- `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts
-t 'OAuth|oauth'` passed 14 tests.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts`
passed 5 tests.
- The complete server group passed 3,686 tests with 4 skipped.
- The complete UI group passed 3,656 tests.
- The full local runner found one environment-only CLI failure because
this agent runtime injects static AWS credentials into a test that
expects `AWS_PROFILE` only. `env -u AWS_ACCESS_KEY_ID -u
AWS_SECRET_ACCESS_KEY pnpm exec vitest run
cli/src/__tests__/secrets.test.ts` passed all 8 tests.
- The prior UI verification passed 55 focused tests, `pnpm
check:token-gates`, the Storybook build, and review of six 1440 x 1000
screenshots.
- OAuth request sequence: protected-resource metadata `GET
https://mcp.notion.com/.well-known/oauth-protected-resource/mcp`;
authorization metadata `GET
https://mcp.notion.com/.well-known/oauth-authorization-server`; dynamic
registration `POST https://mcp.notion.com/register`; authorization `GET
https://mcp.notion.com/authorize`; token exchange and refresh `POST
https://mcp.notion.com/token`; MCP traffic `POST
https://mcp.notion.com/mcp`.
- The live metadata and registration probe confirmed that Notion accepts
HTTPS and loopback HTTP redirects. It rejects a plain HTTP private
hostname.
- A later QA task owns the full browser consent and managed gateway
tool-list dry run against a configured HTTPS deployment.
## Risks
- Notion can add tools. Unrecognized names use the generic classifier,
and new or changed risky tools stay quarantined after connection
activation.
- A deployment that uses a private non-loopback hostname must configure
HTTPS before it can connect Notion.
- Dynamic registration creates a provider-side client. Paperclip reuses
it because registration does not provide a standard delete operation.
- Refresh coordination uses a database CAS lease across service
instances. An unclean crash leaves an uncertain lease and requires
reconnect instead of risking refresh-token replay.
- The current Apps surface overlaps with PR #10910. Merge order can
require a small conflict resolution if that PR lands first.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex on a GPT-5 runtime. The exact deployment ID and context
window are not exposed. The runtime used reasoning, repository tools,
code execution, and network tools.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task view shows a threaded conversation between the user and the
agent, with agent replies rendered as bubbles that carry a "✓ Worked · N
tools" summary line.
> - The conference-room chat already offers per-message copy and
thumbs-up / thumbs-down feedback, but the redesigned task thread dropped
these controls from the agent bubble footer.
> - Users lose a quick way to copy an agent reply or send feedback on
it, and the redesign silently ignored the feedback-vote props it was
already given.
> - This pull request prepends a copy · thumbs-up · thumbs-down cluster
to the bubble summary line and wires the existing feedback-vote props
through.
> - The benefit is a consistent feedback surface across both chat views,
with no new API.
## Linked Issues or Issue Description
No public GitHub issue exists for this change; the underlying issue is
described inline below following the feature template
(`.github/ISSUE_TEMPLATE/feature_request.yml`).
#### Problem or motivation
The redesigned task thread renders each agent reply with a "✓ Worked · N
tools · <timestamp>" summary line, but it dropped the copy and thumbs-up
/ thumbs-down controls that the conference-room chat still shows. Users
can no longer copy an agent reply or vote feedback from the task thread.
The redesign component already received `feedbackVotes` and `onVote`
props but ignored them.
#### Proposed solution
Prepend a copy · 👍 · 👎 cluster to the summary line, leading the
always-visible timestamp, reusing the shared `IssueChatFeedbackButtons`
so both chat views speak the same feedback language. Anchor the cluster
to the turn's summary row (a sibling of the expandable tool-history
fold) so it stays on the summary line whether the tool history is
collapsed or expanded.
#### Alternatives considered
Placing the cluster inside the expandable fold — rejected because
expanding the tool history then re-centered the actions to the middle of
the tall fold.
#### Roadmap alignment
UI polish to the task thread; no core-roadmap overlap.
## What Changed
- Add `TaskChatBubbleActions`: a copy · thumbs-up · thumbs-down cluster
built on the shared `IssueChatFeedbackButtons`.
- Render the cluster on the agent bubble's "✓ Worked · …" summary line,
leading the timestamp; runless agent replies get the same cluster with
the timestamp trailing. Human and system bubbles are unchanged.
- Add a `leading` slot to `TaskChatTurn` so the actions sit on the
summary row, a sibling of the tool-history fold, and stay anchored when
the fold expands.
- Wire the redesign to the `feedbackVotes` / `onVote` props it already
received.
- Add a demo binding in the `TaskChatLab` dev harness.
## Verification
- `pnpm check:token-gates` — 3/3 gates CLEAN.
- `pnpm typecheck` — clean across all packages.
- `cd ui && pnpm vitest run
src/components/task-chat/TaskChatBubble.test.tsx
src/components/task-chat/TaskChatTurn.test.tsx` — 31/31 pass.
- Manual: open a task thread, confirm the copy / 👍 / 👎 cluster shows on
the agent bubble summary line before the timestamp, copy works, votes
toggle, and the cluster stays on the summary line when the tool history
is expanded.
Visual change: snapshot baselines are intentionally not updated, per the
`doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification
demoted to dormant (Jul 13 2026)".
## Risks
Low risk. UI-only change scoped to the redesigned task-chat bubble
footer. It reuses an existing shared feedback component and existing
vote props; no API, schema, or server change. Human and system bubbles
are untouched.
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
thinking enabled, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the control plane for autonomous AI companies
> - One core subsystem runs agent work inside sandboxes
> - The Daytona provider uses that path to run user commands
> - The current one-shot model does not keep a shell alive across
commands
> - This pull request adds an opt-in persistent session model for
Daytona
> - The benefit is faster command dispatch with the same sandbox
boundaries
## Linked Issues or Issue Description
### Subsystem affected
Cross-cutting. This change touches `packages/adapters`, provider tests,
span names, and sandbox command behavior.
### Problem or motivation
The Daytona provider needs a persistent shell for repeated command
dispatch.
The old advisory wrapper path does not reach that goal.
It also adds cost and removes the session speed gain.
### Proposed solution
Add a `useSessions` driver flag.
Keep it off by default.
Open one Daytona session per lease when the flag is on.
Send each user command into that session.
Read stdout and stderr from the session logs endpoint.
Run each command in a subshell so `exit` does not stop the shell.
Remove the advisory `bwrap` wrapper path and its lease metadata.
Add session setup and teardown spans.
Keep a hard delete on teardown.
### Alternatives considered
Keep the advisory `bwrap` wrapper.
That path does not give a real persistent session.
It also keeps extra command overhead.
Keep a one-shot fallback for user commands.
That would weaken the session model and hide a missing session case.
### Roadmap alignment
This work fits the `Cloud / Sandbox agents` milestone in `ROADMAP.md`.
It also supports the control plane goal of safe remote sandbox
execution.
### Additional context
The handoff verification reported `tsc --noEmit` clean and 119 Daytona
unit tests passing.
The handoff also reported a clean host span allowlist test and five
expected commits on the branch.
The security review gate remains required before merge.
## What Changed
- Added an opt-in persistent session model for the Daytona sandbox
provider.
- Routed user commands through `executeSessionCommand` when sessions are
enabled.
- Removed the advisory `bwrap` command wrapper path and the lease
metadata it used.
- Added session lifecycle spans and span allowlist coverage.
- Documented the leak bound in `DIRECTORY-CONSTRAINT-FINDINGS.md`.
## Verification
- `tsc --noEmit` clean for the Daytona plugin, per handoff verification.
- Daytona unit suite passes, with 119 tests, per handoff verification.
- Host span allowlist test passes, per handoff verification.
## Risks
- Persistent sessions can leak if teardown fails.
- Session logs must keep stdout and stderr separate.
- The flag stays off by default to limit rollout risk.
## Model Used
OpenAI GPT-5, Codex, tool use enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks and issues carry a `priority` field that renders across many
product surfaces: the detail header, the Triage properties panel, Kanban
and thread cards, the New Task composer, list Sort/Group/Filter menus,
search filters, and the dashboard chart.
> - Product feedback found the priority level adds visual noise and
decision cost without clear value in day-to-day task flow.
> - We want to remove priority from the interface, but keep the data
model, API, validation, and search DSL fully intact so the choice is
reversible with no migration.
> - This pull request hides every priority indicator and control behind
one compile-time flag, `SHOW_TASK_PRIORITY_UI`, set to `false`.
> - The benefit is a calmer, simpler UI now, with a single-boolean
revive path and zero data loss.
## Linked Issues or Issue Description
<!-- No public GitHub issue exists. Describing in-PR per the feature
template. -->
**Subsystem affected**
The web UI (`ui/src`): issue detail, properties panel, Kanban/thread
cards, New Task dialog, issues list Sort/Group/Filter menus, search
filter bar/sheet, dashboard charts, and the design-guide showcase.
**Problem or motivation**
The task/issue priority level appears across many surfaces and adds
visual clutter and decision overhead without pulling its weight in
normal task flow. We want it gone from the interface without discarding
the underlying data or breaking anything that depends on it.
**Proposed solution**
Add a single compile-time UI flag, `SHOW_TASK_PRIORITY_UI` (default
`false`), and gate every priority indicator and control behind it. Leave
the data model, API params, Zod validation (including the `"medium"`
default), and the search filter DSL untouched. Reviving priority is a
one-line flip of the flag back to `true`.
**Alternatives considered**
Deleting the priority code and schema outright. Rejected: it is
irreversible, needs a data migration, and throws away a field the API
and search still support. A gated flag keeps the change reversible and
low risk.
**Roadmap alignment**
UI simplification. This is a presentation-only change; it does not alter
core agent or data behavior.
## What Changed
- Added `ui/src/lib/ui-flags.ts` exporting `SHOW_TASK_PRIORITY_UI:
boolean = false` (typed `boolean` so gated branches are not flagged as
dead code).
- Gated the priority row in the Triage properties panel and the editable
priority control in the issue detail header (plus its skeleton seed).
- Gated the per-card priority icon in `KanbanBoard` and in
`IssueThreadInteractionCard`.
- Hid the priority chip and the mobile "more" menu priority section in
the New Task dialog. The submit path still sends the `"medium"` default.
- Removed the Priority options from the issues list Sort and Group-by
menus; the comparator and grouping logic stay dormant.
- Hid the Priority sections in the issue filters popover and in the
search filter bar and sheet. The `priority:` search DSL and filter state
stay functional at the data layer.
- Suppressed active-filter priority pills for consistency.
- Gated the "Tasks by Priority" dashboard chart and the design-guide
priority showcase subsection.
- Left activity-feed "changed priority" history text intact as a
historical record.
- Updated call-site tests to assert priority UI is absent while the flag
is off, added focused hidden-surface tests, and added a test that proves
creating a task still persists `priority: "medium"`.
## Verification
- `pnpm check:token-gates` — all 3 gates clean.
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/ui exec vitest run` on the touched
surfaces (IssueProperties, IssueFiltersPopover, IssuesList,
NewIssueDialog, IssueDetail, PriorityIcon and its interaction test) —
all green under `TZ=UTC`.
- Manual: with the flag off, priority does not appear in the detail
header, Triage panel, New Task composer, Sort/Group/Filter menus, or the
dashboard chart. Creating a task still persists `priority: "medium"`,
and the `priority:` search token still filters at the data layer.
## Risks
Low risk. The change is presentation-only and additive: no data model,
API, validation, or search-DSL changes. The priority code paths remain
compiled and tested; flipping `SHOW_TASK_PRIORITY_UI` to `true` restores
the full UI. Visual snapshot baselines are intentionally not updated per
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
thinking enabled, with tool use / code execution in an agentic coding
harness.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runs execute inside sandbox environments. Sandbox provider
plugins (Daytona, Modal, exe.dev, and others) declare a JSON-Schema
`configSchema`. The Environment configuration form renders from that
schema.
> - The sizing and image fields open empty. Users must guess working
values. The Modal form cannot submit at all until the user types an app
name and an image by hand.
> - The form renderer already pre-fills every field that declares a
JSON-Schema `default`. The manifests do not use this mechanism for
sizing or image fields.
> - This pull request adds optional `default` values to the Daytona,
Modal, and exe.dev manifest schemas.
> - The benefit is a form that opens with known-good values. Users can
create a working environment without provider research.
## Linked Issues or Issue Description
**Current behavior**
The Environment configuration form opens with empty sizing and image
fields for the Daytona, Modal, and exe.dev sandbox providers. Users must
find working values in provider documentation. Modal declares `appName`
and `image` as required with no default, so the form blocks submission
until the user invents both values.
**Proposed behavior**
The provider manifests declare JSON-Schema `default` values. The
existing form renderer pre-fills them:
- Daytona: CPU `4`, memory `4` GiB, disk `10` GiB, image
`daytonaio/sandbox:0.8.0`
- exe.dev: CPU `4`, memory `4GB`, disk `20GB`
- Modal: app name `paperclip`, image `node:22`
Secret-ref fields (API keys, tokens) get no defaults on purpose. The
form persists a raw string in a secret-ref field as a company secret on
save. A placeholder default would become a stored secret with a bogus
value. Each plugin test suite now guards this invariant.
**Reason and benefit**
New users can create a working sandbox environment without guessing. The
defaults stay optional: users can clear or change every value, and the
schema marks no new field as required. The Modal image default `node:22`
satisfies the sandbox runtime contract in `SANDBOX-REQUIREMENTS.md`
(`node`, `sh`, and `tar` on PATH).
**Subsystem affected**
Sandbox provider plugins (`packages/plugins/sandbox-providers/*`):
environment driver `configSchema` manifests.
**Breaking changes**
None. Defaults only seed the create-mode form. Saved environments keep
their stored config. E2B, Novita, Cloudflare, and Kubernetes manifests
do not change: E2B and Novita already default to their base templates,
and the Cloudflare bridge and Kubernetes cluster fields have no sensible
universal value.
## What Changed
- Add `default` values for `cpu`, `memory`, `disk`, and `image` in the
Daytona manifest. Trim the memory description to match.
- Add `default` values for `cpu`, `memory`, and `disk` in the exe.dev
manifest.
- Add `default` values for `appName` and `image` in the Modal manifest.
Extend the image description with the runtime-contract rationale.
- Add manifest tests in all three plugins: defaults match expected
values, defaults satisfy their own schema constraints, and no secret-ref
field declares a default.
- Bump plugin versions: daytona and modal `0.1.0` → `0.1.1`, exe-dev
`0.1.1` → `0.1.2`.
## Verification
- Run `pnpm test` in `packages/plugins/sandbox-providers/daytona`,
`.../modal`, and `.../exe-dev`. The new `* manifest form defaults`
suites pass.
- Run `./node_modules/.bin/tsc --noEmit` in each of the three packages.
Typecheck passes.
- Manual: rebuild the plugins (`pnpm build` in each package), let the
plugin dev-watcher refresh the manifest, then open Environments → New
environment. The Daytona form shows CPU 4, Memory 4, Disk 10, and image
`daytonaio/sandbox:0.8.0`. The Modal form shows `paperclip` and
`node:22`. The API key fields stay empty.
- Verified live on a local instance: the served `configSchema` in the
plugin registry carries the new defaults, and existing environments are
unchanged.
## Risks
- Low risk. The change touches only manifest schema metadata and tests.
No runtime code path changes.
- New environments created with untouched forms now request 4 CPU / 4
GiB / 10 GiB from Daytona instead of provider minimums. This can raise
cost per sandbox for users who previously saved empty fields.
- The Daytona image default pins `daytonaio/sandbox:0.8.0`. The default
needs a manual bump when Daytona ships new sandbox images.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), via the Claude
Code CLI, with extended thinking and tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip can run as a self-hosted app or as a Cloud-managed tenant.
> - These modes need different sign-out sequences because Cloud owns
three sessions.
> - Several visible controls implemented sign-out separately and could
choose different paths.
> - This pull request adds one Cloud-aware sign-out action and moves
every visible control to it.
> - The benefit is one safe sign-out path in Cloud and unchanged local
sign-out in self-hosted deployments.
## Linked Issues or Issue Description
Refs #2073. That older PR adds a separate company-settings sign-out
surface. This change centralizes the existing account, company, and
instance-settings surfaces and preserves the self-hosted behavior
described there.
**What happened?**
Visible sign-out controls used separate implementations. A Cloud-managed
control could call the app-local sign-out endpoint and open the local
auth page. That path did not enter the Cloud-owned logout sequence for
the tenant, Cloud, and identity sessions.
**Expected behavior**
Every visible sign-out control must use one action. Cloud-managed
instances must navigate the top-level window to the same-origin
`/cloud/logout` route without a local sign-out call first. Authenticated
self-hosted instances must keep the local sign-out API and cache
invalidation behavior.
**Steps to reproduce**
1. Open a Cloud-managed tenant.
2. Use the account menu, company menu, or instance-settings sign-out
control.
3. Observe that independently implemented controls can enter different
sign-out paths.
**Paperclip version or commit**
The problem reproduces at `656ecfa585b31938e2685ffab3db22e794474803`.
**Deployment mode**
Cloud-managed tenant built from source. The regression tests also cover
authenticated self-hosted mode.
## What Changed
- Added `useSignOut` as the shared Cloud-aware sign-out action.
- Navigated Cloud-managed sessions to `/cloud/logout` exactly once
without calling local auth first.
- Preserved local API sign-out and cache invalidation for authenticated
self-hosted sessions.
- Migrated the account menu, company menu, and instance general settings
to the shared action.
- Added focused tests for mode selection, menu closure, pending state,
failure state, and settings behavior.
## Verification
- `pnpm exec vitest run ui/src/hooks/useSignOut.test.tsx
ui/src/components/SidebarAccountMenu.test.tsx
ui/src/components/SidebarCompanyMenu.test.tsx
ui/src/pages/InstanceGeneralSettings.test.tsx` — 26 tests passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm check:token-gates` — passed.
- `git diff --check origin/master..HEAD` — passed.
## Risks
- Low risk. The change centralizes existing behavior and adds no schema,
API, telemetry, or style-token changes.
- Cloud mode depends on the existing health decision. Tests pin both
mode branches.
- The change does not alter Fetch Metadata, CSRF, cookie, prefetch, or
return-URL protections owned by the Cloud logout route.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex, GPT-5. The runtime does not expose a more specific model
ID or context-window size. The agent used high-reasoning mode, shell
tools, and API tools.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
<!-- ASD-STE100 -->
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server keeps fleet health with periodic sweeps and shows a
dashboard with run activity
> - The Paperclip instance became slow again after the first round of
recovery-sweep indexes landed
> - Live profiling found four steady-state hot paths that read much more
data than they use
> - This pull request bounds the dashboard recursion, adds the missing
taskKey index, and narrows two wide reads
> - The benefit is a large drop in constant database load and a
responsive server
## Linked Issues or Issue Description
**Describe the bug**
The server becomes slow while agents work. Live query sampling shows
four hot paths:
1. The dashboard run-activity recursive CTE reads every run a company
ever had on each call. One call takes 2.85 seconds. The UI calls it
after almost every fleet event through the dashboard and sidebar-badges
routes.
2. The productivity-review sweep runs each 30 seconds. Its run-scope
filter is `issueId OR taskId OR taskKey` on the run context JSONB. No
index exists for `taskKey`. The planner must detoast every run snapshot
for the agent. One query takes 444 ms and the sweep makes one for each
of ~152 candidate issues.
3. The attention failed-run section selects the full `context_snapshot`
for every run newer than the oldest exhausted run. That fetch moves 29
MB for each feed build.
4. The retention sweep pages the attention feed with a cursor. Each page
makes a full feed rebuild.
**Expected behavior**
Periodic sweeps and dashboard queries read only the data they use, and
use indexes.
**Actual behavior**
The database stays saturated. Users see a slow server.
## What Changed
- `server/src/services/dashboard.ts`: bound both arms of the
`recovered_runs` recursive CTE to the chart window. A retry is always
newer than the run it retries, so the bound cannot change visible chart
data. Live time went from 2,852 ms to 54 ms.
- `packages/db/src/migrations/0210_heartbeat_context_taskkey_index.sql`:
add the `taskKey` expression index that completes the
issueId/taskId/taskKey trio. With all three, the planner uses a
BitmapOr. Live time for the productivity run-scope query went from 444
ms to 1.9 ms.
- `packages/db/src/schema/heartbeat_runs.ts`: mirror the new index in
the Drizzle schema.
- `server/src/services/productivity-review.ts`: select only the seven
run fields the evidence code reads. Before, the query pulled full rows
with `result_json` (up to 43 kB per row, 100 rows per issue).
- `server/src/services/attention.ts`: project `issueId`/`taskId` text
fields instead of the full `context_snapshot` in the failed-run
newer-runs query (29 MB per feed build before).
- `server/src/index.ts`: the retention sweep now builds the attention
feed once per company with `all: true` instead of one full rebuild per
cursor page.
- `packages/db/src/heartbeat-context-snapshot-index-migration.test.ts`:
cover the new index and re-run migration 0210 statements to prove
idempotency.
## Verification
- `pnpm --filter @paperclipai/db typecheck` (includes migration
numbering and safety checks) — pass.
- `npx tsc --noEmit` in `server/` — pass.
- `npx vitest run
packages/db/src/heartbeat-context-snapshot-index-migration.test.ts` —
pass (embedded Postgres, full migration chain, planner assertions,
idempotent re-run of 0209 and 0210).
- `npx vitest run` on attention, dashboard, productivity-review,
decision-retention, issue-blocker-attention, and issue-review-attention
test files — 72/72 pass.
- Live EXPLAIN ANALYZE before/after numbers are in the What Changed
list.
## Risks
- Migration 0210 builds one btree index without CONCURRENTLY inside the
transactional migration runner. The table is not in the large-table
bucket. The 0209 twin built in seconds on a 100k-row live table.
- The CTE bound excludes retry ancestors that are older than the chart
window. Those rows are not visible to the chart query, so chart output
does not change.
- The attention projection changes JSONB scalar handling in one edge
case: a non-string `issueId`/`taskId` value now casts to text instead of
reading as absent. These keys are always strings in practice.
- The retention sweep now holds one full feed in memory per company. The
cursor loop already accumulated all pages into one array, so peak memory
is unchanged.
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic, Mythos-class tier, extended
thinking + tool use) via Paperclip agent runtime.
- [x] I searched existing PRs and issues and this change is not a
duplicate.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The task chat thread renders issue comments, system notices, and run
transcripts
> - The server routes comment payloads through the run-secret redaction
walker before it sends them
> - The walker rebuilds each object with `Object.entries`, and this
collapses `Date` instances to `{}`
> - The chat renderer then calls `.toISOString()` on an invalid date and
throws, and the thread falls back to the error banner
> - This pull request keeps `Date` instances intact in redacted
responses and makes the renderer safe against bad timestamps
> - The benefit is that task threads with system notices render
correctly again
## Linked Issues or Issue Description
**What happened**
Task threads that contain a system notice showed the banner "Chat
renderer hit an internal state error." in place of the conversation.
This occurred on many tasks.
**Expected behavior**
The thread renders all comments and system notices with correct
timestamps.
**Steps to reproduce**
1. Open a task that has at least one system notice comment (for example
a "Workspace ready" notice).
2. `GET /api/issues/{id}/comments` returns `createdAt: {}` for every
comment because the secret-redaction walker collapses `Date` objects.
3. The system-notice row calls `new Date({}).toISOString()`. This throws
`RangeError: Invalid time value` and trips the thread error boundary.
**Version / deployment**
Regression from #9934 (`e43f187ca`). It applies to all deployments that
include that commit.
## What Changed
- `server/src/services/run-secret-redaction.ts`:
`redactRegisteredSecretValues` now returns `Date` instances as-is. Dates
hold no redactable text, and the `Object.entries` rebuild turned them
into `{}`.
- `ui/src/components/IssueChatThread.tsx`: the system-notice row formats
its timestamp with a new `toValidIsoString` helper. A value that does
not parse as a date now degrades to "no timestamp" instead of a render
crash.
- Regression tests at three layers:
- Walker unit tests: `Date` values survive with and without registered
secret values.
- Route test: `GET /issues/:id/comments` serializes `createdAt` /
`updatedAt` as ISO strings.
- Render test: a system notice with a malformed `createdAt` renders
without the error boundary.
## Verification
- `npx vitest run --root server
src/__tests__/run-secret-redaction.test.ts` — 5 passed.
- `npx vitest run --root server
src/__tests__/issue-comment-redaction.test.ts` — 4 passed (embedded
Postgres route test).
- `cd ui && npx vitest run src/components/IssueChatThread.test.tsx
src/components/IssueChatThreadSystemNotice.test.tsx
src/lib/issue-chat-messages.test.ts` — 121 passed.
- Each new test was run against the unfixed code and failed there, which
confirms it guards the regression.
- A local sweep rendered 47 real issue threads through
`IssueChatThread`: 7 tripped the boundary before the fix, 0 after.
## Risks
- Low risk. The server change only preserves `Date` objects that the
walker destroyed before. String redaction behavior does not change, and
the registry-key stripping does not change.
- The UI change only affects the timestamp of system-notice rows and
omits it when the value is invalid.
## Model Used
- Claude Fable 5 (`claude-fable-5`), Anthropic. Agentic coding session
with tool use (file edits, shell, Vitest). No extended-context or
special reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server runs a recovery sweep every 30 seconds. The sweep makes
sure assigned issues do not stall.
> - The sweep reads the latest heartbeat run for each candidate issue.
It filters `heartbeat_runs` on `context_snapshot ->> 'issueId'`. No
index covers this expression.
> - Each lookup scans ~26k rows and detoasts each row's JSONB context
(~645 ms per query, measured with EXPLAIN ANALYZE). With ~460 candidate
issues, one sweep schedules ~500 seconds of database work every 30
seconds.
> - The database saturates. Users see the whole server as slow.
> - This pull request adds expression indexes so these lookups become
index scans.
> - The benefit is that the recovery sweep drops from ~500 seconds of
database work per tick to milliseconds, and the server becomes
responsive again.
## Linked Issues or Issue Description
**What happened?**
The server became slow for all users. `pg_stat_activity` sampling showed
the same two recovery-sweep queries active continuously (15/15 samples).
Each `getLatestIssueRun` call filtered `heartbeat_runs` on the unindexed
expression `context_snapshot ->> 'issueId'`, scanned ~26k rows, and
detoasted a 2.1 GB TOAST region. `heartbeat_runs` accumulated 10.8
billion sequential tuples read. `hasActiveExecutionPath` also scanned
`agent_wakeup_requests` (1.8M rows) on the unindexed expression `payload
->> 'issueId'`.
**Expected behavior**
Per-issue run lookups in the recovery sweep complete in milliseconds.
The sweep finishes well inside its 30-second interval. Background
maintenance does not degrade interactive latency.
**Steps to reproduce**
1. Run a board with several hundred issues in `todo` / `in_progress` /
`in_review` and a large `heartbeat_runs` table with big
`context_snapshot` payloads.
2. Let the heartbeat scheduler run its 30-second recovery sweep.
3. Observe `pg_stat_activity`: the per-issue `heartbeat_runs` lookups
run continuously; `EXPLAIN ANALYZE` shows a filter on `context_snapshot
->> 'issueId'` removing tens of thousands of rows per call.
**Paperclip version or commit**
master (814cb336)
## What Changed
- Add migration `0209_heartbeat_context_snapshot_indexes.sql` with three
expression indexes:
- `heartbeat_runs (company_id, (context_snapshot ->> 'issueId'),
created_at DESC)`
- `heartbeat_runs (company_id, (context_snapshot ->> 'taskId'),
created_at DESC)`
- `agent_wakeup_requests (company_id, (payload ->> 'issueId'))` — with a
`large-create-index-not-concurrently` safety pragma and justification,
following the migration 0206 precedent.
- Mirror the three indexes in the drizzle schema files
(`heartbeat_runs.ts`, `agent_wakeup_requests.ts`).
- Add `heartbeat-context-snapshot-index-migration.test.ts`. The test
boots a fresh embedded Postgres, applies the full migration chain,
asserts the indexes exist, and asserts with `EXPLAIN` that the planner
selects them for the exact hot query shapes.
## Verification
- `pnpm --filter @paperclipai/db exec tsx
src/check-migration-numbering.ts` passes.
- `pnpm --filter @paperclipai/db exec tsx src/check-migration-safety.ts`
passes.
- `npx vitest run
src/heartbeat-context-snapshot-index-migration.test.ts` passes (fresh
embedded Postgres, full chain 0000→0209, planner uses all three
indexes).
- `npx vitest run src/check-migration-safety.test.ts` passes (25/25).
- `tsc --noEmit` clean for `packages/db`.
## Risks
- The index builds run inside the transactional migration (no
`CONCURRENTLY`). The `heartbeat_runs` build must read its 2.1 GB TOAST
once; expect the migration step to add roughly 1–3 minutes to one deploy
while writes to the two tables wait. This is a one-time cost at startup,
before the server accepts traffic.
- Three new indexes add small write overhead to two hot-write tables.
The read savings are several orders of magnitude larger.
- No query or API behavior changes; the planner simply gains a better
access path.
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, Postgres EXPLAIN/ANALYZE against the live instance for
measurement, vitest for verification).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task detail view shows the conversation as chat bubbles. The
requester's own messages sit in a solid accent-colored bubble.
> - The bubble container sets `text-white`, but the message body renders
through `MarkdownBody`. Tailwind prose tokens (`--tw-prose-body`) win
over the inherited container color.
> - `prose-invert` only lightens the prose text in dark mode. In light
mode the prose body stayed its default dark color, so the text read as
near-black on the blue bubble and was hard to read.
> - This pull request maps the human bubble's prose tokens to the
inherited text color in both themes.
> - The benefit is that the requester's chat text is readable
white-on-blue in light mode, and dark mode stays exactly as it was.
## Linked Issues or Issue Description
<!-- No public GitHub issue exists; described in-PR per the bug report
template. -->
**What happened?**
In light mode, the text inside the user's own chat bubbles in the task
detail view rendered as dark (near-black) on the solid blue accent
background. This made the requester's messages hard to read.
**Expected behavior**
The text inside the user's accent-colored chat bubbles should be white
in light mode, matching the bubble's `text-white` intent. Dark mode
already rendered correctly and should not change.
**Steps to reproduce**
1. Open a chat-style task detail view in light mode.
2. Post a message as the requester (human) so it renders in the solid
blue accent bubble.
3. Observe the body text renders dark on blue instead of white.
**Paperclip version or commit**
Reproduces on `master` (branched from `814cb3367`).
**Agent adapter(s) involved**
Not adapter-specific (core UI bug).
## What Changed
- Add the existing `paperclip-markdown-on-accent` class to the
human-branch `MarkdownBody` in `TaskChatBubble.tsx`. This class (already
used by `IssueChatThread` for the same accent bubble) maps prose
body/heading tokens to `currentColor`, so the text follows the bubble's
`text-white` in both themes.
- Apply the same class to the human-branch `MarkdownBody` in
`TaskChatDescriptionBubble.tsx` (the description-as-first-bubble
surface) for consistency.
- Add unit tests covering that the human accent bubble carries the
on-accent class and the agent/neutral bubbles do not.
## Verification
- `pnpm check:token-gates` → 3/3 CLEAN.
- `pnpm --filter ./ui vitest run
src/components/task-chat/TaskChatBubble.test.tsx` → 9/9 passing.
- Manual: in light mode, the requester's chat bubble text renders white
on blue; agent/neutral bubbles unchanged; dark mode unchanged.
This is a visual change. Snapshot baselines are intentionally not
updated, per `doc/design/DECISION-SHEET.md` → "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".
## Risks
Low risk. The change is scoped to the human-branch `MarkdownBody`
className on two chat-bubble components and only remaps prose color
tokens to the inherited text color. Agent and neutral bubbles are
untouched, and dark mode behavior is unchanged.
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, ~200K context
window, extended thinking mode, with tool use / code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Issue reviews use thread confirmations to record explicit verdicts
> - The server allowed users to resolve review confirmations but
rejected all agent actors
> - This one-way rule prevented an eligible agent reviewer from
completing a review
> - The existing review policy already defines which actor can submit a
verdict
> - This pull request applies that policy to agent confirmation verdicts
on writable issues
> - The benefit is a consistent review gate for users and agents with
preserved audit attribution
## Linked Issues or Issue Description
- Builds on: #10931 (merged into master before this PR)
- Refs #8617
## What Changed
- Allow eligible agents to accept or reject pending review confirmations
on issues they can write.
- Allow a creator agent to withdraw its own pending review confirmation
when the review policy permits it.
- Reuse the review verdict policy check for users and agents.
- Require an explicit, same-run review-confirmation binding so unrelated
board-only confirmations stay protected.
- Preserve board-only tool action confirmations and existing user
attribution.
- Add route and service tests for agent accept, reject, withdrawal,
human-only denial, and user attribution.
## Verification
- `pnpm exec vitest run packages/shared/src/validators/issue.test.ts
server/src/__tests__/issue-execution-policy-routes.test.ts
server/src/__tests__/issue-review-policy.test.ts
server/src/__tests__/issue-thread-interaction-routes.test.ts
server/src/__tests__/issue-thread-interactions-service.test.ts
server/src/__tests__/issue-stalled-review-decision-routes.test.ts` (256
passed after rebasing onto master and the atomic binding fix)
- `pnpm --filter @paperclipai/shared typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm -r typecheck`
- `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run`
- `pnpm build`
## Risks
- The change expands who can resolve pending review confirmations. The
existing issue write checks and review policy limit this access.
- Tool action confirmations remain board-only.
- The pull request depends on the review policy helper from #10931,
which is now merged into master.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
The roadmap marks Agent Reviews and Approvals as shipped. This pull
request fixes a narrow server behavior gap in that shipped capability.
## Model Used
- OpenAI Codex, model `gpt-5.6-sol`, with reasoning, tool use, and code
execution. The runtime does not expose the context window size.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Issues use `in_review` to request a final decision from an
authorized writer
> - The server rejected an assignee agent that tried to close its own
review, even when the issue had no independent-review rule
> - This rejection stopped the default agent workflow and did not
represent the configured execution-stage rules
> - Paperclip needs an open default and explicit issue-level constraints
for teams that require an independent or human verdict
> - This pull request removes the unconditional rejection and adds
`anyone`, `not_creator`, and `human_only` review policies
> - The benefit is a working default path with opt-in, authenticated
verdict controls
## Linked Issues or Issue Description
Refs #10635, #4429, and #10671.
The related public work covers execution-stage independence,
self-approval fallback behavior, and durable review paths. This change
is distinct. It controls who can resolve an issue review verdict. It
keeps configured execution stages active.
## What Changed
- Added a nullable `review_policy` issue column. Null has the same
meaning as `anyone`. The migration does not backfill existing issues.
- Added shared create, update, response, and compact issue contracts for
`anyone`, `not_creator`, and `human_only`.
- Removed the unconditional agent self-approval rejection for
`in_review` issues.
- Added one reusable verdict-actor check for terminal status changes and
pending interaction accept or reject actions.
- Used the authenticated principal type for `human_only`. Agent keys and
run tokens remain agent principals.
- Used the latest transition into `in_review` to identify the requester
for `not_creator`.
- Added actionable 403 responses that name the policy, the allowed
actor, and the next step.
- Kept the configured execution-stage transition and signoff behavior.
- Added focused contract, helper, status-route, interaction-route, and
execution-stage regression tests.
- Updated the implementation specification for the new issue field.
## Verification
- `pnpm exec vitest run packages/shared/src/validators/issue.test.ts
server/src/__tests__/issue-review-policy.test.ts
server/src/__tests__/issue-stalled-review-decision-routes.test.ts
--reporter=dot` passed: 42 tests.
- `pnpm --filter @paperclipai/shared typecheck` passed.
- `pnpm --filter @paperclipai/db typecheck` passed, including migration
numbering and safety checks.
- `pnpm --filter @paperclipai/server typecheck` passed.
- `pnpm run typecheck:build-gaps` passed across server, CLI, plugin
SDK/examples, plugin wiki, and UI.
- `git diff --check origin/master...HEAD` passed.
- SecurityEngineer review approved the authenticated-principal checks
and accepted policy-relaxation tradeoff with no required changes.
- Greptile reviewed the latest head at 5/5 with zero inline comments or
follow-ups.
- The latest-head GitHub rollup passed build, typecheck,
server/workspace tests, serialized suites, canary, e2e, and external
security checks.
## Risks
- The migration adds one nullable text column. It has no default and no
backfill.
- `not_creator` reads the latest recorded transition into `in_review`.
It denies the verdict when it cannot identify the requester.
- Agents can change or relax `reviewPolicy` when they have issue write
access. This is intentional for this issue-level control.
- Null and `anyone` do not add a database query to the verdict path.
- Configured execution-stage checks still run after the issue-level
policy check.
> This work aligns with the completed "Agent Reviews and Approvals" and
"Enforced Outcomes" roadmap items. It does not add a new roadmap
capability.
## Model Used
- OpenAI Codex, GPT-5. The exact deployment ID and context-window size
are not exposed to the agent. The run used reasoning, repository tools,
code execution, and GitHub CLI access.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents ask humans for decisions through interaction blocks: plan
confirmations, structured questions, suggested tasks, and checkbox
prompts
> - Each agent writes these decision prompts in its own style, so
operators can get long or unclear text at the exact moment they must
decide
> - There is no instance-level control that makes agents use a
controlled language for these decision points only
> - This pull request adds an experimental setting that tells agents to
write all user-interaction content in ASD-STE100 Simplified Technical
English, with the context the user needs and the effect of each choice
> - The benefit is faster, clearer human decisions, with agent thinking
and normal responses unchanged
## Linked Issues or Issue Description
Refs #10410 (optional `/simplified-english` skill in the skills catalog;
this PR adds the instance-level toggle for interactions).
**Problem or motivation**
Agent-posted user interactions (plan confirmations, structured
questions, suggested-task proposals, checkbox prompts) are written in
each agent's default style. Operators who want fast, unambiguous
decisions have no way to ask agents to use a controlled language for
exactly those decision points.
**Proposed solution**
Add an experimental instance setting,
`enableSimplifiedEnglishInteractions` ("Simplified English
Interactions"). When it is on, the server sets
`simplifiedEnglishInteractions: true` in the heartbeat wake payload. The
shared wake-prompt renderer, used by every adapter, then emits a
directive: write all user-interaction content in ASD-STE100 Simplified
Technical English, state what information the user needs to decide, and
state what happens for each choice. The directive applies to interaction
content only. Thinking, comments, documents, and other responses keep
their usual style.
**Alternatives considered**
Per-agent instructions work today, but someone must maintain them on
every agent. An instance-level toggle applies uniformly and turns off in
one place. Server-side rewriting of interaction payloads was rejected:
post-hoc translation is lossy and cannot add the decision context that
only the agent has.
**Roadmap alignment**
Extends the experimental settings surface with another opt-in
agent-behavior refinement, consistent with existing prompt-side flags.
## What Changed
- Added `enableSimplifiedEnglishInteractions` to the experimental
instance-settings zod schema, mirror type, and feature catalog (default
off, tier preference) in `packages/shared`.
- Server `instance-settings.ts` normalizes the flag on both read
branches; `heartbeat.ts` reads it once and passes
`simplifiedEnglishInteractions` into `buildPaperclipWakePayload`.
- Shared adapter renderer
(`packages/adapter-utils/src/server-utils.ts`): added the field to
`PaperclipWakePayload`, normalization, and an `- interaction language
(experimental): ...` directive emitted in both fresh and resume prompt
lanes, so one injection point covers all adapters.
- UI: new experimental settings card "Simplified English Interactions"
in `ui/src/pages/InstanceExperimentalSettings.tsx`, in alphabetical card
order; fixtures updated.
- Tests: renderer coverage for flag on/off in both lanes, plus
schema/catalog/UI fixture updates.
## Verification
- From the repo root: `node_modules/.bin/vitest run
packages/adapter-utils` (88/88), `packages/shared` validators (25/25),
server instance-settings + heartbeat suites (47/47 and 70/70 consumer
tests), `ui` settings tests (32/32).
- Typecheck is clean in all four touched packages.
- Manual check: turn the flag on in Settings → Experimental, wake an
agent, and confirm the wake prompt contains the interaction-language
directive; turn it off and confirm the directive is absent.
## Risks
- Low risk: the flag defaults to off, and the only behavior change is
one extra directive line in the wake prompt when an operator turns it
on.
- The directive is advisory to the agent; models can still deviate from
STE. No data or API shape changes; no migration.
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended
thinking, agentic tool use via Claude Agent SDK.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the control plane people use to manage AI-agent
companies.
> - Agents can encounter credentials during work.
> - Directly creating live secrets or bindings would bypass human
governance.
> - Proposal records must remain inert and separate from live secret
resolution until an authorized human approves them.
> - Approval must reuse the existing secret-create and protected
agent-config write paths.
> - This pull request adds the propose, review, approve, and reject
lifecycle.
> - The benefit is that agents can safely hand credentials into
Paperclip without exposing plaintext or gaining authority to activate
them.
## Linked Issues or Issue Description
Follow-on to #9921, which established run-bound agent secret access.
**Problem / motivation:**
Agents can receive credentials during work. There is no governed way for
them to propose a credential or binding without exposing plaintext in
work artifacts or immediately creating live access.
**Proposed solution:**
Store agent-authored proposals outside live secret tables. Encrypt each
proposed value and register exact-value redaction when Paperclip
receives it. Require an authorized human to approve or reject each
proposal. Approval executes through the normal write paths as the human
approver. Binding proposals can target only the proposer or its downward
reporting chain under the restrictive V1 policy.
**Alternatives considered:**
We rejected live secrets with a `proposed` status. That design would put
untrusted rows in resolver, list, and sync paths. It would also allow
uniqueness squatting. We rejected direct agent binding writes because a
binding is an agent-config write and must keep the existing human
permission gate.
**Roadmap alignment:**
This change extends the run-bound agent secret-access foundation in
#9921 with a governed proposal workflow.
## Security Verdict
Q0 SecEng verdict: **PASS-with-required-changes**. The review accepted
the separate proposal-table design and required the implementation to:
- fail closed unless both encryption and exact-value run redaction
registration succeed;
- scrub ciphertext idempotently on reject, withdraw, and expiry, with
audit-visible state;
- treat agent justification as hostile input and foreground action,
target, provenance, and approver permissions;
- snapshot and re-check the target agent plus reports-to chain at
approval to prevent org-chart laundering;
- make cascade approval atomic and fail closed if either secret creation
or binding authorization fails;
- deny low-trust, `skill_test`, `task_bridge`, and non-run-bound sources
consistently; and
- execute approval through the normal human secret/config write paths,
including protected-change gates.
Those requirements are implemented and covered by focused service,
route, and UI tests. Residual V1 risk remains the accepted 14-day
encrypted retention window. Proposal-time redaction also cannot clean a
value that leaked before the propose call.
## What Changed
- Added `company_secret_proposals`, migration `0207`, shared proposal
contracts, and a state-machine service for create, approve, reject,
withdraw, cascade, expiry, and ciphertext scrubbing.
- Added run-bound agent proposal routes and board review routes. The
routes derive provenance from authentication and enforce source
restrictions, company isolation, chain-of-command checks,
approval-as-approver, wake-on-resolution, and dual audit trails.
- Added durable per-run exact-value redaction registration so proposal
values remain redacted on later read surfaces.
- Added the Secrets **Proposals** tab and agent configuration **Proposed
access** rows. The UI shows fingerprint and length only. It also frames
agent justification as untrusted input, runs permission preflight,
supports approve and reject actions, and confirms cascades.
- Updated OpenAPI, agent skill guidance, API reference documentation,
and focused server and UI regression coverage.
- Rebased the branch onto current `master` and renumbered the proposal
migration after `0206`.
## QA Acceptance Results
Q5 QA verdict: **PASS — 9/9 acceptance criteria met**, with one Minor
non-blocking follow-up.
- **AC1:** proposed values never echo, never appear in live
lists/resolvers, and expose only fingerprint + length to board
reviewers.
- **AC2:** restrictive `self_and_reports` matrix passes: self/downward
allowed; upward/lateral denied.
- **AC3:** secret approval uses the normal create path, honors rename
overrides, records proposer/approver provenance, and scrubs ciphertext.
- **AC4:** approved bindings materialize and resolve through the target
agent's runtime list/fetch routes.
- **AC5:** pending-secret bindings require cascade; cascade succeeds
atomically and permission failures leave nothing applied.
- **AC6:** reject, withdraw, dependent rejection, and expiry paths scrub
ciphertext and preserve reasons/audit state.
- **AC7:** token/source and approver denial matrix passes through live
checks plus focused route tests.
- **AC8:** proposal lifecycle events and reused
`secret.created`/config-write events form the required dual audit trail;
origin-issue notification and wake are queued.
- **AC9:** both review surfaces render and execute correctly; UI
approval materializes the binding.
QA also confirmed zero plaintext occurrences for all exercised proposal
values in server logs. The single finding is that the company-level
`bindingTargetPolicy` toggle is not wired yet. V1 is hardcoded to the
restrictive `self_and_reports` policy. The matrix is correct and the
follow-up is tracked separately, so QA classified it as non-blocking.
## Verification
- Focused server proposal and redaction suite: 83 tests pass.
- Focused proposal review UI suite: 54 tests pass.
- Embedded-Postgres migration reapply test: 1 test passes with the
documented 30-second timeout.
- `pnpm --filter @paperclipai/db typecheck` passes, including migration
numbering and safety checks.
- `pnpm --filter @paperclipai/shared typecheck` passes.
- `pnpm --filter @paperclipai/ui typecheck` passes.
- `pnpm check:token-gates` passes with all gates clean.
- Q5 exercised the complete propose, review, approve, bind, and
runtime-resolve flow over real HTTP, JWT, and database paths. It
verified 9/9 acceptance criteria.
## Risks
- Proposal ciphertext is retained encrypted for up to 14 days while
pending. Terminal-state and expiry scrub paths reduce but do not remove
server-compromise risk during that window.
- The V1 target policy is restrictive but not yet company-configurable.
A separate follow-up owns that change.
- A new migration can require another renumber if another migration
lands before maintainers merge this pull request.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex coding agent. The exact runtime model ID and
context-window size are not exposed. The agent used reasoning,
repository editing, terminal execution, Paperclip API, and GitHub CLI
capabilities. Q3 UI work also records Claude Opus 4.8 assistance in its
commit trailers.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Chat-style tasks show an agent's plan in a dedicated "Plan" pane,
and a plan confirmation lets the user accept or request changes to that
plan
> - When an agent asked for confirmation but never actually published
the plan document (it only wrote the plan in a comment or a question),
the Plan pane rendered empty, and the confirmation call-to-action that
used to sit pinned at the bottom of the pane had disappeared
> - A user asked to confirm a plan they cannot see, with no visible CTA,
is stuck — the feature silently fails
> - This pull request closes the gap on both sides: it prevents plan
confirmations that don't point at a real, latest plan revision, it
teaches agents to publish the plan document before confirming, and it
restores the sticky confirmation action bar and an explanatory empty
state so the pane never goes silently blank
> - The benefit is that when a plan is expected, it reliably shows up in
the right pane with reachable accept/revise actions
## Linked Issues or Issue Description
<!-- No public GitHub issue exists; describing in-PR per the bug
template. -->
**Bug report**
- **What happened:** A task in planning mode could present a plan
confirmation while the Plan pane stayed empty (no plan document
rendered), and the plan-card confirmation CTAs that were previously
pinned to the bottom of the Plan pane no longer appeared.
- **Expected behavior:** When a plan is expected, the plan document
appears in the Plan pane; when a plan is genuinely missing, the pane
explains why rather than showing nothing; and the accept/request-changes
CTAs stay visible and reachable while the plan scrolls.
- **Steps to reproduce:** Put a task in planning mode with the
chat-style task view enabled, have an agent create a plan confirmation
without first publishing the `plan` document, and open the Plan tab —
the pane is blank and the confirmation actions are missing.
- **Deployment mode:** Local dev and self-hosted; UI + server.
Related PR (not a duplicate): #9609 "Pin pending confirmations by
composer" pins confirmations in a different surface (the composer); this
PR restores the Plans-pane action bar and the server/agent guarantees
behind it.
## What Changed
- **Server:** Reject a `request_confirmation` whose target is a plan
document unless a plan document exists and the target points at its
*latest* revision, so a confirmation can never reference a plan the pane
cannot render (`readPlanTarget` is now exported for reuse).
- **Agent instructions:** The CEO and default agent instruction bundles
now spell out a plan-publish contract — publish the `plan` document,
re-`GET` it and capture `latestRevisionId`, then create the confirmation
targeting that revision; never present a plan only in a thread comment
or via `ask_user_questions`.
- **UI — sticky CTAs:** Restore the plan confirmation action bar pinned
to the bottom of the Plans tab so accept/revise stay reachable while the
plan scrolls.
- **UI — diagnostics:** Keep the Plan tab visible whenever an issue is
in planning mode (even before a plan document exists) and show an empty
state explaining why the pane is empty instead of rendering nothing.
- **UI — annotations:** Add a `panelPlacement="inline"` mode so the
plan-document annotation panel renders in document flow instead of as a
floating side panel when hosted in the narrow task properties pane.
## Verification
- `pnpm check:token-gates` → 3/3 CLEAN
- `pnpm typecheck` → clean (all packages)
- UI: `pnpm --filter @paperclipai/ui exec vitest run
src/components/issue-properties/IssuePlanConfirmationActionBar.test.tsx
src/components/IssueProperties.test.tsx
src/components/IssueDocumentAnnotations.test.tsx` → 69 passed
- Server: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-thread-interaction-routes.test.ts
src/__tests__/agent-skills-routes.test.ts` → 60 passed
- Manual: with a planning-mode task, the Plan tab stays visible, shows
the plan document (or a diagnostic empty state), and the confirmation
CTAs stay pinned at the bottom.
Visual note: snapshot baselines are intentionally not updated — per
`doc/design/DECISION-SHEET.md` "Per-change snapshot verification demoted
to dormant (Jul 13 2026)". The `storybook-visual` label is intentionally
not added.
## Risks
Low-to-moderate. The server change adds a validation gate on
plan-document confirmations: an interaction that targets a stale or
nonexistent plan revision is now rejected with a 422 instead of being
created. This is the intended guarantee, but any caller that relied on
creating such confirmations will now need to publish the plan document
first (which the updated agent instructions cover). UI changes are
additive to the Plans tab and gated by the existing chat-style-task
experimental flag.
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
thinking enabled, with tool use (file editing, shell, test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task detail page shows a breadcrumb header with the task status
glyph and the task title.
> - The breadcrumb did not show the task identifier, so a reader could
not name the task without opening extra context.
> - Agents and people refer to tasks by identifier, so the identifier
belongs next to the title.
> - This pull request renders the task identifier in the breadcrumb
header, between the status glyph and the title.
> - The benefit is faster reference: a reader sees the task key and the
title together at the top of the page.
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The task detail breadcrumb header. It shows the status glyph and the
task title, but not the task identifier.
**Subsystem affected**
Web UI — the breadcrumb bar on the task detail page
(`ui/src/components/BreadcrumbBar.tsx`,
`ui/src/context/BreadcrumbContext.tsx`, `ui/src/pages/IssueDetail.tsx`).
**Current behavior**
The breadcrumb header renders the status glyph and then the task title.
The task identifier does not appear in the header.
**Proposed behavior**
The breadcrumb header renders the task identifier between the status
glyph and the title. The identifier uses gray monospace styling from
design tokens (`font-mono text-muted-foreground`).
**Reason and benefit**
A reader can name and reference the task from the header without opening
more context. The identifier and the title appear together.
**Breaking changes**
None. The identifier field is optional. Crumbs without an identifier
render as before.
## What Changed
- Add an optional `identifier` field to the `Breadcrumb` type and
include it in the `breadcrumbsEqual` comparison so an identifier change
triggers a fresh render.
- Add a `CrumbIdentifier` helper in `BreadcrumbBar` that renders the
identifier in gray monospace (`font-mono text-muted-foreground`), placed
after the leading status glyph in each crumb variant.
- Wire the issue identifier onto the task crumb in `IssueDetail`.
- Add unit tests that cover the identifier field in `breadcrumbsEqual`
(fresh render on change, no-op on identical value).
## Verification
- `pnpm check:token-gates` → 3/3 gates CLEAN (color literals, arbitrary
bracket values, raw font-size).
- `pnpm --filter @paperclipai/ui exec vitest run
src/context/BreadcrumbContext.test.tsx` → 4/4 tests pass.
- `pnpm typecheck` → the four changed files typecheck clean.
- Manual: open a task detail page. The breadcrumb header shows the
status glyph, then the task identifier in gray monospace, then the
title.
Visual change. Snapshot baselines are intentionally not updated, per
`doc/design/DECISION-SHEET.md` → "Per-change snapshot verification
demoted to dormant (Jul 13 2026)".
## Risks
Low risk. The change is additive and the identifier field is optional.
It touches only the breadcrumb header rendering and the equality check.
No data model or API change.
## Model Used
Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
thinking enabled, tool use enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>