Commit Graph

4325 Commits

Author SHA1 Message Date
Dotta 4e488f9f42 fix(storybook): package avatar PNGs for static branch previews
Build finite avatar presets with the API renderer and resolve preview images under each build path. Verify published PNG bytes and density dimensions before considering deployment successful.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 10:06:56 -05:00
Dotta ad735fbb94 fix(ci): resolve branch dependencies before Storybook deployment
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 09:19:59 -05:00
Dotta 535ff2bce0 Merge origin/master for branch Storybook deployment
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 09:17:22 -05:00
Dotta a20ecce409
feat: publish CODEOWNER-approved Storybook branch previews (#13226)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Maintainers use Storybook to review the board UI.
> - Reviews need public previews of selected repository branches.
> - Each branch needs its own URL so previews do not replace each other.
> - This pull request adds manual, CODEOWNER-controlled publishing to S3
and CloudFront.
> - The action returns stable branch links and permanent build links in
its summary and a Markdown artifact.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The existing Storybook build and manual visual-review workflow.

**Current behavior**

The repository has no manual branch-preview publisher. A single GitHub
Pages site cannot support independent publishers without combining their
output.

**Proposed behavior**

A CODEOWNER selects a source branch and approves publication. Each
branch has a stable CloudFront URL. A completed build becomes the branch
target only after its upload succeeds. The action attaches
`storybook-deployment.md` with the preview links and source commit.

**Reason and benefit**

Maintainers can share multiple branch previews at the same time. Branch
builds have no repository token permissions or AWS credentials.
Dependency caching and install hooks are disabled. The publisher cannot
write runner dashboard files or delete objects.

**Breaking changes**

None. Normal visual checks keep their existing behavior. This does not
change application code or GitHub Pages settings.

**Additional context**

Searched public issues and PRs for Storybook deployment work. No
duplicate deployment proposal was found. This is maintainer
infrastructure, not a roadmap-level core feature.

## What Changed

- Add `Storybook Deploy` with a source-branch input and a manual entry
through `Storybook Visual`.
- Check the original actor and rerunner against default-branch
CODEOWNERS. Require a protected deployment environment with CODEOWNER
reviewers.
- Separate public-source builds with no repository permissions from an
OIDC publisher restricted to the Storybook S3 prefix.
- Publish distinct branch URLs and retain build URLs. Preserve Storybook
deep links across the branch redirect.
- Add the run summary, a downloadable Markdown deployment report,
focused tests, and operator setup docs and IAM policies.

## Verification

- `node --test scripts/__tests__/storybook-deploy.test.mjs`: 19 tests
pass.
- `actionlint .github/workflows/storybook-deploy.yml
.github/workflows/storybook-visual.yml`: passes.
- [Feature branch live publication and deployment-only
rerun](https://github.com/paperclipai/paperclip/actions/runs/34533202273):
passed.
- [Master branch live
publication](https://github.com/paperclipai/paperclip/actions/runs/34533204743):
passed.
- Both public branch URLs render a component story without browser
errors. A deployment-only rerun updates only the selected branch entry
and preserves the previous build URL.
- AWS policy simulation allows Storybook uploads and denies dashboard
writes and object deletion.
- Full local typechecking passes. Full local tests, build, and
current-head PR checks are running.
- [Revised build and Markdown artifact
validation](https://github.com/paperclipai/paperclip/actions/runs/34605623088):
passed. Downloaded the report and verified its branch URL, build URL,
and source commit.
- The public verifier also checks that the stable branch URL points to
this build and rejects stale targets.

## Risks

- Storybook previews are public. Maintainers must publish only public UI
fixtures.
- Retained builds accumulate until an operator prunes them.
- Environment reviewers must stay synchronized with CODEOWNERS. The
workflow fails closed if its environment loses required protection.
- The existing CloudFront distribution is shared with runner reports.
Separate S3 prefixes and a dedicated role prevent the publisher from
overwriting those reports.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, shell tools, and browser
verification. The exact runtime model ID and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 09:13:55 -05:00
Dotta d10cbde815
fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path

> - Paperclip manages agent work through tasks and runs.
> - Recovery continues assigned work when no live execution path
remains.
> - A recovery sweep can read an in-progress task before its run
completes.
> - The sweep can then observe the successful run after completion has
changed the task status.
> - This pull request checks current status and assignment under the
existing enqueue lock.
> - A stale continuation leaves a skipped wake receipt and creates no
run.
> - Task chat also omits an empty continuation cancelled before it
started because its task had become terminal.

## Linked Issues or Issue Description

Related public work: #10779 and #8419. Those older open changes address
terminal disposition across other recovery paths. This change uses the
existing scheduler guard for productive successful-run continuation and
adds real database lock contention coverage.

**What happened?**

Recovery could combine an old in-progress task snapshot with a newer
successful run. It queued an automatic continuation after the task was
done. Dispatch cancelled that run before it started, but task chat
displayed “Couldn't start” below the successful answer. This can happen
after the native runner's finish result has already been accepted. It
does not require a missing comment.

**Expected behavior**

Productive continuation must remain eligible when enqueueing acquires
the task lock. Completion, cancellation, reassignment, or a move away
from in-progress must prevent creation of the run. Actual execution
stops must remain visible.

**Steps to reproduce**

1. Let recovery select an assigned in-progress task whose latest run
succeeded with productive progress.
2. Hold the task row lock in another transaction and change the task to
done.
3. Let recovery attempt to enqueue while that transaction holds the
lock.
4. Commit completion. Before this fix, recovery creates a redundant run
from the stale snapshot.

**Paperclip version or commit**

Reproduced against master at `4042eb1c4` with deterministic integration
tests.

**Deployment mode**

Built from source with PostgreSQL. The bug is in core recovery and is
not adapter-specific.

## What Changed

- Pass the existing status-and-assignee guard for productive terminal
continuation recovery.
- Preserve a skipped wake receipt with the expected and actual task
state, without creating a run.
- Test actual PostgreSQL lock contention for native and legacy
completion, cancellation, backlog, review, blocked state, and
reassignment.
- Omit empty redundant pre-start cancellations from native and legacy
task chat. Preserve stop markers for runs that started.
- Document recovery eligibility at enqueue time.

## Verification

- All seven new race cases failed before the guard was connected.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Task chat suite: 98 tests passed.
- Recovery integration suites: 290 tests passed, including 31
stale-queue tests.
- Full CI verification passed on `7ea71f04d`: all 31 active checks
succeeded, including all test shards, browser tests, build, typecheck,
and canary release dry run. Storybook visual regression was skipped by
its path filter.
- Greptile reviewed this commit at 5/5 with no review threads.
- The local `pnpm test:run` aggregate reported a setup failure in the
unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed
all 231 tests without edits. The duplicate aggregate was stopped after
the complete CI matrix passed; it is not counted as a successful local
full-suite run.

## Risks

Low risk. The backend guard applies only to productive successful-run
recovery. It requires the task to remain in-progress with the same
agent. Other wake sources keep their current policy. The UI change only
suppresses empty redundant cancellations; run records remain available.
No schema change or migration is required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
edits, and local test execution. The exact deployment snapshot and
context window are not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 08:51:50 -05:00
Dotta a05b828bcd
Reduce run polling and workspace inspection amplification (#13174)
## Thinking Path

> - Paperclip manages agent work and shows run progress to operators.
> - Run lists, live events, transcripts, and workspace details must
remain responsive as usage grows.
> - Run-list redaction rereads the full context for every run. Hidden
tabs can still trigger requests through live events and manual timers.
> - Workspace detail reads repeat Git inspection even when concurrent
callers request the same state.
> - This pull request batches registry reads, pauses hidden-tab
refreshes, and caches Git inspection for display.
> - Cleanup keeps fresh Git checks, and redaction keeps company and run
boundaries.

## Linked Issues

**What happened?**
Run-list responses perform one extra database read per run and parse
full context JSON to obtain small secret registries. Hidden tabs
continue transcript reads and event-triggered refetches. Workspace
detail requests repeat Git scans.

**Expected behavior**
A run list reads registries once. Hidden tabs stop recurring run reads
and reconcile when visible. Concurrent workspace detail reads share a
short-lived Git result.

**Steps to reproduce**
1. Open run lists and task transcripts in several tabs while agents run.
2. Hide some tabs and observe transcript and event-triggered requests.
3. Request a 200-run list and count redaction database queries.
4. Request the same workspace detail concurrently and count Git
inspections.

Related: #5255 adjusts polling cadence. This change addresses hidden-tab
lifecycle, batched registry reads, and workspace inspection reuse. No
duplicate with this scope was found.

## What Changed

- Batch heartbeat and live-run redaction into one company-scoped
registry query. Select only registry JSON for run and issue redaction.
- Resolve duplicate secret values once per request. Preserve each run's
registry and remove registry material from responses.
- Suspend company event sockets and transcript reads while hidden.
Refresh active queries and resume transcript offsets on return.
- Prevent queued event invalidations and developer health polling from
fetching in hidden tabs. Gate legacy run-log readers in both UI
variants.
- Exclude legacy plugin placeholder connections from remote health
probes. Select only due connection IDs in SQL before the sweep limit.
Preserve existing plugin records.
- Cache concurrent Git display inspections for five seconds, with at
most 256 entries. Leave close-readiness and cleanup checks uncached.
- Add regression coverage and document the performance behavior.
- Stabilize the existing Rust descendant-lineage fixture: allow a
bounded 30 seconds for 300 durable notifications under concurrent test
load, retaining every correctness assertion and adding timeout
diagnostics.

## Verification

- Regression coverage verifies one registry query for 200 runs, per-run
isolation, request-local secret resolution, decryption failures, Git
cache expiry/bounds, hidden-tab pause, and visibility recovery.
- Real PostgreSQL redaction/run-route suites passed all 57 tests;
workspace-service coverage passed. The health-sweep regression verifies
plugin placeholders and chat connections remain untouched and do not
consume the sweep limit.
- Both legacy transcript viewers retain history and resume their byte
offset after visibility changes. The related visibility/progress/chunk
suites passed all 29 tests. Other focused UI suites and token gates
passed.
- Full `pnpm -r typecheck` and `pnpm build` passed. Affected-package
typechecks/builds passed after review fixes. The concurrent Rust
provider suite passed 84 tests (two ignored), and Rust formatting
passed.
- Full local `pnpm test:run` stopped after the general-server group:
10,538 passed, 65 skipped, four failed. Fresh chat-delivery and
health-sweep reruns passed; building the debug runner fixture cleared
the native-event test. One unchanged native-session recovery assertion
still fails locally with a semantic-digest error instead of the expected
settled-session message. The full local command is therefore not green.
CI runs the later groups separately and skips the two native-session
tests requiring a prebuilt runner binary (confirmed in its 37-test
native-session suite).
- All CI gates pass on final head `ee610e737`: typechecking, general and
serialized tests, browser tests, runner verification, build, and canary
dry run. One server shard passed on its single retry after exposure
fixtures encountered port 42001 where they assumed 42000; that suite
also passed locally (25 passed, three platform-specific skips).
- Greptile reviewed the final head at 5/5 with no actionable findings.

## Risks

- Workspace delivery display can lag local Git changes by five seconds.
Destructive operations still inspect current state.
- Hidden tabs do not receive company live-event notifications until
visible. Active queries refresh on return.
- This change preserves legacy plugin records and does not repair
instance-specific workspace rows. There is no database migration.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser inspection. The exact model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted regressions;
full-suite limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 08:34:24 -05:00
Devin Foley 932c8bec56
fix(ci): bake the managed runtime identity into cloud images (#13210)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed deployments start from the image built by the Cloud
workflow.
> - The managed runtime requests user and group 1001.
> - The image currently builds the node user as 1000.
> - Startup must remap that user, which can walk a large mounted home
directory.
> - This pull request uses the existing Docker build arguments to bake
user and group 1001 into Cloud images.
> - Matching the runtime identity removes that startup work and helps
avoid health-check retries.

## Linked Issues or Issue Description

Refs #13208, #1923, and #7861. Searched open and closed PRs for the
Cloud UID change. The older #7861 addresses build context and volume
ownership repair. This change uses the existing identity arguments in
the Cloud workflow and preserves ownership repair.

**What happened?**

A measured rollout had a container log `Updating node UID to 1001` after
startup. The container stayed at this step for at least 2 minutes 55
seconds before rollback stopped it. The baked node identity was 1000,
while the managed runtime requested 1001. A health check timed out and
the target required a second deployment attempt.

**Expected behavior**

Cloud images should already have the managed runtime identity. A
matching image should skip user and group remapping. Fresh or mismatched
volumes must still receive ownership repair.

**Steps to reproduce**

1. Build the current Cloud image with its default build arguments.
2. Start it with `USER_UID=1001`, `USER_GID=1001`, and a populated home
volume.
3. Observe the startup user remap before the application starts.

**Paperclip version or commit**

`fc06f7f05f42c675be71ff0927b6334405d520ed`

**Deployment mode**

Docker on managed hosts.

## What Changed

- Pass `USER_UID=1001` and `USER_GID=1001` to the Cloud image build.
- Check the pushed digest's baked identity before the entrypoint can
repair it. Then check the normal entrypoint's effective identity and
writable home before publishing the verified full-SHA tag.
- Add a workflow regression and two entrypoint cases for a matching
Cloud identity, including a mismatched volume.
- Document the runtime identity and the first-build cache cost.

## Verification

- Focused workflow and artifact tests: 27 passed.
- Entrypoint tests: 11 passed. Actionlint passed. Full local `pnpm -r
typecheck` passed. Full local `pnpm build` passed. The manual [Cloud
image
build](https://github.com/paperclipai/paperclip/actions/runs/34575473213)
passed on the exact PR head. It checked Sentry, baked and effective
identity, writable home, orphan reaping, and full-SHA publication. The
new identity check took one second. All 30 PR checks passed; the
Storybook workflow was intentionally skipped. Greptile reviewed commit
`114d408f637a0b53e2e2b1339c263779b1e4ae54` at 5/5 with no findings or
open threads.
- The full local suite for the same application source was already run
in #13205. Its macOS general-server phase had 10,471 passes and 70
failures in seven unchanged files. Those failures included missing
Runner fixtures, filesystem errors, timeouts, a port conflict, and a
load-count mismatch. After configuring Cargo and rebuilding fixtures, 37
of 38 native tests passed; one unchanged native-resume assertion still
failed. Linux PR CI passed. This change adds entrypoint tests and does
not change application code.

## Risks

- The first build must rebuild layers that depend on the base image
identity. Later builds can reuse them.
- A future managed runtime identity change must update these build
arguments and checks together.
- The Dockerfile's self-hosted defaults remain 1000. Runtime overrides
and mounted-volume ownership repair remain supported.
- The observed startup delay supports this change, but fleet timing also
includes provider startup, image pull, canary order, and retries. No
fixed end-to-end gain is claimed before a live rollout.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and tool
use. The exact serving model ID and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused workflow tests;
full-suite limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 00:58:27 -07:00
Devin Foley fc06f7f05f
fix(ci): isolate chaos verification by caller workflow (#13208)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments require verified artifacts for the merged source
commit.
> - Cloud readiness and the npm release independently run the same
source checks.
> - Their shared chaos workflow used only the source ref as its
concurrency key.
> - One caller could cancel the other caller's required job for the same
commit.
> - This pull request scopes that key to the caller workflow and source
ref.
> - Both callers can finish their checks without blocking deployment
readiness.

## Linked Issues or Issue Description

Refs #13192 and #13205. Searched for related open issues and PRs; no
duplicate fix was found.

**What happened?**

The master push for `398d304e15739d1ee6105633bd8a0e42c929d33f` started
Cloud readiness and Release together. GitHub cancelled the Cloud
readiness chaos job before it acquired a runner. Its annotation reported
a higher-priority waiting request for the same concurrency group. The
required readiness gate cannot pass after that cancellation.

**Expected behavior**

Cloud readiness and Release must each finish source verification for the
same SHA. Standalone chaos evals must also have a separate group.

**Steps to reproduce**

Merge a commit to master while the npm release queue is empty. Both
callers reach the reusable chaos workflow with the same source SHA. See
[the cancelled
job](https://github.com/paperclipai/paperclip/actions/runs/34569569760/job/103168603926).

**Paperclip version or commit**

`398d304e15739d1ee6105633bd8a0e42c929d33f`.

**Deployment mode**

GitHub Actions on master.

## What Changed

- Add the caller workflow name to the chaos workflow concurrency group.
Retain source isolation and cancellation of duplicate calls within the
same workflow.
- Add a regression test that evaluates the group for Cloud readiness,
Release, and standalone evals at the same source SHA.
- Document the concurrency boundary in the readiness runbook.

## Verification

- `node --test scripts/preview-artifacts.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs` passed: 26 tests.
- The new regression test fails against the previous concurrency key and
passes with this fix.
- `actionlint -shellcheck= -pyflakes=
.github/workflows/runner-chaos-evals.yml
.github/workflows/release-verify.yml
.github/workflows/cloud-readiness.yml` passed.
- `git diff --check` passed.
- The full local typecheck passed for the same application source in
#13205. Its macOS general-server test phase had 10,471 passes and 70
failures in seven unchanged application test files: missing Cargo/Runner
test binaries, filesystem permissions, timeouts, a port conflict, and a
load-test count mismatch. Linux CI test checks passed. The full local
build passed with Cargo on PATH. This PR changes workflow configuration,
its test, and documentation only.
- All CI checks pass on the final head, including typecheck, tests,
browser suites, build, and canary dry run. Greptile is 5/5 with no open
findings. After merge, verify both callers' chaos jobs complete for the
same master SHA and record the resulting readiness time.

## Risks

- Two callers may now run chaos tests at the same time. This uses two
existing GitHub runners, which is the intended cost of independent
verification.
- Renaming a caller changes its concurrency group. The fixed prefix
keeps this child group separate from caller-level concurrency groups.
- The readiness gate continues to require every verification
prerequisite. No gate is bypassed.

## Model Used

- OpenAI GPT-6 / Codex, with reasoning, repository editing, and
command/API tools. Exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (26 focused
workflow/artifact tests)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 23:59:44 -07:00
Devin Foley 398d304e15
docs: measure cloud deployment through target health (#13205)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Hosted deployments need a verified image and migrator for the same
source commit.
> - The cloud readiness workflow certifies those inputs before a
deployment consumer acts.
> - Its completion time does not show when a tenant runs the new commit.
> - This pull request documents each milestone from merge through target
health and fleet completion.
> - Operators can use the evidence to find the slow stage and measure a
complete deployment.

## Linked Issues or Issue Description

Refs #13192, #13188, and #13189. Searched related issues and PRs; no
duplicate timing documentation change was found.

**Issue type**

Missing documentation.

**Where is the issue?**

`doc/cloud-build-readiness.md`, Timing and rollout.

**What's wrong?**

The timing instructions stop at the readiness job. That omits consumer
queues, artifact resolution, and target deployment. An image can be
ready while the tenant still runs an older commit.

**Suggested fix**

Record separate merge, image, readiness, canary health, and fleet
completion timestamps for the same full source SHA. Keep
preparation-only runs out of deployment results.

## What Changed

- Define the evidence needed for each merge-to-deployment milestone.
- Explain how consumer queues can hide upstream build gains.
- Require target source identity as well as health, and report
exclusions, retries, cache state, and queue conditions.

## Verification

- `git diff --check` passed.
- `node --test scripts/preview-artifacts.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs` passed: 25 tests.
- Cross-checked the readiness identity and artifact prerequisites
against the current workflows and consumer contract.
- Full local `pnpm -r typecheck` passed using the session's installed
Rust toolchain. The full local test suite and subsequent build are still
running.
- All CI checks pass and Greptile is 5/5 on the exact head, with no
unresolved findings. This changes one documentation file and adds no
runtime behavior.

## Risks

- Low risk: documentation only. Timing must still use trusted run
evidence and the actual target commit. A single measured run is not a
latency guarantee.

## Model Used

- OpenAI GPT-6 / Codex, with reasoning, repository editing, and
command/API tools. Exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (25 focused
workflow/artifact tests; full checks pending)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 23:21:17 -07:00
Devin Foley d56be3f3fc
fix(ci): verify deployable cloud artifacts independently (#13192)
Verify source, build the cloud image, and wait for exact-source migrator packages concurrently. Emit Cloud deployable v1 only when every prerequisite succeeds for the merged full SHA.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:08:29 -07:00
Devin Foley 5cc51fad06
fix(release): publish exact-source cloud migrators on merge (#13188)
Publish exact-source shared and database migrator packages for each master merge through the existing trusted Release workflow, independently of the full release and image build.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:07:38 -07:00
Devin Foley 6a7025ebe3
ci: skip cloud runner cleanup when disk headroom is ample (#13191)
Skip cloud runner disk cleanup when both the Docker and workspace filesystems have at least 64 GiB free. Preserve the existing cleanup for low, unavailable, or invalid measurements and verify the actual shell behavior across eight scenarios.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:02:42 -07:00
Devin Foley 5c660a32f3
ci: build cloud images independently for each merge (#13189)
Build cloud images independently for each master commit through a reusable workflow. Preserve production release dependencies and image runtime checks, and write cloud registry caches per commit with bounded ancestor imports to prevent overlapping builds from replacing each other's cache.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:42:18 -07:00
Devin Foley 59d74b68b2
ci: cache the native Runner in a separate Docker stage (#13195)
Compile the native Runner from its complete Cargo and protocol inputs in a separate cached Docker stage. Preserve Cargo validation and generated-contract checks during the normal application build, normalize input timestamps across checkouts, and compile the isolated target in PR CI.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:36:55 -07:00
Devin Foley 42961b6ef1
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:30:30 -07:00
Devin Foley 9e970df4c5
ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:10:36 -07:00
Devin Foley 3fd556b8f6
ci: preserve weekly Docker tool cache across commits (#13190)
Keep stable Docker tool installation layers independent of application build version and commit metadata. Preserve the existing weekly tool refresh and runtime build stamp.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:10:31 -07:00
Devin Foley 399daa1f25
ci: publish verified full-SHA cloud image tags (#13187)
Publish the full-source-SHA cloud image tag only after the pushed digest passes runtime, revision, and platform checks. This makes verified images directly resolvable by Cloud.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:10:02 -07:00
Tonio f70accd3a4
fix(onboarding): answer the Claude paste at once, and show the code as dots (#13193)
The Claude card's button only moved to Connecting once the login was stored - a submit, a status poll and a completion read after the paste - so for about a second the customer had done their part and the button still read Waiting for code. The panel now reports the submit as it starts (onCodeSubmitted) and the step shows Connecting from that moment.

That could not simply move the phase earlier: the two-second hold started when Connecting did, so it would have hired whether or not a credential existed. The hire now waits for both the stored login and two seconds of Connecting counted from the paste. onSubmitFailed gives the button back when a submitted code does not become a stored login, the field locks while a code is out, and Cmd+Enter no longer hires mid-connect.

Reports that land after Back are ignored. The panel stays mounted through Back's exit, so a late failure reopened the card being left, and a late success hired a customer who had backed away. The second predates this change; its test fails the same way against master.

The authorization code shows as dots. The OpenAI card is untouched.
2026-09-10 19:52:31 -07:00
Nicky Leach 9effe51b63
fix(server): let the cloud-harness sandbox environment self-heal past operator-drift protection (#13177)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every cloud-harness-managed stack gets one platform-owned "Paperclip
Computer" sandbox environment, reconciled from
`PAPERCLIP_MANAGED_CONFIG` on boot
> - That reconciler deliberately refuses to overwrite a row it
classifies as operator-modified, to protect a self-hosted operator's
hand-edited environment (#10979)
> - But for the cloud-harness-managed row specifically, no operator has
any path to hand-edit it at all — so a hash mismatch there can only be
drift between two platform-driven reconciliation passes, never a real
customization
> - Roughly 10 staging stacks got stuck on a broken sandbox image
because of exactly this: a `sandbox_image` campaign correctly delivered
a fixed snapshot, but the reconciler classified the row as
operator-modified and silently skipped applying it
> - This pull request adds an explicit `platformFullyManaged` flag so
the cloud-harness caller can assert that guarantee and let its own drift
self-heal, without weakening the protection for every other caller
(self-hosted kubernetes-execution-mode, tests, admin routes) where an
operator genuinely can edit the row
> - The benefit is that a sandbox-image rollout can no longer get
silently stuck fleet-wide, while self-hosted operator customization
keeps exactly the protection #10979 built

## Linked Issues or Issue Description

No public issue exists for this internal-instance-discovered bug;
opening directly per CONTRIBUTING.md path B, following the bug report
template fields.

**What happened?**
After a `sandbox_image` campaign delivered a fixed Daytona snapshot
fleet-wide, ~10 of 79 active staging stacks kept booting agents against
the old, broken snapshot. Their `PAPERCLIP_MANAGED_CONFIG` env var and
the reconciler's own bookkeeping (`built_in_managed_resources`) both
correctly showed the new snapshot — but `environments.config.snapshot`,
the field the runtime actually reads to acquire a sandbox lease, was
never updated on those rows.

**Expected behavior**
A `sandbox_image` campaign (or any `PAPERCLIP_MANAGED_CONFIG` delivery)
to the cloud-harness-managed sandbox environment should always converge
that row's `config` to the newly-desired value, since no operator can
have a competing edit to protect.

**Steps to reproduce**
1. Boot a cloud-harness-managed stack; let the reconciler create the
managed sandbox row and record its stock hash in
`built_in_managed_resources`.
2. Somehow cause the row's live content hash to no longer match the
recorded binding hash without an operator ever touching it (in the
field, this happened via drift between two platform-driven
reconciliation passes carried out across a catalog-version bump — the
exact trigger wasn't fully pinned down, but is irrelevant to the fix).
3. Deliver a new `PAPERCLIP_MANAGED_CONFIG` (e.g. via a `sandbox_image`
campaign).
4. Observe `ensureManagedSandboxEnvironment` classify the row
`operator_modified` and skip writing `config`, even though
`updateAvailable: true` is reported and the binding itself already
advanced to the new stock hash.

**Paperclip version or commit**
`master` as of this PR.

**Deployment mode**
Cloud-managed stacks with `enableManagedSandboxOnly` declared (any
Paperclip Cloud–provisioned staging or production stack).

Related PR for context (not a duplicate — this is additive to it, not a
revert): #10979, which introduced the `operator_modified` classification
this PR narrowly opts the cloud-harness path out of.

## What Changed

- `server/src/services/environments.ts`: added `platformFullyManaged?:
boolean` to `ManagedSandboxEnvironmentInput`. When set, a plain
content-hash mismatch against a real prior binding (i.e.
`operator_modified` that isn't an archive-reaffirmation) is reclassified
as `stock_update_available` before the skip-vs-apply branch, so it flows
through the normal update path instead of being frozen.
- `server/src/services/managed-environments.ts`: pass
`platformFullyManaged: true` from both `ensureManagedSandboxEnvironment`
call sites — the main boot ensure and the provider-recovery reactivation
path. These are the *only* two callers driven by
`PAPERCLIP_MANAGED_CONFIG`; `ensureKubernetesEnvironment` (self-hosted
`kubernetes-execution-mode` bootstrap) and every other caller are
untouched and keep the original protective default.
- `server/src/services/managed-environments.test.ts`: updated the two
`toHaveBeenCalledWith` assertions that now include the flag.
- `server/src/__tests__/environment-service.test.ts`: two new tests —
one confirming the bypass applies drift under `platformFullyManaged`,
one confirming archive-reaffirmation still wins even under the flag.

Archive-reaffirmation is deliberately *not* bypassed even under
`platformFullyManaged`: a `sandbox_image` update must never resurrect a
row something else deliberately kept archived after Paperclip's own
provider-unavailability archival. That's a distinct, still-real signal,
orthogonal to config drift.

## Verification

- `vitest run` on `managed-environments.test.ts` and
`managed-resource-drift.test.ts`: 26/26 pass, including the two updated
assertions.
- `environment-service.test.ts` — the file both new tests live in, and
the file holding the two pre-existing tests this change must not regress
("classifies operator drift, preserves the row, and exposes the pending
stock update" and "preserves an existing unmanaged sandbox row holding
the desired name") — requires a real embedded-Postgres instance
(`describeEmbeddedPostgres`) not available in the sandbox this was
developed in; `getEmbeddedPostgresTestSupport()` reports unsupported
there, so the whole file is skipped locally. I traced the reconciliation
logic by hand against all four relevant tests (the two new ones plus the
two pre-existing ones) line by line to confirm the expected outcomes,
but **CI running this suite for real is the actual gate here**, not this
description — please don't merge on a green run of everything else alone
if this suite doesn't show as executed.
- `tsc --noEmit`: zero errors in any of the four touched files. The
pre-existing ~229 errors elsewhere in `server` are unrelated
missing-module issues from packages needing a build step first,
confirmed unchanged by this diff.
- Manually reproduced the underlying bug against real staging data (a
`paperclip-cloud`-managed stack whose `environments` row was stuck
exactly this way) before writing the fix, and confirmed via direct SQL
inspection that the recorded `built_in_managed_resources` baseline
already held the correct desired snapshot on every affected stack — i.e.
the reconciler already *knew* the right answer, it was just refusing to
apply it. That data point is what ruled out "the campaign didn't
actually deliver the update" as the cause.

## Risks

- Scope is intentionally narrow: only the two
`PAPERCLIP_MANAGED_CONFIG`-driven call sites pass the new flag; every
other caller of
`ensureManagedSandboxEnvironment`/`ensureKubernetesEnvironment` is
byte-for-byte unchanged. The two pre-existing regression tests that
specifically cover self-hosted operator-edit protection don't pass this
flag and are unmodified.
- The main residual risk is the unresolved root cause of *why* the hash
drifted in the first place (a race between two close-together
reconciliation passes, or a catalog-version-dependent change to what
gets hashed, most likely) — this PR makes that drift self-healing rather
than fixing whatever produces it. If the drift is being caused by a
genuine concurrency bug (rather than an expected, occasional side effect
of a stock-field/catalog-version change), that bug still exists and
could recur; it just no longer gets stuck when it does.
- Low risk of behavior change for real self-hosted deployments: none of
them can reach the new code path, since only the two now-flagged call
sites exist inside `managed-environments.ts`, itself gated to
`PAPERCLIP_MANAGED_CONFIG` (which self-hosted
`kubernetes-execution-mode` explicitly refuses to run alongside — see
the existing mutual-exclusivity check this PR does not touch).

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, with tool use
(file edits, shell/git, `gh` CLI, direct Postgres inspection of live
staging data via `psql`/`pg`, Railway SSH for on-host diagnosis). No
extended-thinking mode. Standard Claude Code context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — see Verification: the
file holding the four load-bearing tests can't run in this sandbox (no
embedded-Postgres support); traced by hand instead, CI is the real gate
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes — none
applicable beyond the inline doc comments this PR adds
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — pending CI run on this PR
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending review
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 19:23:13 -07:00
Devin Foley c5c80e1feb
ci(release-verify): split server tests five ways like pr-trusted (#13185)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every master push publishes a canary through release.yml, gated by
release-verify.yml — the fleet's staging deploys and the
nightly/beta/stable chain all start from those canaries
> - release-verify splits the server test suite across three shards with
a 20-minute job cap, while pr-trusted splits the same suite across five
> - The server suite grew on 2026-09-10 and the three shards moved to
17-19 minutes; that evening every push-triggered canary run was
cancelled by the 20-minute cap mid-verify, and no canary published after
18:50 UTC
> - This pull request mirrors pr-trusted's five-way server split in
release-verify, putting shards back at the 10-15 minute range with real
headroom
> - The benefit is a canary lane that reports test verdicts instead of
dying on an infrastructure cap

## Linked Issues or Issue Description

**What happened?**

Push-triggered Release runs stopped publishing canaries on 2026-09-10.
Runs at 19:34, 22:30, and 22:37 UTC were all cancelled by "The job has
exceeded the maximum execution time of 20m0s" on a `verify_canary /
General tests (server (N/3))` shard. No canary published after 18:50
UTC, which also starves the staging fleet's continuous deploys.

**Expected behavior**

release-verify's server shards finish well inside the 20-minute cap and
runs conclude with a test verdict, as pr-trusted's five-way split of the
same suite does (10-15 minutes per shard).

**Steps to reproduce**

1. Compare server shard durations in the `verify_canary` job across
2026-09-10: 11-14 minutes in the morning, 17-19 minutes from 15:06 UTC,
over 20 minutes by evening.
2. Observe runs 34521169020, 34537798488, and 34538332689 cancelled at
the cap.

**Paperclip version or commit**

`master` at `d1ba17eec` (current tip; its canary run was one of the
cancelled ones).

## What Changed

- `release-verify.yml`: the `general-server` matrix goes from three
shards to five, byte-for-byte the shape `pr-trusted.yml` already runs,
with a comment recording why.

## Verification

- The identical five-way split runs green on every pr-trusted run (10-15
minutes per shard today, including on PRs merged this evening).
- The suite's own growth (slower chat-connector tests) is being
addressed separately; this PR only removes the artificial cliff.

## Risks

- Low risk: two more runners per verify run; no test content changes. If
shard durations regress further, the cap fires again — which is the
correct signal once shards have honest headroom.

## Model Used

- Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5),
extended thinking, tool use via Claude Code CLI.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-10 17:47:00 -07:00
Devin Foley 2585ed0550
test(server): settle three contention flakes that killed canary verifies (#13186)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every master push publishes a canary through release-verify; the
staging fleet and the nightly/beta/stable chain start from those
canaries
> - The server suite grew substantially on 2026-09-10 and now runs under
real contention in CI, where three tests assert timing properties that
only hold on an idle machine
> - Each of the three failed a release-verify canary run that day (runs
34497348802 and 34517515849), and together with the shard timeouts
(#13185) they kept any canary from publishing after 18:50 UTC
> - This pull request makes the three assertions contention-tolerant
without weakening the invariants they prove
> - The benefit is a canary lane whose verdicts reflect the code, not
the load on the runner

## Linked Issues or Issue Description

**What happened?**

Three server tests failed release-verify canary runs on 2026-09-10 under
CI load:

1. `chat-channels.integration.test.ts › returns a retryable webhook
failure when the delivery insert fails before durable receipt` — the
duplicate-redelivery request drew the retryable 503 instead of an
immediate 200 (run 34517515849).
2. `chat-channels.integration.test.ts › returns ephemeral guidance for
exact Slack controls in channels without creating tasks or actions` — a
`provider_effect` row was read before its async settlement reached
`processed` (run 34497348802).
3. `runner-connection-eval-fixtures.test.ts › resets paired attempts…` —
the fixture's `TRUNCATE companies CASCADE` was chosen as a deadlock
victim (40P01) against the helper app's own background sweeps (run
34497348802).

**Expected behavior**

Verify runs fail only for real regressions. A momentary-contention 503
on a duplicate redelivery, an in-flight settlement row, and a
deadlock-victim reset are all recoverable states the code handles by
design.

**Steps to reproduce**

Run the three tests under a loaded 3-shard release-verify split; the
timing assertions flake. Under `pr-trusted`'s lighter shards they
usually pass, which is why the PRs that introduced them were green.

**Paperclip version or commit**

`master` at `d1ba17eec`.

## What Changed

- The duplicate-redelivery assertion retries on 503 the way Slack itself
would (bounded, 250 ms apart), then asserts the 200 and the unchanged
dedup invariants: duplicate count increments, still exactly one issue.
- The channel-controls settlement read is wrapped in a bounded
`vi.waitFor`, the same pattern the file's durable-receipt paths already
use.
- The runner eval fixture retries its TRUNCATE on Postgres error 40P01,
bounded at five attempts, and rethrows anything else.

## Verification

- All three run green locally: the two chat tests via `-t` filters, the
eval fixtures file in full (6 tests).
- Each change is assertion-shape only; no product code is touched.
- Observation for a follow-up, not this PR: the chat integration file
costs ~15 s transform + ~28 s import per vitest worker before any test
executes — splitting it would give back real shard time.

## Risks

- Low risk: the retries and waits are bounded, so a genuine regression
(permanent 503, settlement that never lands, persistent deadlock) still
fails within the same timeouts as before.

## Model Used

- Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5),
extended thinking, tool use via Claude Code CLI.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-10 17:24:10 -07:00
Nicky Leach d1ba17eeca
fix(adapter-utils): fail fast when the sandbox control channel is lost mid-turn (#13158)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Adapter utilities run agent turns and report their results to the
control plane
> - A lost sandbox control channel can leave an agent turn without a
result
> - The host then waits for the full adapter timeout instead of
reporting the loss
> - This pull request adds a push loss signal and a bounded host wait
> - The benefit is a prompt failure terminal when the agent stops
answering

## Linked Issues or Issue Description

**What happened?**

A sandbox control channel loss during an Agent Client Protocol turn left
the host waiting for the four-hour adapter execution timeout.

**Expected behavior**

The host should detect the terminal channel loss, stop the turn, and
report a safe failure without waiting for the agent.

**Steps to reproduce**

1. Start an Agent Client Protocol turn through a sandbox adapter.
2. Close the duplex control channel while the turn remains active.
3. Observe the host response before the adapter timeout expires.

**Paperclip version or commit**

Test the pull request commit set at
`10b6bbc5525a79fd575298607dd5a25ae448fc8a`.

**Deployment mode**

The change applies to sandbox-backed adapter execution.

## What Changed

- Add `onLoss(listener)` to the duplex bridge handle.
- Register the loss listener at turn start and read losses latched
before turn start.
- Cancel the turn on loss and arm a 30-second host deadline.
- Close the stream locally when the deadline wins and create a host
terminal.
- Derive the public error from the closed `DuplexLossReason` enum.
- Add tests for loss order, cancellation, timeout, and safe error
output.

## Verification

- Run `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`.
- Run `pnpm --filter @paperclipai/adapter-utils exec vitest run
src/acpx-engine/execute.test.ts -t "run-disposition seam"`.
- Confirm that the full pull request workflow passes.

## Risks

The new deadline changes a lost-channel path from a long wait to a
host-built failure after 30 seconds. Orderly completion keeps its
existing behavior. The deadline race against a pending `turn.result` has
no direct test.

## Model Used

OpenAI Codex, GPT-5, with tool use and code execution. The runtime does
not expose a more specific deployment version or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 15:37:05 -07:00
Michael Nguyen 60ee13a0f7
feat: allow operator UI snippets on Cloud instances (#13168)
Adds an optional Cloud-only HTML snippet so operators can load Plain’s
standard chat bubble. **6 files, 16 implementation lines added; 102
additions including tests and docs.**

## Thinking Path

> - Paperclip serves Cloud and self-hosted users.
> - Closed beta users need a way to report problems.
> - Plain provides a ready-made chat widget.
> - Cloud operators can load it through a generic deployment setting.
> - Self-hosted instances ignore that setting.

## Linked Issues or Issue Description

**Subsystem affected**

Server-served UI HTML.

**Problem or motivation**

Enable a chat bubble in Cloud without adding a support feature to the
React app.

**Proposed solution**

Insert trusted `PAPERCLIP_CLOUD_UI_SNIPPET` HTML before `</body>` when
the existing Cloud-managed predicate is true. The setting is off by
default. Related Cloud-gated integration: #12190.

## What Changed

Review the [final
diff](https://github.com/paperclipai/paperclip/pull/13168/files) in this
order:

1. `server/src/cloud-ui-snippet.ts`: the eight-line Cloud gate and HTML
insertion.
2. `server/src/static-index-html.ts` and `server/src/app.ts`: apply it
to static root/index, SPA routes, and Vite HTML.
3. Two test files and `doc/cloud-ui-snippet.md`: boundary checks and
setup instructions.

React UI, customer identity, and database behavior are unchanged. The
existing feedback flag remains. Plain chat is anonymous; no Paperclip
name, email, or organization is supplied.

## Verification

- **Greptile: 5/5**, no actionable findings, reviewed commit
`04bb44515`.
- **[CI
passed](https://github.com/paperclipai/paperclip/actions/runs/34535763243)**,
including build, typecheck, server tests, and end-to-end tests.
- Local: six focused tests, full typecheck, and build passed. The full
local suite has not produced a final result; CI is the completed full
verification.
- Staging deployment and live chat testing remain to be done.

### Staging setup

Set **one server environment variable**, `PAPERCLIP_CLOUD_UI_SNIPPET`,
to:

```html
<script>
(function(d) {
  var script = d.createElement('script');
  script.src = 'https://chat.cdn-plain.com/index.js';
  script.onload = function() {
    Plain.init({ appId: 'liveChatApp_01M26J213F6RR53YRARZVAFCZZ' });
  };
  d.head.appendChild(script);
})(document);
</script>
```

This is the public staging app ID. **No API key or signing secret is
needed.** Deploy to staging and restart the app with this setting. Test
`/`, `/index.html`, and an organization dashboard; send a message and
confirm a support reply returns. Production rollout is separate.

[Plain embed docs](https://www.plain.com/docs/product/channels/chat) ·
[Configuration and
rollback](04bb445151/doc/cloud-ui-snippet.md)

## Risks

Only trusted operators should set this value. The HTML is public and
scripts execute in the app origin; do not include secrets or
user-provided HTML. Plain owns the anonymous browser session, with no
Paperclip account-switch integration. To roll back, unset the variable,
restart, and refresh open tabs.

## Model Used

OpenAI Codex (GPT-6), with repository inspection and code execution.
Exact runtime model identifier and context size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-10 15:30:36 -07:00
Dotta 9915a9e3d3 Merge native recovery fixture verification fix
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/agent-personas-foundation:
  test(runner): accept startup detection of damaged resume input
2026-09-10 17:09:19 -05:00
Dotta b31c96d9c5 test(runner): accept startup detection of damaged resume input
Keep the archive, retained-input, same-task and successful-recovery assertions while accepting either startup or attach as the point that detects the deliberately corrupted fixture.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 17:09:15 -05:00
Dotta 4110f8001b test(runner): accept startup detection of damaged resume input
Keep the archive, retained-input, same-task and successful-recovery assertions while accepting either startup or attach as the point that detects the deliberately corrupted fixture.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 17:09:11 -05:00
Dotta 70a73dc728 Merge avatar stream disposal fix
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/agent-personas-foundation:
  fix(avatars): await stream disposal before cache cleanup
2026-09-10 16:56:08 -05:00
Dotta d78a595063 fix(avatars): await stream disposal before cache cleanup
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:56:00 -05:00
Dotta 148fc40bde fix(avatars): await stream disposal before cache cleanup
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:55:57 -05:00
Dotta a663ef56db test(ui): seed dashboard live personas before visual capture
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:42:55 -05:00
Dotta 81749fbdba Merge reviewed avatar foundation fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/agent-personas-foundation:
  fix(agents): bound cold avatar admission and retain company drafts
2026-09-10 16:37:57 -05:00
Dotta 91471a8f31 fix(ui): retain personas in onboarding and routine rows
Serve on-demand avatars in the standard Storybook test server and inherit the configured browser origin in persona tests.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:37:57 -05:00
Dotta 1293364d3c fix(agents): bound cold avatar admission and retain company drafts
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:37:21 -05:00
Dotta a13ce70a7a fix(agents): bound cold avatar admission and retain company drafts
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:37:08 -05:00
Dotta 0b7d006dd4 feat(ui): use agent personas throughout the app
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:24:23 -05:00
Dotta f5cd8de1b3 feat(agents): persist personas and render cached avatar URLs
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 16:24:00 -05:00
Devin Foley 4042eb1c48
test(release-smoke): follow the connect-step source question and the first-task chat (#13166)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release pipeline promotes canary → nightly → beta → stable, and
the nightly lane is gated by the Docker release smoke, a Playwright walk
of first-run onboarding against the exact published artifact
> - Onboarding changed twice since the smoke was last updated: the
connect step now opens as a model-source question (#12796, #12801), and
the seeded first task now opens as a chat with the lead that
deliberately creates no run until the user answers (#13068)
> - The smoke still waited for an immediate "Connect" button and then
polled for an assignment-triggered heartbeat run, so it failed every
scheduled nightly since 2026-09-03 and blocked all nightly and beta
promotions
> - This pull request updates the smoke to follow the current arc: pick
the Claude source tile, press Connect, launch, then assert the seeded
chat greeting, the opening question card, and the absence of heartbeat
runs
> - The benefit is a release pipeline that can promote current master
again, with the smoke asserting the product's current contract instead
of a removed one

## Linked Issues or Issue Description

**What happened?**

The scheduled nightly lane of `release.yml` has failed every night since
2026-09-03. The `smoke_nightly / smoke` job fails in
`tests/release-smoke/docker-auth-onboarding.spec.ts` at
`expect(connectButton).toBeVisible()`. No nightly has published since
`2026.902.0-nightly.0`, so no beta can promote recent master.

**Expected behavior**

The release smoke follows the current onboarding arc and passes against
a healthy published artifact. The nightly lane promotes the newest green
canary each night.

**Steps to reproduce**

1. Run `PAPERCLIPAI_VERSION=2026.910.0-canary.5 SMOKE_DETACH=true
./scripts/docker-onboard-smoke.sh`.
2. Run `pnpm run test:release-smoke` against the container with the
previous spec.
3. The spec times out waiting for a "Connect" button. The step now shows
a model-source tile row first, and after launch the seeded task is a
chat with no heartbeat run.

**Paperclip version or commit**

Reproduced against published `paperclipai@2026.910.0-canary.5`; spec
updated on current `master`.

## What Changed

- The spec answers the connect step's model-source question: it asserts
the "Connect a model" heading, picks the Claude tile from the "Model
source" radiogroup, and only then waits for the "Connect" footer button
(#12796, #12801 rebuilt the step around that question).
- The spec replaces the assignment-run poll with the first-task chat
contract from #13068: it asserts the deterministic greeting ("Welcome to
Paperclip!"), the opening question card ("What would you like to do?"),
and that the lead has zero heartbeat runs, because launch must not wake
the assignee before the user answers.

## Verification

- Launched the CI harness locally: `scripts/docker-onboard-smoke.sh`
with `PAPERCLIPAI_VERSION=2026.910.0-canary.5` (the newest canary, the
one the next nightly would promote).
- `pnpm run test:release-smoke` against that container: 1 passed.
- The previous spec against the same container reproduces the CI failure
mode first (Connect-button wait), and after the connect-step fix, the
run-poll failure — both match the nightly logs.

## Risks

- Low risk: the change touches only the release smoke spec. If
onboarding's copy for the greeting or the question card changes, the
smoke fails loudly at that assertion, which is this suite's job.

## Model Used

- Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended
thinking and tool use enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-10 12:33:57 -07:00
github-actions[bot] 6728e133f8
chore(lockfile): refresh pnpm-lock.yaml (#13165)
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.

Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
2026-09-10 11:57:40 -07:00
Devin Foley daea92b647
feat(server): accept a Cloud control assertion on the task-drain endpoint (#13125)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server has a task-drain admission hold so operators can stop new
agent work and wait for quiescence before maintenance
> - Cloud deploys restart tenant containers, but the Cloud control plane
has no sanctioned credential for the drain routes, so agent runs are
killed mid-restart
> - The only Cloud credential this server trusts is the runtime identity
assertion, deliberately scoped to the one-time bootstrap health call
> - This pull request adds a disjoint, action-bound Cloud control
assertion accepted only on the task-drain endpoint
> - The benefit is that Cloud can hold new work and drain a stack before
it restarts the container, through the same authorization and audit
paths a human operator uses

## Linked Issues or Issue Description

Refs #12485 (the task-drain admission hold this makes reachable for the
Cloud control plane).

**Problem or motivation**

Cloud deploys restart the container without stopping agent work first.
The task-drain hold from #12485 exists for exactly this, but its routes
require instance-admin board authority. The Cloud control plane holds no
such credential: the runtime identity assertion is accepted only on `GET
/api/health`, by design. So in-flight runs die at every deploy.

**Proposed solution**

A second, deliberately disjoint use of the same Cloud signing key
(`PAPERCLIP_CLOUD_RUNTIME_IDENTITY_JWKS`): a control assertion with its
own JWS type (`paperclip-cloud-control+jwt`), its own audience, an
`action` claim, a request id, and a short maximum lifetime. A new
middleware accepts the `x-paperclip-cloud-control` header only on
`/api/instance/task-drain`, binds each method to one exact action
(`task-drain:read` / `task-drain:start` / `task-drain:stop`), verifies
the assertion against the configured JWKS and
`PAPERCLIP_CLOUD_STACK_ID`, and installs a synthetic instance-admin
board actor so the existing route authorization, validation,
transactional audit, and activity publishing run unchanged (audit rows
record actor id `paperclip-cloud`). The header is rejected with 400
anywhere else, so it can never become an ambient credential. The board
mutation guard exempts the new `cloud_control` source exactly like the
other non-browser lanes.

**Alternatives considered**

Widening the existing runtime identity middleware would conflate a
one-time bootstrap claim with a repeatable management credential and
weaken both. A per-stack minted instance-admin API key would work with
no auth change but adds a long-lived privileged credential per tenant to
store and rotate. The action-bound short-lived assertion keeps
authorization per-call and stateless.

**Additional context**

Self-hosted instances have no `PAPERCLIP_CLOUD_STACK_ID` and reject
every assertion — the feature is inert off Cloud. A runtime identity
token cannot replay as a control token or vice versa (disjoint `typ` and
`aud`, covered by tests). The Cloud-side caller (drain before deploy,
bounded quiescence wait) lands separately in the Cloud control plane.

## What Changed

- `server/src/services/cloud-runtime-identity.ts`:
`verifyCloudControlAssertion` plus the control
header/audience/type/action constants, reusing the existing JWKS
resolution, JWS parsing, and lifetime discipline.
- `server/src/middleware/cloud-control.ts` (new): accepts the header
only on the task-drain endpoint, per-method action binding, installs the
synthetic instance-admin actor on success, 401 on invalid assertions,
400 anywhere else.
- `server/src/app.ts`: mounts the middleware directly after the actor
middleware, so a valid assertion replaces whatever actor the request
otherwise resolved to.
- `server/src/middleware/board-mutation-guard.ts`: `cloud_control` joins
the non-browser exemptions.
- `server/src/types/express.d.ts`,
`server/src/services/authorization.ts`: `"cloud_control"` added to the
actor source unions.

## Verification

- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/cloud-control-task-drain.test.ts
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/heartbeat-task-drain.test.ts
server/src/__tests__/heartbeat-scheduling-suppression.test.ts
server/src/__tests__/cloud-runtime-identity.test.ts` — 87 tests, all
passing.
- `pnpm --filter @paperclipai/server exec tsc --noEmit` reports no new
errors against the base commit's known pre-existing set.
- The new suite covers: acceptance per method, cross-action rejection,
unknown-action rejection, runtime-identity-token replay rejection,
wrong-audience rejection, wrong-stack and self-hosted rejection, expiry
and oversized-lifetime rejection, unknown-key rejection, request id
validation, endpoint containment (400 elsewhere, 400 on unbound
methods), pass-through without the header, and the mutation-guard
exemption.

## Risks

Low risk, additive. No behavior changes without the header; the header
grants nothing outside the one endpoint; each assertion authorizes one
action for at most five minutes; the existing route-level validation,
queued transitions, and audit writes are unchanged. The browser-facing
Cloud proxy strips Cloud headers, and possession of the shared
tenant-session token cannot mint an assertion (signing key never leaves
Cloud).

## Model Used

Claude (Anthropic) — Fable 5 (`claude-fable-5`), extended thinking,
agentic tool use via Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(module doc comments carry the contract)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-10 11:56:28 -07:00
Nicky Leach c1b55537ba
fix(paperclip-runner): bump claude-agent-acp pin to 0.73.0 (#13162)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Claude local adapter can run agent turns through an ACP (Agent
Client Protocol) server, `claude-agent-acp`, instead of the plain CLI
> - Two separate packages each pin their own copy of that dependency:
`packages/adapters/claude-local` (the server-side adapter) and
`packages/paperclip-runner` (which builds the provider pack baked into
every managed sandbox image)
> - `claude-local` moved to `^0.73.0` in #12730, but `paperclip-runner`
was never bumped past `0.70.0` — nothing keeps the two in sync when only
one changes
> - That split means a sandbox image built from `paperclip-runner`'s
provider pack ships a `claude-agent-acp` the server-side adapter was
never actually compatible with
> - This pull request bumps `paperclip-runner`'s pin to `0.73.0`, the
only version that satisfies both packages' declared ranges at once, and
fixes the matching hardcoded version assertion in
`docker/daytona-runner/Dockerfile`
> - The benefit is one consistent, compatible `claude-agent-acp` version
across both the server host and every sandbox image built from this
source, instead of a silent split that only surfaces as a runtime
failure

## Linked Issues or Issue Description

No public issue exists for this specific split; opening directly per
CONTRIBUTING.md path B, following the bug report template fields.

**What happened?**
`packages/paperclip-runner/package.json` pins
`@agentclientprotocol/claude-agent-acp` at an exact `0.70.0`.
`packages/adapters/claude-local/package.json` requires `^0.73.0` (added
in #12730, 2026-09-02). Nobody re-synced `paperclip-runner`'s pin after
that change — the two packages' dependency graphs are independent, so a
bump in one doesn't propagate to the other. `paperclip-runner`'s copy is
what the fleet sandbox image's provider pack actually ships, so every
managed sandbox built from current source carries a `claude-agent-acp`
version the server-side adapter's own declared compatibility range
excludes.

**Expected behavior**
The two packages' `claude-agent-acp` pins should stay within a mutually
compatible range, so a sandbox image built from this source always ships
a version the server-side adapter actually supports.

**Steps to reproduce**
1. Check `packages/adapters/claude-local/package.json`'s
`@agentclientprotocol/claude-agent-acp` range (`^0.73.0`).
2. Check `packages/paperclip-runner/package.json`'s pin for the same
package (`0.70.0` before this PR).
3. Note that `^0.73.0` on a `0.x` version only admits patch releases
(`>=0.73.0 <0.74.0` per semver caret rules), so `0.70.0` falls outside
it.

**Paperclip version or commit**
`master` as of this PR (paperclip-runner still at `0.70.0` prior to this
change; claude-local's `^0.73.0` requirement landed in #12730).

**Deployment mode**
Any deployment that runs `claude_local` agents through the ACP engine
against a sandbox image built from `packages/paperclip-runner`'s
provider pack (managed cloud sandboxes in particular).

Related PRs for context (not duplicates — none of these touch
`paperclip-runner`'s pin):
- #12730 — introduced the `^0.73.0` requirement in `claude-local`
- #11873 — the last time `paperclip-runner`'s pin moved (`0.69.0` →
`0.70.0`)
- #13105 — separately made an unavailable ACP engine a hard failure
instead of a silent CLI fallback, which is what turned this version
split into a visible, run-blocking error rather than a quiet downgrade

## What Changed

- Bump `@agentclientprotocol/claude-agent-acp` from `0.70.0` to `0.73.0`
(exact pin, matching this package's existing pin style for its other
agent-CLI dependencies) in `packages/paperclip-runner/package.json`.
- Update the corresponding hardcoded version assertion (`test
"$(claude-agent-acp --version)" = "0.70.0"`) in
`docker/daytona-runner/Dockerfile` to `0.73.0`, so its own build-time
check stays accurate instead of failing on the next build for an
unrelated reason.
- `pnpm-lock.yaml` is intentionally **not** included —
`pr-trusted.yml`'s `Validate dependency resolution and regenerate stale
lockfile` step already regenerates it for the merge tree and hands it to
downstream `--frozen-lockfile` jobs as an artifact, so a manual lockfile
commit here would just be stale the moment CI runs.

## Verification

- `0.73.0` is a real published version on npm (confirmed via `npm view
@agentclientprotocol/claude-agent-acp versions`), and it's the *only*
version satisfying claude-local's `^0.73.0` range, so this isn't a guess
at compatibility — it's the unique intersection of both packages'
declared ranges.
- `grep -rn "0\.70\.0" docker/ packages/paperclip-runner/package.json`
after this change shows no remaining stale references to the old pin.
- I did not run a full local install/test pass against a hand-updated
lockfile, since regenerating one locally would conflict with leaving
`pnpm-lock.yaml` untouched per the note above; CI's own
lockfile-regeneration step is the intended verification path for a
manifest-only dependency bump like this one.
- Downstream/full verification (does a sandbox image actually built with
this pin work end-to-end) is tracked separately in `paperclip-cloud` —
an unrelated internal-only repo, so not linked here — where a sibling
fix restores the ACP servers to the runtime `PATH` in the fleet sandbox
image itself; both fixes are needed together for a working sandbox, but
this PR is scoped to the version pin alone.

## Risks

- Low risk: single-line dependency version bump plus a matching
test-assertion update, no code changes. `0.73.0` is a patch release
within claude-local's own already-declared-safe range, so there's no
reason to expect it changes behavior tenants depend on.
- The main risk is unknown breaking changes between `claude-agent-acp`
0.70.0 and 0.73.0 that aren't caught by the version-string assertion
alone (that check only confirms the binary reports the right version,
not that its behavior is unchanged). I have not audited that package's
own changelog between those versions.
- `docker/daytona-runner/Dockerfile` is a parallel/reference image (per
its own header comment, meant to stay aligned with the private
`paperclip-cloud/fleet-sandbox-image/Dockerfile`, which is out of scope
here) — this PR does not touch that other Dockerfile.

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, with tool use
(file edits, shell/git, `gh` CLI, `npm view` for version verification).
No extended-thinking mode. Standard Claude Code context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — see Verification: a
manifest-only bump with the lockfile intentionally left to CI's own
regeneration step; no local test run applicable
- [x] I have added or updated tests where applicable — version-pin bump
only, no new behavior to test
- [x] I have updated relevant documentation to reflect my changes — none
applicable
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — pending CI run on this PR
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending review
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 11:35:33 -07:00
Dotta e9828f8bf4
fix: reuse saved model connections during agent setup (#13161)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent setup connects each agent to a model provider.
> - The organization can already hold subscription logins and API keys.
> - The simplified setup flow did not consistently offer those saved
credentials.
> - This pull request restores reuse and selects a saved connection by
default.
> - Agents keep secret references, so reuse does not copy or rotate
credentials.

## Linked Issues or Issue Description

Related change: #13011. Searched public issues and PRs; no duplicate fix
found.

**What happened?**

Onboarding and new-agent setup could ask for a new API key or sign-in
despite an existing saved connection. A general environment auth signal
could also be mistaken for the owner's saved Claude subscription.

**Expected behavior**

Offer saved credentials from the selected organization. Default to a
saved subscription when one exists. Otherwise select a saved API key.
Keep the option to enter a new key or sign in to another account.

**Steps to reproduce**

1. Save a Claude or OpenAI API key, or complete a supported subscription
login.
2. Add another agent with the same provider.
3. Open the provider connection step.
4. Check whether the saved credential is available and selected.

**Paperclip version or commit**

Reproduced on 5cb4f061d after #13011. This branch is rebased onto
current master.

**Deployment mode**

Built from source. Tested in an isolated local test drive with embedded
storage and board access.

## What Changed

- Add a shared saved-credential lookup and picker for active personal
and organization keys.
- Reuse saved Claude subscriptions and saved Codex account homes. Select
an existing connection by default.
- Preserve secret references through connection tests and agent
creation, including the native Claude and Codex runner setup paths.
- Store newly entered onboarding keys separately. Do not rotate another
agent's key.
- Keep explicit choices during metadata refresh. Prevent refreshes from
remounting an active login panel.
- Add integration tests and production-component Storybook stories.
Document connection reuse.

## Verification

- All 5,628 UI tests passed before rebase.
- Twenty targeted server credential tests passed.
- UI typecheck, UI build, token gates, and diff whitespace checks
passed.
- Browser walkthroughs covered onboarding and new-agent setup, saved
keys, saved subscription fixtures, and new sign-in screens.
- Live Claude and Codex API-key probes succeeded. Created both agents
and confirmed that each retained its saved-secret reference. Both secret
versions remained unchanged. Codex passed after one retry.
- Live subscription authentication was not repeated. Subscription flows
use fixture browser tests and integration tests.
- After rebase and the cache fix, all 109 focused onboarding and
agent-creation tests passed.
- Full repository `pnpm build` and `pnpm -r typecheck` passed.
- The full local test attempt encountered timeouts and embedded
PostgreSQL startup failures under parallel load. All four affected
suites passed in isolation: 20 tests, with no code changes. The complete
CI matrix passed, including all workspace, general server, serialized
server, browser end-to-end, build, typecheck, and canary dry-run checks.
- Greptile reviewed commit d53ddf6b82c101d35894587afc9b0d135a5abc55:
5/5, successful check, no review threads.

## Risks

- The default connection mode changes when saved credentials exist. A
saved subscription takes priority over saved API keys; personal keys
appear before organization keys.
- A listed credential can be expired or unavailable in the selected
environment. The existing connection test still checks it.
- No database migration or API contract change is required.

## Model Used

OpenAI Codex, GPT-6. The exact runtime model identifier and
context-window size are not exposed in this session. Used reasoning,
code execution, repository tools, and browser automation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 12:57:53 -05:00
Nicky Leach 86c2e0ac4a
feat(server): log an activity row for each queued-comment queue mutation (#13159)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server records actions that change issues and their queued
comments
> - The queued-comment edit, reorder, and discard routes changed queue
state without activity rows
> - Operators could not inspect these queue mutations in the activity
feed
> - This pull request adds one identifier-only activity row for each
successful queue mutation
> - The benefit is a durable audit trail with no comment text in the
activity log

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The queued-comment edit, reorder, and discard routes now record their
successful mutations in the activity feed.

**Subsystem affected**

Cross-cutting (server and ui).

**Current behavior**

The three queue mutation routes change queued comments but do not write
an activity row. The activity feed has no label for these actions.

**Proposed behavior**

Each successful route writes one activity row with the actor fields,
entity fields, queue identifiers, and queue revision. The discard row
also includes the cancelled run identifier. The activity feed shows a
label for each action.

**Reason and benefit**

Operators need a durable record of queue changes. Identifier-only
details support audit and troubleshooting without storing comment text.

**Breaking changes**

None. The routes keep their existing response and authorization
behavior.

**Additional context**

Each mutation writes its activity row on the same locked transaction
that applies the mutation, so the two commit or roll back together. The
route publishes the live activity event only after that transaction
commits. The separate comment-cancel route opts out of this write and
keeps its existing single activity row.

## What Changed

- Add activity rows for queued-comment edit, reorder, and discard
mutations.
- Include queue identifiers, revisions, ordered comment identifiers, and
cancelled run identifiers as applicable.
- Add activity-feed labels for the three new actions.
- Add route and activity-format tests for the new behavior.
- Write each activity row on the same transaction as the mutation it
records, through a new port method that the adapter implements.
- Keep the comment-delete route opted out of that write, so a
cancellation does not log two rows.

## Verification

- [x] `npx vitest run
server/src/__tests__/issue-queued-comments-routes.test.ts` passes.
- [x] `npx vitest run ui/src/lib/activity-format.test.ts` passes.
- [x] `pnpm --filter @paperclipai/server typecheck` exits 0.
- [x] `pnpm --filter @paperclipai/ui typecheck` exits 0.
- [x] `node scripts/check-module-boundaries.mjs` passes.
- [x] The full CI suite is green.

## Risks

Low risk. The change adds activity rows after successful mutations and
does not change route responses, authorization, or stored comment text.

## Model Used

OpenAI Codex, GPT-5 Codex. The model used repository inspection, Git
operations, and command execution. The context window and reasoning mode
are not exposed by this runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:47:17 -07:00
Nicky Leach 0d8bbf7cf4
refactor(server): move the queued-comment queue mutations into the wake-queue module (#13145)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server coordinates issue execution and agent wake events
> - Queued comment mutations belong to the wake queue that owns their
state
> - Route-local database writes split queue rules across two layers
> - This pull request moves those mutations into the wake-queue module
and keeps route authorization and response mapping
> - The benefit is one transaction boundary with company-scoped writes
and a shared checked response contract

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The queued-comment edit, reorder, and discard endpoints write queue
state directly from the route layer.

**Subsystem affected**

server/ — REST API and orchestration services.

**Current behavior**

The route layer owns database transactions, locks, queue writes, and
wake-row writes for queued comments.

**Proposed behavior**

The wake-queue module owns these operations. The routes keep
authorization, input checks, error mapping, and response mapping.

**Reason and benefit**

The module gives all queued-comment callers one transaction boundary and
applies company predicates to every adapter read and write.

**Breaking changes**

None. The endpoints keep their existing paths and response behavior.

## What Changed

- Move queued-comment edit, reorder, and discard operations into the
wake-queue module.
- Add company predicates to seven queue writes.
- Use the shared queue contract type for mutation responses.
- Add module tests and route tests for the moved operations.

## Verification

- `server/src/modules/wake-queue`: 128 tests pass across 6 files.
- `server/src/__tests__/issue-queued-comments-routes.test.ts`: 19 tests
pass.
- The server TypeScript check reports the same 141 pre-existing errors
before and after this change.
- GitHub Actions must pass the required pull-request checks.

## Risks

The change moves transaction and lock ownership across module
boundaries. The new adapter, use-case, and route tests cover the moved
behavior. No database schema changes occur.

## Model Used

OpenAI Codex, GPT-5, current agent runtime, tool use and code review
support.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 08:41:59 -07:00
Dotta 889947c238
feat: add experimental native chat connectors (#13038)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also ask agents for work in their existing chat tools.
> - Each external conversation needs one task and a current authorized
source.
> - Retries, Stop, and provider failures must not duplicate work or
expose private data.
> - The first chat PR establishes the opt-in provider and data
contracts.
> - This PR adds experimental channel integration and its durable
control plane.
> - Users can request work from connected channels and inspect delivery
in Paperclip.

## Linked Issues or Issue Description

Refs #13100 and #13092. This is the second of exactly two chat PRs.
Foundation #13100 is merged and changed 143 files. Runner prerequisite
#13092 is also merged. This PR changes 400 files against master, below
the 500-file review limit. It contains no wireframe images or HTML
galleries.

## What Changed

- Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat
connections. Keep chat disabled unless the operator enables experimental
chat connectors. Preserve the production GitHub tool connection and its
normal setup path.
- Bind each provider bot identity to one immutable Paperclip agent. Bind
each admitted external conversation to one task. Paperclip owns tasks,
runs, permissions, and audit records.
- Add durable admission, per-conversation queues, questions, task
controls, progress, final replies, images, files, and delivery receipts.
Board comments remain internal unless explicitly sent to the channel.
- Check current identity, provider reach, resource access, credentials,
runtime generation, and exact source before provider effects. Keep
private responses private. Never send raw reasoning, private logs,
credentials, or tool arguments.
- Hold uncertain sends for explicit audited resolution. Make Board
Send-to-channel atomic and idempotent. Keep reconnect and setup
credentials in Paperclip secret storage.
- Preserve current native-runner authority across retries, lost
acknowledgements, and recovery. Keep immutable input and completion
contracts separate from newer user input. Receipt reconciliation cannot
launch a provider.
- Reconcile chat close/new ordering and provider-effect lock order.
Audit resource access changes in the same transaction. Submit only the
selected resource from each UI toggle so stale pages cannot undo
unrelated access changes.
- Drain Codex stdout before certifying process exit. Bound the drain
with the existing shutdown grace. Preserve observed terminal authority
without treating an undrained process as successful or reusable.
- Incorporate master `018ca5da` with its ACP Stop, mobile task layout,
runner packaging, and official lock changes. Preserve dedicated
chat-answer continuations in both directions when ordinary queued
comments are adopted after Stop.
- Fence late adapter readiness behind an earlier Stop for the same run.
Preserve verified cleanup for registered adapters. Handle single Stop,
agent pause, duplicate Stops, and failure release without creating a
false cancellation receipt.
- Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact
failed-chat retry authorization and lineage, retired question-source
suppression, and the block on generic recovery that would discard the
admitted source. Fresh deferred input retains its separate promotion
path.
- Incorporate master `2a05b5ed3` and its queue-admission extraction,
simplified transaction ports, and separate runner CI job. Preserve exact
durable receipts, actor separation, and dedicated-answer isolation
through the new module. A failed receipt insert rolls back the
accompanying deferred-wake merge.

## Verification

Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating
master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are
resolved. This successor fixes two test-harness boundaries exposed by
CI: per-case route-module preparation and actual durable-save completion
before intentional runner termination. Production code and all existing
test/turn deadlines are unchanged. [Exact-head Greptile
review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594)
is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable
findings or open review threads. [Fresh exact-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341)
passes **all 24 jobs**, including Build and both required aggregates.
Normal exact-head guarded merge was attempted and rejected by the
remaining branch approval policy: CODEOWNER review is required and no
human approval is present. Normal **squash auto-merge is enabled** as of
September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified;
no approval bypass or self-approval was used. Earlier-head results below
remain historical evidence, not qualification of this successor.

- Final exact-head Linux evidence: 995/995 chat integration cases; 36/36
agent-skills routes; 35/35 runner live-session cases, including real
process kill/resume; 1948 runner Vitest cases with three existing
benchmark/platform guards; 870/870 API-authority cases; and 104 browser
cases with four existing optional skips. Rust, conformance/replay, full
repository build, typecheck, canary, all server/workspace shards, and
both required aggregates pass with normal CI concurrency. Earlier failed
attempts remain recorded below.

- Latest test-only qualification: 141/141
route/permissions/authentication cases pass in separate cold forks, with
plain server types and independent review clear. The real-runner suite
passes 35/35, with plain runner types and independent review clear. A
controlled premature-save acknowledgement fails as expected; matching
ownership/effect/process evidence, rejected saves, real turn outcome,
test abort, and pre-kill liveness are covered. No local reproduction of
the original CI scheduling failure is claimed. The preceding [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34479680858)
passes 21/24 jobs, including all 995 Linux chat cases and browser
aggregate (104 passed, four existing optional skips); only Build, the
skills serialized shard, and the required verification aggregate fail.
Its exact-head Greptile review was 5/5. Both failed job logs are
retained.

- Final fixture qualification: all eight focused Discord cases and all
995 chat integration cases pass. The exact modal statement/PID is
observed before taking the real connection lock; the test then proves
its actual blocking relationship before mutation. Original SQL
execution, provider behavior, negative assertions, and 1s/15s timeouts
remain unchanged. Independent review is clear and test/production hashes
remain frozen. The preceding [CI
attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777)
passed 22 jobs, including Build/runner, typecheck, canary, all other
test shards, and browser aggregate (104 passed, four existing optional
skips); the two fixture failures and failed verification aggregate
remain recorded, not relabeled as a pass.

- Current queue-module composition: 308/308 recovery/batching/queue/Stop
tests; 995/995 full chat integration; 89/89 module tests, including real
PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary
tests; plain server and UI types. All four actual local process/ACP
browser paths pass in 1.4 minutes. Fresh databases, no skips or retries,
stable reviewed source hashes. The initial boundary failure is retained;
its no-op service wrapper was removed without changing recovery context
or weakening the check. An exploratory standalone test-directory
typecheck fails because its new upstream transformation config is not a
standalone typechecking project; standard CI/build does not invoke it,
and no configuration was weakened to suppress those diagnostics.

- The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed
[all 24 CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958)
and exact-head Greptile review at 5/5. Required CODEOWNER review
prevented its normal merge before master advanced again.

- Final extracted-module composition: 307/307 recovery, batching, queue
and Stop-control tests; 995/995 full chat integration; 49/49 module
tests including eight PostgreSQL adapter cases; and 19/19 issue-update
tests. Plain server types pass. All four actual local process/ACP
browser paths pass in 1.3 minutes. Fresh databases, no skips or retries
in these cohorts, frozen source hashes, and independent review clear.

- The preceding head `3e4e1c1c` passes [all PR CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820),
including Build and required `ci / verify` and `ci / e2e`. Both the
original Rust failure and the previously load-sensitive lineage fixture
pass with unchanged Linux concurrency. Master advanced afterward and
required this reconciliation.
- Final master composition: 448/448 focused UI tests, 186/186 adapter
tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI,
server, shared, and adapter types pass. Token gates and diff checks
pass. Independent server and UI reviews are clear.
- Stop-registration regression: both real-service cases fail against
exact `a95` source and pass with the fix. The full corrected
recovery/control suite passes 265/265. Duplicate-owner and failed-Stop
controls also pass. Plain server types pass. The readiness barrier
prevents provider startup without adding an acknowledgment to an already
terminal run.
- Final qualification strengthens terminal-field equality and repeats
both affected cases successfully on a fresh database. All four actual
local process/ACP browser paths pass again in 1.3 minutes, without skips
or retries. The final screenshot shows Cancelled, a paused subtree,
retained input, and no error toast.
- Two new actual-service regressions fail before the merge fix. They
prove that queued-comment adoption could consume a dedicated chat answer
or add unrelated input to that answer. The fixed four-case cohort
passes, including ordinary upstream continuation and adapter Stop
controls. Full recovery passes 257/257. All four actual local
process/ACP Stop browser flows pass in 1.4 minutes, without skips or
retries, on a fresh database.
- The unchanged runner artifact was qualified with 171/171 transport
tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11.
Six controlled reader tests prove the exit/drain repair. Its local
serial Rust workspace passed 546 top-level cases plus two invoked
helpers; the later passing Linux CI supplies default-concurrency
evidence.
- Prior exact-source full chat integration passes 995/995. Settings
regressions cover concurrent stale pages, 501 destinations, pending
state, rejected updates, and explicit retry. These deterministic tests
do not prove live provider behavior.
- Retained failed attempts and their causes are in the [qualification
log](afe19299d0/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md).
The first merge adapter run timed out while macOS slept for 290 seconds.
Its unchanged repeat passed with a temporary sleep guard. No assertion,
deadline, or CI gate was weakened.

Review commands include `pnpm --filter @paperclipai/server exec vitest
run src/__tests__/heartbeat-process-recovery.test.ts
src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec
playwright test --config tests/e2e/playwright.config.ts
tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh
disposable databases. See the [browser
runbook](afe19299d0/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md)
for provider setup and separate live acceptance steps.

## Risks

- This remains experimental. Deterministic tests and bounded live
evidence do not establish every provider feature, tenant, permission
layout, or media shape. Teams work-tenant qualification is still open.
- Failed and uncertain provider effects remain visible and can require
operator action. A transport receipt does not prove recipient
visibility.
- Native controller and runner artifacts must remain compatible.
Preserve lease ownership, terminal authority, source binding, and
quarantine during future changes.
- Access and audit rows commit together, but activity notifications
remain best-effort. This is not a new durable event outbox.
- The PR operation does not deploy a live server, replace its runner, or
change provider permissions. Remaining live qualification is documented
in the [temporary
handoff](afe19299d0/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md).

## Model Used

OpenAI Codex assisted with implementation, tool execution, testing, and
review. The work records `gpt-6-astra` assistance. The environment does
not report a context-window size. No private reasoning traces are
included.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:06:45 -05:00
Devin Foley bce976d60d
feat: bind an agent to a Codex login whose account differs from the company default (#13067)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The `codex_local` adapter signs agents in to OpenAI, and a company
keeps one default Codex identity in its shared company home
> - A login with a DIFFERENT account than the company default is
deliberately kept out of the shared home — one agent's sign-in must not
switch every unbound agent's credentials — but that left the
cross-account login inert: nothing connected the agent the operator was
configuring to the credential the login stored
> - The stored credential and its company secret already exist; only the
last mile — an agent actually using them — was missing
> - This pull request reports a non-secret binding claim on the
authenticated login and lets the agent page bind that one agent's
`CODEX_HOME` to the account's secret, exactly and only when the
identities differ
> - The benefit is that multi-account Codex becomes one click on the
agent that needs it, with company-wide identity untouched

## Linked Issues or Issue Description

**What happened?**

On an agent's detail page, "Sign in with Codex" using a different OpenAI
account than the company default succeeds but changes nothing for that
agent. The credential lands in the per-identity store and a company
secret names it, but the agent keeps using the company default. The Test
keeps reporting that authentication is needed, and no repeat login
helps.

**Expected behavior**

When the operator deliberately signs an agent's page in with a different
account, that agent starts using that account. Agents that were not part
of the action keep the company default. A same-account login keeps
working through the shared company home with no per-agent pinning.

**Steps to reproduce**

1. Configure a company whose Codex home holds account A.
2. Open a `codex_local` agent's detail page with a sandbox environment
and complete "Sign in with Codex" using account B.
3. Press Test. Before this change the agent still resolves account A and
the authentication-needed check returns.

## What Changed

- `packages/adapters/codex-local` — the prerequisite shield:
`isCodexAuthCachePath` recognizes per-identity credential-store entries,
and `seedManagedCodexHome` refuses to symlink, heal, or
API-key-overwrite an entry's `auth.json`. The seeding pass runs before
every probe and execute; without the shield, an agent bound to an entry
would have its stored login silently swapped for the host credential.
Static shared config files still copy in. Rotation already survives
binding: the sandbox copy-back writes rotated credentials into the
identity-keyed store slot.
- `server` — the promotion records whether the company default home
ended on a different account than the login (any read failure degrades
to `false`, so the client can never be told to bind wrongly). After the
terminal commit, the routes layer remembers a non-secret claim — the
opaque account-home secret id plus that verdict — in a bounded in-memory
map, and merges it into the owner read of an `authenticated`
`codex_local` session. A restart drops the claim; the panel then shows
plain success.
- `packages/shared` — `CodexAccountBindingClaim` on the owner session
response. It carries no account identifier and no credential byte.
- `ui` — the login panel reports the claim upward once. The edit-mode
form binds the agent's `CODEX_HOME` to the secret and saves in one step,
only when `companyIdentityDiffers` is true. Same-account logins bind
nothing on purpose: the company-home refresh already carried them, and
an unbound agent keeps following the company default across rotations.
Create mode is unchanged.

## Verification

- Adapter suite: 381 passed, 1 skipped (includes the new store-entry
shield tests and the path-predicate cases).
- Server suites (8 files): 130 passed, 15 skipped — including two new
route tests that drive a login to `authenticated` and assert the claim
with both identity verdicts.
- UI render suite: 85 passed — including a panel test that the claim is
reported upward exactly once.
- `tsc --noEmit` clean in `packages/shared`, the adapter package, and
`ui`; `server` clean for the touched file.

## Risks

- The bind changes one agent's configuration through the normal
agent-update patch, initiated by the operator's own login on that
agent's page. The failure direction of every fallback is "offer
nothing": a missing claim, a restart, or an unreadable company home all
degrade to no bind.
- The seed shield narrows what the seeding pass may touch; homes outside
the credential store behave exactly as before, covered by the existing
seed tests.
- Builds on the sign-in credential-resolution fix (#13064), now merged;
this branch is rebased onto master and the diff contains only the
binding feature. Supersedes #13066, which GitHub auto-closed when its
stacked base branch was deleted on merge.

## Model Used

Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended
thinking, agentic tool use in Claude Code (terminal).

**Related PRs (searched; no duplicates found):** #12740, #12082, and
#9621 touch adjacent Codex credential sync paths; #8495 is the standing
hardening effort for probe auth seeding.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
standalone docs cover this flow; the behavioral contracts are documented
in-line at each changed site)
- [x] I have considered and documented any risks above
2026-09-10 07:15:01 -07:00
Nicky Leach e25a6b797f
fix(runner): close two timing windows in the capability-live suspend path (#13143)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner manages live agent sessions and their suspend
path.
> - A turn-timeout promise could reject before `reconcileActiveTurn()`
attached its handler.
> - A short provider-drain budget could reject a valid suspend on a
loaded continuous-integration host.
> - This pull request closes both timing windows and adds deterministic
regression tests.
> - The benefit is a fail-closed suspend path that does not report false
failures under load.

## Linked Issues or Issue Description

**What happened?**

Capability-live tests failed under load. A turn-timeout promise could
raise an unhandled rejection during a slow interrupt round trip. An idle
provider drain could also exceed its one-second proof budget during
suspend.

**Expected behavior**

The suspend path must observe turn-timeout rejections and allow enough
time for one provider command round trip. It must still fail closed when
the runner does not prove durable suspension.

**Steps to reproduce**

1. Run the capability-live tests on a loaded four-vCPU
continuous-integration host.
2. Delay a `turn/interrupt` reply beyond the turn timeout.
3. Close a live session and observe the suspend barrier.

**Paperclip version or commit**

`01b442b2926ac2010a2fcfbda6592802472ef07f`

**Deployment mode**

Built from source with the Paperclip Runner test suite.

## What Changed

- Attach the turn-timeout rejection handler inside `armTurnWaiter()` at
promise creation.
- Remove the redundant per-call-site guard in `sendMessage()`.
- Use a uniform five-second provider-drain proof budget, capped by the
outer preparation deadline.
- Remove the unused boolean return from the provider-turn-stop helper.
- Add deterministic tests for the delayed interrupt and the short close
grace period.

## Verification

- `npx vitest run
packages/paperclip-runner/src/live/live-session.test.ts` passed locally
with zero skipped tests.
- `npx vitest run
packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts`
passed locally with zero skipped tests.
- The two new tests appeared in the local run output and were not
skipped.
- The package type-check passed locally.
- GitHub Actions must confirm the full continuous-integration suite,
including `Verify Paperclip Runner`.

## Risks

- The provider-drain wait now allows up to five seconds before the outer
deadline caps it.
- The fail-closed suspend barrier remains unchanged.
- The test suite still depends on the Rust runner binary for
capability-live tests.

## Model Used

OpenAI Codex, GPT-5. The runtime provided tool use and code execution.
The runtime did not provide a context-window value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 07:04:55 -07:00
Nicky Leach 2a05b5ed34
ci: split runner verification from build (#13142)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip uses GitHub Actions to verify changes before release.
> - The Paperclip Runner has a separate verification boundary.
> - The build job currently runs this verification before the workspace
build.
> - This pull request moves runner verification into its own parallel
job.
> - The benefit is clearer CI results and less wait time for independent
work.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The trusted PR and release verification workflows run Paperclip Runner
verification inside the Build job.

**Subsystem affected**

Cross-cutting (GitHub Actions CI workflows).

**Current behavior**

The Build job runs `pnpm --filter @paperclipai/paperclip-runner
check:all` before it builds the workspace. A runner verification failure
appears as a Build failure. The workspace build cannot run in parallel
with runner verification.

**Proposed behavior**

Each workflow has a `Verify Paperclip Runner` job with the same
checkout, dependency install, and command. The Build job only builds its
required outputs. Both jobs run after the same gate and policy jobs.

**Reason and benefit**

The runner command is an independent verification boundary. A dedicated
job gives it a clear status and allows it to run in parallel with Build.

**Breaking changes**

None. The same runner verification command still runs in both workflows.

## What Changed

- Added a dedicated `Verify Paperclip Runner` job to the trusted PR
workflow.
- Added a dedicated `Verify Paperclip Runner` job to the release
verification workflow.
- Kept the Build jobs independent and retained their existing build
commands.
- Updated the trusted-workflow policy test for the additional
dependency-install job.

## Verification

- Ran `git diff --check`.
- Ran `node --test ./scripts/__tests__/e2e-shard.test.mjs`.
- Ran `pnpm exec prettier --check .github/workflows/pr-trusted.yml
.github/workflows/release-verify.yml`.
- Confirmed both jobs retain their prior runner, dependency, and policy
prerequisites.

## Risks

Low risk. The runner verification job repeats the existing setup. It
adds one parallel GitHub Actions runner to each affected workflow.

## Model Used

OpenAI Codex, GPT-5.6, 128k context window, reasoning and tool-use
capabilities.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 01:18:06 -07:00
Nicky Leach 92c5c1ac3d
fix(build): give two orphaned test setup files a governing tsconfig (#13141)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip uses separate server and UI TypeScript projects to build
and test the app
> - The test transform tool finds the nearest tsconfig that includes
each test file
> - Two test setup files had no governing tsconfig inside the repository
> - The lookup then read a tsconfig outside the repository and stopped
test runs when that file was invalid
> - This pull request adds a server test tsconfig and includes the UI
setup file in the UI tsconfig
> - The benefit is stable test configuration in every repository
worktree

## Linked Issues or Issue Description

**What happened?**

The test transform tool walked outside the repository because two test
setup files had no tsconfig that included them. A stale or invalid
parent checkout then stopped server and UI test runs with
`TSCONFIG_ERROR`.

**Expected behavior**

Each test setup file must use a governing tsconfig inside the
repository. Test runs must not depend on a tsconfig outside the
repository.

**Steps to reproduce**

1. Run the server test command in a clean worktree.
2. Run the UI test command in the same worktree.
3. Observe that the transform tool searches above the repository for the
setup files when no local tsconfig includes them.

**Paperclip version or commit**

`04eb274fa12007392e468fc808d3ad12fbdcb02e`

**Deployment mode**

Built from source with the local test commands.

**Installation method**

Built from source with pnpm.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Database mode**

Not database-related.

## What Changed

- Add `server/src/__tests__/tsconfig.json` for the server test setup
directory.
- Add `vitest.setup.ts` to the `include` array in `ui/tsconfig.json`.
- Keep `server/tsconfig.json` unchanged, so the build graph does not
change.

## Verification

- Run `pnpm exec vitest run --project @paperclipai/server
server/src/modules/wake-queue/domain/policy.test.ts`.
- Run `pnpm exec vitest run --project @paperclipai/ui
ui/src/adapters/adapter-display-registry.test.ts`.
- Run `pnpm --filter @paperclipai/server run typecheck`.
- Run `pnpm --filter @paperclipai/ui run typecheck`.
- Confirm that the test commands report no `TSCONFIG_ERROR`.
- Confirm that the full pull request workflow passes.

## Risks

This change adds one scoped server tsconfig and expands one UI tsconfig
include list. It does not change application runtime code, database
schema, or production build settings. Risk is low.

## Model Used

OpenAI Codex, GPT-5, accessed through the Codex agent with tool use and
repository execution. The model used reasoning and code inspection to
assist this change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 00:57:02 -07:00
Nicky Leach ae0c1fbbd5
refactor(server): move the admission half of the deferred wake state machine into the wake-queue module (#13136)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service admits wake requests while an issue has an
active execution run.
> - That admission branch mixes wake policy, database reads, and
database writes in one service.
> - This structure makes the wake-queue boundary hard to test and
extend.
> - This pull request moves the admission policy and its database
adapter into the wake-queue module.
> - The result keeps heartbeat orchestration small and makes the
admission behavior testable in isolation.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The heartbeat service now delegates deferred wake admission to the
wake-queue module. The module keeps the existing merge, defer, and
ordinary-wake outcomes.

**Subsystem affected**

server/ — REST API and orchestration services.

**Current behavior**

The heartbeat service contains a 146-line branch that reads wake state,
chooses an outcome, and writes the result.

**Proposed behavior**

The wake-queue module owns the pure admission decision and the adapter
reads and writes. The heartbeat service calls one module method.

**Reason and benefit**

This boundary reduces service coupling and lets module tests cover the
admission policy. The change keeps the existing reason strings and
outcomes.

**Breaking changes**

None. The change preserves the current behavior and public API.

**Additional context**

This pull request follows [PR
#13132](https://github.com/paperclipai/paperclip/pull/13132), which
merged the first slice of this refactor. I searched GitHub for duplicate
and related pull requests before opening this pull request.

## What Changed

- Move deferred wake admission policy into
`server/src/modules/wake-queue`.
- Add module ports and a PostgreSQL adapter for the admission reads and
writes.
- Keep the existing wake outcomes and stored reason strings.
- Extend the module boundary check to reject service imports from the
application layer.
- Add unit and adapter tests for the moved behavior.

## Verification

- `node --test scripts/check-module-boundaries.test.mjs` passes.
- The `server/src/modules/wake-queue` suite passes 58 tests.
- The eight pinned heartbeat and queued-comment tests remain unchanged
and require CI verification.
- Every continuous-integration check must reach a terminal green state
before merge.

## Risks

The refactor changes the location of wake admission logic. A missed
adapter condition could change deferred wake behavior. The tests cover
the policy outcomes and the adapter writes. The residual tenant-scope
risk remains documented in the review record.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution. The implementation
author ran the tests and prepared the commit set.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 00:55:48 -07:00