Commit Graph

1677 Commits

Author SHA1 Message Date
Dotta bf95a7eae2
fix(ui): stabilize active-run steering queue (#12834)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The issue detail page shows a live agent run and accepts follow-up
instructions.
> - A follow-up must stay in a stable queue until the user sends,
reorders, or removes it.
> - Native runners can receive a steering event in the active run.
> - Legacy runners must interrupt the active run and start a follow-up
run.
> - The current UI moved comments between the queue and the transcript
and could show duplicate text or ambiguous chronology.
> - This pull request makes the queue projection durable, keeps each
message in one clear place, and labels when queued input was actually
steered or delivered.
> - The benefit is predictable steering with stable ordering, no
duplicate messages, and visible causal timing.

## Linked Issues or Issue Description

Refs #11374.
Refs #12591.

**What happened?**

During an active run, a new follow-up could first appear as a transcript
bubble and then move into the steering queue. After a steer or remove
action, it could appear again. Progress text could also repeat the final
response text. Once consumed, a queued bubble displayed only its
original submission time even though it moved to its later causal slot,
and a native run split by steering looked like two unrelated runs.

**Expected behavior**

An active-run follow-up must appear in the queue immediately. A native
steer must move it once into the active run. A legacy interrupt must
move it once into the follow-up run. A removed item must stay removed.
Progress text that is identical to the final response must appear once.
Consumed follow-ups must show both queue and steer/delivery times, and
post-steer native segments must identify themselves as continuations of
the same run.

**Steps to reproduce**

1. Start a long-running task.
2. Send two or more follow-up messages while the agent is active.
3. Reorder the messages and remove one message.
4. Send the first queued message as steering.
5. Observe the queue and transcript during and after both runs.

**Paperclip version or commit**

The problem reproduced on commit `da1e40302`.

**Deployment mode**

Local development with the embedded database.

## What Changed

- Project queued comments into the steering well for native and legacy
live runners.
- Send native steering to the active run and use interrupt-and-follow-up
for legacy runners.
- Keep optimistic queue order stable across refreshes and roll back
failed actions.
- Remove discarded comments from the transcript cache and keep them
removed when the queue becomes empty.
- Collapse only the final progress occurrence matching the durable
response, including across steered transcript segments.
- Show `Queued … · Steered …` for same-run input and `Queued … ·
Delivered …` for successor-run input at their causal positions.
- Label settled and live post-steer segments `Continued after steering`
and time them from the steer boundary.
- Add regression tests for queue display, steering, fallback interrupt,
reorder, remove, rollback, duplicate text, causal timestamps, and
live/settled continuation headers.

## Verification

- Ran the final focused steering/chronology UI suite with 233 passing
tests.
- Ran the activity-service regression suite with 5 passing tests.
- Ran the broader queue-focused UI suite with 298 passing tests before
the final chronology refinement.
- Ran `pnpm -r typecheck` successfully.
- Ran `pnpm build` successfully.
- Ran `pnpm check:token-gates` successfully.
- Tested native steering in a real browser with a 90-second baseline
wait and a three-second steering correction.
- Confirmed that the old final response did not appear before the
steered response.
- Tested three queued messages in a real browser.
- Confirmed that reorder changed delivery order and that the removed
message was never sent or shown again.
- Tested a legacy runner in a real browser.
- Confirmed that it used the interrupt fallback and showed the follow-up
once.
- Reloaded a saved mixed-steer/successor-run thread and confirmed the
causal timestamps and continuation header render in the correct
positions.
- The complete macOS suite reaches five unrelated platform assertions in
workspace-runtime tests. Two compare `/var` with `/private/var`. Three
require Linux `/proc` listener data. GitHub Actions provides the
authoritative Linux run.

## Risks

- Low risk. The change is limited to issue-chat queue projection and
transcript presentation.
- The server run-history API adds only a read-only `contextIssueId`
projection; the database schema does not change.
- Optimistic actions restore the prior UI state when a request fails.

## Model Used

- OpenAI Codex with GPT-5, extended reasoning, browser automation, shell
tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-04 12:54:00 -05:00
Dotta b84964e5a2
fix(runner): stabilize local paid E2E recovery (#12836)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paid runner E2E tests verify the complete runner, control-plane, and
UI path.
> - A server restart could load a fresh task page while Playwright still
waited on an unsettled Vite navigation lifecycle.
> - The current server also ignored the isolated Vite cache path and
skipped Vite's per-request HTML transform from the known-green runner
snapshot.
> - A one-cell paid run then exposed that download-artifact v8 removes
the artifact-name directory for one pattern match.
> - This pull request restores the Vite contract, proves a fresh
document after restart, and accepts only the exact singleton artifact
layout.
> - The benefit is reliable local runner qualification without weaker
UI, source, or artifact checks.

## Linked Issues or Issue Description

Refs #12769
Refs #12828
Refs #12829
Refs #12833

**What happened?**

The structured-question restart test could time out after the
replacement server returned the task route and rendered the durable
pending interaction. A focused one-cell rerun passed the paid test but
failed aggregation because download-artifact v8 flattened its single
artifact.

**Expected behavior**

The test must prove that a new document loaded after the server restart
and that the same pending interaction survived. The aggregate must
accept the exact documented singleton download layout while it continues
to reject ambiguous or foreign artifacts.

**Steps to reproduce**

1. Run the local ACPX-Codex structured-question restart-resume cell.
2. Restart the isolated server while the question waits for an answer.
3. Observe that the route and task UI can reload before Playwright
settles the navigation promise.
4. Run a paid campaign with one selected cell.
5. Observe download-artifact v8 extract the sole campaign directory
directly into the requested path.

**Paperclip version or commit**

The local campaign reproduced the navigation failure at
`3586956a1b794b3cb4a9c5f57ffb7355e2b0c46d`. The one-cell aggregate
reproduced the singleton layout at
`f487660c0a06ba06ca140b57386f21ed39f13120`. This fix is
`de4ccceff453a4b39436bf9a2eb8f03924151af7`.

**Deployment mode**

Local development and paid GitHub Actions.

**Installation method**

Built from source.

**Agent adapter(s) involved**

ACPX-Codex. The Vite and aggregate fixes are provider-neutral.

## What Changed

- Prove a new post-restart browser document with an in-memory sentinel.
- Tolerate only Playwright's navigation timeout before the exact UI and
API checks run.
- Honor `PAPERCLIP_VITE_CACHE_DIR` in the embedded Vite server.
- Limit dependency optimization to the real UI entry.
- Run `vite.transformIndexHtml` for each request while caching only the
branded source template.
- Accept download-artifact v8's flattened layout only for one expected
cell with one unique recognized campaign.
- Keep source SHA, source ref, workflow URL, execution ID, attempt, and
unexpected-entry validation.
- Add focused positive and negative regressions for Vite rendering and
singleton artifact selection.

## Verification

- Exact 45-cell local campaign
https://github.com/paperclipai/paperclip/actions/runs/33888939013 passed
44/45. Its only failure was the post-restart navigation false negative
fixed here.
- Exact focused rerun
https://github.com/paperclipai/paperclip/actions/runs/33891207957 passed
the ACPX-Codex restart cell first attempt with the same session, two
durable runs, the terminal marker once, and cleanup complete.
- The focused Vite renderer suite passed 2/2 tests.
- The focused rerun-artifact selector suite passed 12/12 tests.
- Prettier and `git diff --check` passed.
- An exact-head 45-cell confirmation is pending.

## Risks

Low to medium risk. The Vite change restores known-green per-request
transforms and isolated cache behavior. It can affect all development UI
loads. The paid matrix and ordinary CI will verify that behavior. The
singleton selector remains fail-closed for ambiguous layouts and
validates every result source.

## Model Used

OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code
execution, and parallel focused agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 11:16:20 -05:00
Nicky Leach 184b014c25
feat(telemetry): add the agent.task_run event and emit it at every terminal run transition (#12809)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip records agent run outcomes through telemetry and run
lifecycle services
> - Terminal run transitions need one consistent event for outcome
analysis
> - The current paths do not report every terminal transition through
one event
> - This pull request adds the agent.task_run event and emits it at each
terminal transition
> - The benefit is complete run outcome data without exposing raw task
identifiers

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Paperclip telemetry reports agent activity, but it does not report every
terminal task run through one event.

**Subsystem affected**

Cross-cutting (multiple of the above): packages/shared telemetry and
server run lifecycle services.

**Current behavior**

Several run paths write a terminal status without a matching
agent.task_run telemetry event.

**Proposed behavior**

Each terminal run transition emits one agent.task_run event. The event
records the terminal state and uses the existing pseudonym helper for
the optional task identifier.

**Reason and benefit**

Complete terminal-run data helps operators measure agent outcomes. The
pseudonym helper prevents the raw task identifier from leaving the
installation.

**Breaking changes**

None. The change adds an event and keeps existing event behavior
compatible.

## What Changed

- Add the agent.task_run telemetry contract and client helper.
- Reuse the existing pseudonym helper for the task identifier. The
helper hashes the identifier with a per-installation salt and returns 16
hexadecimal characters. The raw identifier never leaves the
installation. Existing identifiers do not move.
- Emit one event from each legacy, native, recovery, and issue terminal
transition.
- Keep emissions outside database transactions and make delivery
best-effort.
- Add regression tests for event shape, hashing, terminal transitions,
and emission failures.
- Document the event and its privacy rule in the telemetry data
contract.

## Verification

- `npx tsc --noEmit` in `server/` passes at the submitted commit.
- The pull-request CI suite must pass. CI is the authority because local
Vitest has a known dependency artifact.
- The added regression tests cover event output shape, per-installation
hash divergence, raw identifier handoff, omitted identifiers, and
non-throwing emits.

## Risks

- A missed terminal path could reduce event coverage.
- Telemetry delivery remains best-effort and cannot change run
finalization.
- The pseudonym helper uses installation-specific state, so identifiers
differ between installations.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution. Context window details
were not provided.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I have addressed all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-04 08:21:32 -07:00
Dotta 7dfc769f3b
fix(server): honor proxy trust for forwarded host (#12832)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-04 10:02:28 -05:00
Dotta af3023f1e3
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path

> - Paperclip manages AI agents that perform work.
> - Paperclip Runner connects durable task runs to local provider
processes.
> - The full-stack paid matrix exposed failures after the runner
integrity repair.
> - Verified JavaScript entrypoints lost their relative module graph
when Linux executed them through descriptor paths.
> - Returned provider startup errors also remained pending and became
indeterminate after recovery.
> - Sparse Codex tool lifecycle events lost the `write_document`
identity before task transcript projection.
> - This pull request repairs those three boundaries and makes the
structured-question fixture deterministic.
> - The benefit is repeatable provider startup, exact failure replay,
and correct inline Plan placement.

## Linked Issues or Issue Description

Refs #12721 and #12700.

**What happened?**

The paid runner matrix failed ACPX and OpenCode startup before provider
session creation. The runner journal then replaced the original startup
error with an indeterminate recovery result. Native Codex saved a Plan
but rendered it only as a fallback card. A legacy Claude waiting reply
could also echo the reserved terminal marker before the answer arrived.

**Expected behavior**

Verified JavaScript providers must start from immutable
descriptor-backed artifacts. Returned startup failures must persist as
terminal failed command results. Native tool lifecycle updates must
preserve the `write_document` boundary. Pre-answer fixture output must
not contain the reserved terminal marker.

**Steps to reproduce**

1. Run the local provider cells in the Runner Full-Stack E2E workflow.
2. Observe ACPX and OpenCode fail during `session.open` before provider
execution.
3. Observe recovery report `execution_indeterminate` instead of the
original startup error.
4. Run the native Codex Plan cell and observe the fallback Plan card
after the tool activity row.
5. Run the legacy Claude structured-question resume cell and observe an
early marker echo in waiting prose.

**Paperclip version or commit**

`0f9452101740835ce0b1488a204bf48acd5bafc3`

**Deployment mode**

Local development with the paid GitHub Actions acceptance workflow.

## What Changed

- Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM
entrypoints before hashing and verified descriptor launch.
- Anchor ACPX dynamic provider package resolution at a
controller-derived provider-pack root and keep that root out of the
provider child environment.
- Persist executor-returned startup errors as redacted durable failed
command results while retaining indeterminate recovery for true process
death.
- Coalesce sparse native tool items by stable ID so a late
`write_document` name, input, and result reach the transcript boundary
once.
- Forbid the structured-question fixture from spelling or announcing its
reserved terminal marker before the user answers.

## Verification

- Rust and TypeScript regression tests cover durable failed replay, true
crash ambiguity, bundle closure, package-root derivation, environment
filtering, exact Codex tool lifecycle coalescing, and prompt
determinism.
- Local execution is intentionally limited to formatters and static diff
checks. GitHub Actions will run tests, type checks, builds, and security
checks.
- After ordinary CI is green, scoped paid cells will validate one ACPX
launch, one OpenCode launch, native Codex Plan projection, and legacy
Claude structured resume before a complete matrix rerun.
- Prior failing matrix:
https://github.com/paperclipai/paperclip/actions/runs/33682434315

## Risks

- Bundling changes the bytes covered by provider launch hashes.
Provider-pack generation already hashes the final built files.
- ACPX still loads qualified provider packages dynamically. The
controller supplies a normalized package root, while existing version,
digest, path, and descriptor checks remain active.
- Durable `failed` is terminal. Replays return the same redacted result
and do not execute the provider effect twice.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5 with agentic reasoning, repository
inspection, code editing, Git, parallel subagents, and GitHub Actions
coordination. The exact deployed snapshot and context-window size are
not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked related public work or described the bug in
this PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] No documentation change is required for this runtime repair
- [x] I have considered and documented the risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 07:58:44 -05:00
Devin Foley 54dd0f4868
feat(agents): grant new agents hire permission by default (#12814)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent permissions control which agents can create or hire other
agents (`canCreateAgents`)
> - Today only CEO-role agents get this permission by default; every
other agent starts without it
> - Teams that want agents to delegate and build out their own teams
must flip the toggle on each hire, and most operators want delegation to
work out of the box
> - This pull request makes `canCreateAgents` default to enabled for new
standard-trust agents, while low-trust agents keep a disabled default
> - The benefit is that agent teams can grow without per-agent
permission toggling, while low-trust containment and checkout protection
stay intact

## Linked Issues or Issue Description

Related (not fixed by this PR): #8064 also decouples an authority from
`agents:create`.

**Subsystem affected**

Server agent permissions (`server/src/services/agent-permissions.ts`),
authorization (`server/src/services/authorization.ts`), the shared
`agentPermissionsSchema` validator, and the UI trust-preset helper.

**Problem or motivation**

New agents cannot hire other agents unless an operator enables
`canCreateAgents` on each one. Only CEO-role agents get the permission
by default. This blocks delegation-by-default workflows. Operators must
toggle the permission for every hire.

**Proposed solution**

Default `canCreateAgents` to `true` for newly created agents. Apply and
persist the default at creation only. Stored rows without an explicit
value stay fail-closed at read and enforcement time. Keep the default at
`false` when the agent's permissions record marks it low-trust (the
`low_trust_review` preset or a trust boundary). Explicit values always
win. Decouple `tasks:manage_active_checkouts` from `canCreateAgents` so
the default-on flag does not let a peer agent write over another agent's
checked-out issue.

**Alternatives considered**

Granting the default only at the route layer would leave stored rows and
enforcement out of sync. Keeping the checkout authority coupled to
`canCreateAgents` would void the active-checkout write protection once
the flag is default-on. A per-company setting adds configuration surface
without a clear need; explicit per-agent overrides already exist.

**Roadmap alignment**

Governance and trust-preset work already separates standard-trust from
low-trust agents. This change follows that line: capability by default
for standard trust, containment by default for low trust.

## What Changed

- `normalizeAgentPermissions` now takes a `create`/`stored` context.
Creation writes get the new default: enabled unless
`permissionsImplyLowTrust()` detects the low-trust review preset or a
trust boundary. Stored rows without an explicit value normalize to
disabled (fail-closed). The role parameter is gone.
- `agentPermissionsSchema` no longer injects `canCreateAgents: false`
when the field is omitted. The server-side default applies instead.
- `authorization.ts` normalizes raw agent rows for `agents:create`, so
enforcement matches what the API reports for legacy rows.
- `tasks:manage_active_checkouts` no longer rides on `canCreateAgents`.
CEO role, explicit grants, and the manager chain remain the paths.
- `agents:create` is denied outright inside any resolved low-trust
execution context (agent, project, issue, or run policy). The default-on
flag can never reach the legacy creator allow there.
- The UI trust-preset helper sets `canCreateAgents: false` when an agent
is switched to the low-trust preset, instead of carrying the old value
forward.
- `doc/CLI.md` describes the new default for `teams install`.
- Tests pin the default matrix (standard, low-trust, explicit overrides)
on the server and in the UI helper.

## Verification

- `cd server && npx vitest run
src/__tests__/agent-permissions-service.test.ts
src/__tests__/agent-permissions-routes.test.ts
src/__tests__/low-trust-red-team-routes.test.ts
src/__tests__/authorization-service.test.ts` — 143 tests pass.
- Broader sweep: 18 suites that touch `canCreateAgents` (hire,
pending-approval, teams catalog, portability, built-in agents,
plugin-managed agents) pass locally.
- `cd ui && npx vitest run src/lib/trust-policy-ui.test.ts
src/components/TrustPresetSection.test.tsx src/pages/NewAgent.test.tsx
src/pages/Agents.test.tsx` — passes.
- Typecheck is clean for the changed files in `packages/shared`,
`server`, and `ui`.

## Risks

- Behavioral shift: agents created after this change persist
`canCreateAgents: true` unless low-trust. Pre-existing agents keep their
stored value. Legacy or malformed permission records without an explicit
value stay fail-closed at read and enforcement time; they never gain the
authority retroactively.
- Low-trust runs can no longer create agents at all, even when the agent
carries an explicit `canCreateAgents: true`. Before this change, that
combination could hire. The red-team suite and a new authorization test
pin the denial.
- Narrowing: a non-CEO agent with `canCreateAgents: true` loses implicit
`tasks:manage_active_checkouts`. The manager chain and explicit grants
still provide it. This narrowing is deliberate; without it, the
default-on flag would let any peer bypass active-checkout write
protection.
- No migrations. No API shape changes. Low-trust defaults are covered by
the red-team regression suite.

## Model Used

- Claude Fable 5 (`claude-fable-5`), Anthropic — via Claude Code CLI
with extended thinking and tool use (code search, editing, local test
execution).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-03 23:26:51 -07:00
Magnus b98badb246
fix(recovery): exclude hidden issues from stranded recovery and continuation wakes (#5648)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The recovery subsystem watches assigned issues and re-wakes an agent
whose run ended without finishing the work
> - Intake can hide a duplicate issue by setting `hiddenAt` while
leaving its status and assignee in place
> - The stranded-issue query and the terminal-run cleanup both ignore
`hiddenAt`, so a hidden issue is re-woken on every cycle
> - Nothing on the board shows the hidden issue, so the repeated wakes
have no visible cause
> - This pull request adds a hidden-issue guard to both predicates and a
test for each
> - The benefit is that hiding an issue stops recovery work on it, with
no other change in behavior for visible issues

## Linked Issues or Issue Description

**What happened?**

When intake marks an issue as a duplicate it sets `hiddenAt` but leaves
the status at `todo` or `in_progress` with the agent still assigned. The
stranded-issue recovery timer selects that issue on every tick and
queues an `issue_continuation_needed` wake for it. The agent's run on
the hidden issue fails or is cancelled, the terminal-run cleanup queues
immediate recovery for the same issue, and the cycle repeats
indefinitely. Hidden issues are invisible on the board, so nothing a
person can see explains the wakes.

**Expected behavior**

A hidden issue is never a recovery candidate. Stranded-issue
reconciliation skips it, and a failed, timed-out or cancelled run on it
releases the issue without queuing a continuation.

**Steps to reproduce**

1. Assign an issue to an agent and leave it `in_progress`.
2. Hide the issue (set `hiddenAt`, for example by marking it a duplicate
through intake) without changing its status or assignee.
3. Let a run on that issue fail, or wait for the stranded-issue recovery
timer.
4. Observe a new `issue_continuation_needed` heartbeat run queued for
the hidden issue on every cycle.

**Paperclip version or commit**

Reproduced on `master` when this PR was opened (May 2026). The two
predicates are unchanged on current `master`; this branch is rebased
onto it.

**Deployment mode**

Not deployment-specific: both guards are in the server's recovery and
heartbeat services and apply in every mode.

## What Changed

- `server/src/services/recovery/service.ts`: `isNull(issues.hiddenAt)`
added to the `reconcileStrandedAssignedIssues` candidate query, so
hidden issues never enter the stranded set.
- `server/src/services/heartbeat.ts`: `!issue.hiddenAt` added to
`issueNeedsImmediateRecovery`, so terminal-run cleanup releases a hidden
issue instead of queuing a continuation.
- `server/src/__tests__/heartbeat-process-recovery.test.ts`: one test
per guard. A failed run on a hidden issue queues no recovery run, and a
hidden stranded issue is left out of reconciliation.

## Verification

- `heartbeat-process-recovery.test.ts` covers both guards; CI runs it
against embedded Postgres.

## Risks

Low. Both changes narrow an existing predicate to exclude rows that
already carry `hiddenAt`; visible issues take exactly the path they take
today. A hidden issue that genuinely needs recovery would have to be
unhidden first, which matches how hidden issues behave everywhere else
in the board.

## Model Used

The original two-line fix was authored by @im0xMagnus. The rebase onto
current `master`, the two regression tests, and this description were
produced with Claude (claude-fable-5-1, extended thinking, tool use)
driven by a Paperclip maintainer through Prospector's triage flow.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
2026-09-03 22:05:20 -05:00
Dotta f449b05bc5
feat(apps): unify permissions and action testing (#12802)
## Thinking Path

> - Paperclip is the control plane for companies that use AI agents.
> - Apps give humans and agents controlled access to external services.
> - The existing app detail flow split permissions, tests, setup, and
activity across separate pages.
> - The split made access rules harder to understand and made reconnect
work hard to find.
> - New write actions also defaulted to Ask first, which did not match
the intended connection policy.
> - This pull request combines permission control and action testing,
removes the setup page, and moves connection activity into Audit.
> - The benefit is one clear place to configure, test, reconnect, and
review each app.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The installed app Permissions, Test, Setup, and Activity views.

**Subsystem affected**

Cross-cutting. This change updates the React UI, shared app defaults,
server permission behavior, tests, smoke scripts, and connection
documentation.

**Current behavior**

App access and action testing use separate pages. The app detail view
also links to a setup page after installation. Connection activity uses
a separate tab. New write actions default to Ask first.

**Proposed behavior**

Permissions uses the connection access language from the initial flow.
It includes searchable Read and Write sections, a three-state permission
control, and a Test dialog for each action. Reconnect appears below a
Needs attention header on Permissions and Review. Old Setup and Test
links redirect to Permissions. Old Activity links redirect to the
filtered company Audit feed. New write actions default to Allowed.

**Reason and benefit**

A person can understand and test app access without moving between
several pages. Reconnect work stays visible where the person reviews the
connection. Audit events use one consistent feed and filter model. New
connections have the intended default policy.

**Breaking changes**

The Setup, Test, and app Activity tabs are removed. Existing deep links
redirect to their replacement pages. Existing saved action permissions
do not change. Only defaults for new write actions change.

**Additional context**

This builds on the managed app connection work in #12728. A search found
no duplicate open pull request or issue.

## What Changed

- Combined action testing with Permissions.
- Added searchable Read and Write action groups.
- Added Off, Ask first, and Allowed controls with tooltips.
- Added an action Test dialog with agent selection, arguments, and
formatted results.
- Removed the installed-app Setup and Activity tabs.
- Added reconnect guidance to Permissions and Review when a connection
needs attention.
- Routed connection activity into the company Audit feed and preserved
the Apps & tools filter in streamlined Audit.
- Moved connection removal to the Connectors-page management menu.
- Made new write actions default to Allowed across connection creation
paths.
- Updated regression tests, browser suites, smoke scripts, and
connection documentation.

## Verification

- `pnpm check:token-gates`
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-access-service.test.ts
ui/src/components/AppConnectionSidebar.test.tsx
ui/src/pages/apps/AppDetail.test.tsx
ui/src/pages/apps/AppNotConnected.test.tsx
ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx
ui/src/pages/apps/Connections.test.tsx
ui/src/pages/apps/composio-services.test.ts
ui/src/pages/audit/AuditFeed.test.tsx
ui/src/pages/tools/PasteConfigTab.test.tsx` (517 tests passed)
- `pnpm exec vitest run ui/src/pages/apps/app-detail/TestPanel.test.tsx
ui/src/pages/audit/AuditHub.test.tsx
ui/src/pages/audit/AuditFeed.test.tsx
ui/src/pages/apps/AppDetail.test.tsx ui/src/pages/apps/Browse.test.tsx`
(96 tests passed)
- Targeted Playwright verification for connection removal, rename on
Permissions, inline action testing, and Smoke Lab Audit evidence (5
flows passed)
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm test:run` completed with 5,755 passing tests and 20 unrelated
macOS harness failures. The failures use `/tmp` versus `/private/tmp`,
invalid ports above 65535, and workspace fixtures outside this change.

## Risks

- Low migration risk. This change has no database migration.
- Old app-detail URLs depend on redirect compatibility.
- New connections grant write actions by default. Finalization remains
configure-authorized and audited, Ask first and Off remain available per
action, and existing connections keep their saved policy.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected - check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, exact model ID `gpt-5`. The client does not expose the
context-window size. The model used reasoning, repository tools, code
execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 21:23:26 -05:00
Maxxsong7 505e7b40fc
fix: guard listComments against non-UUID afterCommentId to prevent 500 errors (#8695)
## Thinking Path

> - Paperclip is an open-source app for managing AI agents
> - The issue history subsystem stores comments per issue, with
cursor-based pagination via the `after` query parameter
> - `GET /issues/:id/comments?after=<commentId>` looks up the anchor
comment by UUID to get its created_at timestamp
> - When agents store an incorrect or truncated comment ID (e.g.
`670427ab` instead of `670427ab-e0ae-4a54-959e-2b13a2e33d14`), Postgres
throws `invalid input syntax for type uuid` before the anchor-not-found
guard can execute
> - This surfaces as an unhandled 500 and causes agents to fail when
doing incremental comment reads on any issue
> - This pull request adds a UUID validation guard in `listComments`
using the already-imported `isUuidLike` helper
> - The benefit is that invalid cursors get a clean empty-array response
instead of a 500, matching what already happens when a valid UUID simply
isn't found

## Linked Issues or Issue Description

Refs #2612 (a different 500 on the same `after=` cursor path, fixed
earlier; this PR covers the malformed-cursor case that remains).

**What happened?**

`GET /issues/:id/comments?after=<value>` returns a 500 when `after` is
not a UUID. The route trims the query value and passes it straight to
the anchor lookup, so Postgres raises `invalid input syntax for type
uuid: "670427ab"` before the anchor-not-found guard can run. Any agent
that stored a truncated or malformed comment ID as its pagination cursor
gets stuck in a 500 loop on that issue.

**Expected behavior**

A cursor that cannot name a comment behaves like a cursor that names a
missing comment: the endpoint returns `[]`.

**Steps to reproduce**

1. Pick any issue id on a running instance.
2. Call `GET /api/issues/<issue-id>/comments?after=670427ab` (8 hex
characters instead of a full UUID).
3. Observe a 500 with `PostgresError: invalid input syntax for type
uuid: "670427ab"`, where a full-but-unknown UUID such as
`00000000-0000-0000-0000-000000000000` returns `[]`.

**Paperclip version or commit**

`master` at the time this PR was opened (June 2026). The `listComments`
anchor lookup in `server/src/services/issues.ts` is unchanged on current
`master`, so the failure still reproduces there.

**Deployment mode**

Local dev (`pnpm dev`). Not deployment-specific: the failure is in the
server's comment-listing service, so it reproduces in every mode.

## What Changed

- `server/src/services/issues.ts` — added `if
(!isUuidLike(afterCommentId)) return [];` guard in `listComments` before
the DB anchor lookup, using the already-imported `isUuidLike` helper

## Verification

```bash
# Start the dev server
pnpm dev

# Pass a truncated UUID — should return [] instead of 500
curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=670427ab"
# Expected: []

# Pass a valid full UUID that doesn't exist — should also return []
curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=00000000-0000-0000-0000-000000000000"
# Expected: []

# Pass a valid full UUID that exists — should return comments after that cursor
curl -s "http://localhost:3100/api/issues/<any-valid-issue-id>/comments?after=<real-comment-uuid>"
# Expected: array of comments
```

## Risks

Low risk. The change only adds an early-return guard for values that are
provably invalid UUIDs. The code path for valid UUIDs is unchanged. The
existing behavior for anchor-not-found (returning `[]`) is preserved for
invalid UUIDs, which is the correct semantic (cursor not found → no
comments after it).

## Model Used

Claude Sonnet 4.6 (`claude-sonnet-4-6`) via Paperclip CTO agent, tool
use + code execution mode, 200K context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip CTO <cto@paperclip.ai>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
2026-09-03 20:47:19 -05:00
Nicky Leach 333abdd2c2
test(plugin-worker): remove the wall-clock race from the duplex buffered-replay tests (#12799)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The plugin worker manager runs agent plugin workers through duplex
channels.
> - The duplex buffered-replay tests check data that arrives before a
listener attaches.
> - The tests used a fixed 60 ms sleep as the barrier for worker output.
> - Worker startup and output latency can exceed that delay under load.
> - This pull request uses a worker exit frame as a deterministic
barrier.
> - The benefit is stable test results without a product code change.

## Linked Issues or Issue Description

**What happened?**

The duplex buffered-replay tests used a fixed 60 ms sleep before they
attached a data listener. Under load, worker output could arrive after
the sleep. The tests then saw a partial buffer and failed.

**Expected behavior**

The tests must wait until the worker sends all three data frames before
they inspect the pre-bind buffer.

**Steps to reproduce**

1. Run npx vitest run src/__tests__/plugin-worker-manager-duplex.test.ts
in the server package.
2. Add a 200 ms or 800 ms delay to the worker fixture emit path.
3. Repeat the test run and observe the old fixed-sleep barrier fail
intermittently.

**Paperclip version or commit**

b773f0f2e2

**Deployment mode**

Built from source. This change affects tests only.

## What Changed

- Replace the fixed sleep in both buffered-replay tests with an
exit-frame barrier.
- Write the three data frames and the exit frame in one worker output
write.
- Wait for the session to settle before the tests attach listeners.
- Keep the non-batch buffer-then-drain path and the throwing-listener
behavior.
- Remove the retry wrapper from the first test because the drain runs
synchronously.

## Verification

- Run npx vitest run src/__tests__/plugin-worker-manager-duplex.test.ts
in the server package.
- The full file passes 35 of 35 tests.
- Run the full file 15 times. All 15 runs pass.
- Test the new barrier with 200 ms and 800 ms worker-output delays. Both
tests pass.
- The server type check still reports 71 pre-existing errors in
native-runtime and paperclip-runner. No new error appears in the changed
test file.
- Search GitHub for duplicate or related public issues and pull
requests. No duplicate open item exists.
- Check ROADMAP.md. This test-only fix does not duplicate planned core
work.

## Risks

- This change affects test synchronization only.
- The test could become invalid if the worker stops sending the exit
frame. The session wait then fails instead of hiding the problem behind
a clock delay.
- No product code, database schema, or runtime behavior changes.

## Model Used

OpenAI GPT-5, exact model ID gpt-5, API model with code execution and
tool use. The model used a 1M-token context window. No extended
reasoning mode was specified.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the bug report template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 17:55:12 -07:00
dependabot[bot] 1f92011f99
chore(deps): bump dompurify from 3.4.13 to 3.4.14 (#12266)
Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.13 to
3.4.14.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/cure53/DOMPurify/releases">dompurify's
releases</a>.</em></p>
<blockquote>
<h2>DOMPurify 3.4.14</h2>
<ul>
<li>Fixed an issue with possible bypasses when risky tags are
allow-listed, thanks <a
href="https://github.com/AlirezaRouhbakhsh"><code>@​AlirezaRouhbakhsh</code></a></li>
<li>Fixed a couple of edge cases with mixed document contexts, thanks <a
href="https://github.com/fishjojo1"><code>@​fishjojo1</code></a></li>
<li>Added the SVG <code>pointer-events</code> and
<code>vector-effect</code> presentation attributes to the allow-list,
thanks <a
href="https://github.com/Jaybhade"><code>@​Jaybhade</code></a></li>
<li>Conducted another refactoring run, removed dead branches and
duplicated logic, flattened attribute validation</li>
<li>Updated the documentation in several spots, README, wiki, etc.,
thanks <a
href="https://github.com/Akokonunes"><code>@​Akokonunes</code></a></li>
<li>Updated several development dependencies and CI workflow
actions</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="4e6fe24173"><code>4e6fe24</code></a>
release: 3.4.14 (<a
href="https://redirect.github.com/cure53/DOMPurify/issues/1587">#1587</a>)</li>
<li>See full diff in <a
href="https://github.com/cure53/DOMPurify/compare/3.4.13...3.4.14">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 16:43:44 -07:00
dependabot[bot] daa2391021
chore(deps): bump @aws-sdk/client-s3 from 3.1120.0 to 3.1122.0 (#12261)
Bumps
[@aws-sdk/client-s3](https://github.com/aws/aws-sdk-js-v3/tree/HEAD/clients/client-s3)
from 3.1120.0 to 3.1122.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/releases">@​aws-sdk/client-s3's
releases</a>.</em></p>
<blockquote>
<h2>v3.1122.0</h2>
<h4>3.1122.0(2026-08-31)</h4>
<h5>Documentation Changes</h5>
<ul>
<li><strong>client-controltower:</strong> Updated the descriptions for
the AWS Control Tower ListEnabledControls API parameters to make them
more accurate and intuitive. (<a
href="c54ac4e601">c54ac4e6</a>)</li>
</ul>
<h5>New Features</h5>
<ul>
<li><strong>client-pinpoint-sms-voice-v2:</strong> AWS End User
Messaging SMS now returns ConditionalBehavior on
DescribeRegistrationFieldDefinitions, allowing you to programmatically
discover which registration fields are required, optional, or disallowed
based on the values of other fields in the same form. (<a
href="9cbace1398">9cbace13</a>)</li>
<li><strong>client-customer-profiles:</strong> This release introduces
new APIs for segment membership events allowing segment definition
membership events to be exported to a kinesis stream for downstream
processing. Additionally, includes new calculated attribute statistic
and 2 new segment dimension types. (<a
href="be1a9dab42">be1a9dab</a>)</li>
<li><strong>client-sagemaker:</strong> Amazon SageMaker Batch Transform
now supports G6e instances, powered by NVIDIA L40S Tensor Core GPUs. G6e
instances are the most cost-efficient GPU instances for deploying
generative AI models and the highest-performance GPU instances for
spatial computing workloads. (<a
href="b063cf77a9">b063cf77</a>)</li>
<li><strong>client-quicksight:</strong> This release adds support for
managing apps in Amazon QuickSight with ListApps, SearchApps,
DescribeApp, DescribeAppPermissions, UpdateAppPermissions, and DeleteApp
(<a
href="98a49570d5">98a49570</a>)</li>
<li><strong>client-connect:</strong> Added support for global routing on
Amazon Connect Global Resiliency instances. New APIs
GetCrossRegionRouting and UpdateCrossRegionRouting allow you to view and
control cross-region contact routing between linked instances, so both
Regions are active at all times. (<a
href="ce41026342">ce410263</a>)</li>
<li><strong>client-agent-registry-control:</strong> AWS Agent Registry
becomes Generally Available (<a
href="e41244e930">e41244e9</a>)</li>
<li><strong>client-kinesis:</strong> Adds support for data delivery to
Amazon S3 Tables (Apache Iceberg) and general purpose Amazon S3 buckets
with new CreateChannel, UpdateChannel, DeleteChannel, DescribeChannel,
and ListChannels APIs for Amazon Kinesis Data Streams. (<a
href="64ebb058e0">64ebb058</a>)</li>
<li><strong>client-agent-registry:</strong> AWS Agent Registry becomes
Generally Available (<a
href="e60306f198">e60306f1</a>)</li>
<li><strong>client-devops-agent:</strong> Adds support for Slack
bidirectional communication configuration in AWS DevOps Agent agent
spaces. (<a
href="75bc6d6da3">75bc6d6d</a>)</li>
<li><strong>client-kafkaconnect:</strong> Amazon MSK Connect now
supports restarting newly created connectors via the asynchronous
RestartConnector API. Restart all tasks or only failed tasks, while
preserving configuration and committed offsets. This returns a connector
operation ARN that you can track with DescribeConnectorOperation. (<a
href="8771afafd4">8771afaf</a>)</li>
<li><strong>client-support:</strong> AWS Support now allows up to 10
attachments (150 MB each) per case correspondence, up from 3 at 5 MB.
Customers can share large diagnostic logs, heap dumps, and packet
captures directly in cases to reduce back-and-forth and speed up
resolution. Available in US East, US West, and Europe (Ireland). (<a
href="4ddd79c106">4ddd79c1</a>)</li>
<li><strong>client-workspaces-instances:</strong> Amazon WorkSpaces Core
managed instances now support nested virtualization. Customers can
enable nested virtualization with supported instance types at launch via
CpuOptions.NestedVirtualization in CreateWorkspaceInstance to run
hypervisors and virtual machines inside their WorkSpaces Instance. (<a
href="29587d1236">29587d12</a>)</li>
</ul>
<hr />
<p>For list of updated packages, view
<strong>updated-packages.md</strong> in
<strong>assets-3.1122.0.zip</strong></p>
<h2>v3.1121.0</h2>
<h4>3.1121.0(2026-08-28)</h4>
<h5>New Features</h5>
<ul>
<li><strong>client-ecs:</strong> Amazon Elastic Container Service - This
release adds support for early success criteria on ECS rolling
deployments, letting deployment complete once a configurable percentage
of tasks are healthy, with configurable BLOCKING (required) or DEFERRED
(asynchronous) cleanup of previous service revisions. (<a
href="ef22d750f2">ef22d750</a>)</li>
<li><strong>client-healthlake:</strong> New HealthLake API,
RestoreFHIRDatastore, providing the capability to restore active
datastores to a point in time within the last 30 days or recover a
deleted datastore from the delete snapshot. (<a
href="6249174262">62491742</a>)</li>
<li><strong>client-bedrock-agentcore:</strong> AgentCore Memory now
supports direct ingestion into long-term memory via IngestData API (<a
href="20d652de56">20d652de</a>)</li>
<li><strong>client-partnercentral-selling:</strong> Releasing PARC, new
APN Program that lets sellers add solftware revenue details to aws
opportunity summary (<a
href="2b6350f012">2b6350f0</a>)</li>
<li><strong>client-cognito-identity-provider:</strong> Adds two new
operations - GetClientToken which allows M2M auth through the SDK, and
DescribeTermsByClient to find which Terms are associated with a
user-pool client without knowing the Terms resource id. (<a
href="86dffd282f">86dffd28</a>)</li>
<li><strong>client-bedrock-agent:</strong> Adds an optional syncSchedule
field to CreateDataSource and UpdateDataSource for Managed Knowledge
Bases data source connectors, so a data source can sync automatically on
a daily, weekly, or monthly schedule. (<a
href="a8d3714a75">a8d3714a</a>)</li>
</ul>
<hr />
<p>For list of updated packages, view
<strong>updated-packages.md</strong> in
<strong>assets-3.1121.0.zip</strong></p>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/blob/main/clients/client-s3/CHANGELOG.md">@​aws-sdk/client-s3's
changelog</a>.</em></p>
<blockquote>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1121.0...v3.1122.0">3.1122.0</a>
(2026-08-31)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1120.0...v3.1121.0">3.1121.0</a>
(2026-08-28)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="e1cf460a1e"><code>e1cf460</code></a>
Publish v3.1122.0</li>
<li><a
href="e53a25aafb"><code>e53a25a</code></a>
Publish v3.1121.0</li>
<li>See full diff in <a
href="https://github.com/aws/aws-sdk-js-v3/commits/v3.1122.0/clients/client-s3">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 15:16:31 -07:00
dependabot[bot] a0028d7e1b
chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
Bumps
[vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest)
from 4.1.10 to 4.1.11.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/vitest-dev/vitest/releases">vitest's
releases</a>.</em></p>
<blockquote>
<h2>v4.1.11</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li>Revive global concurrency limit for test lifecycle [backport to v4]
 -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> and
<a href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10992">vitest-dev/vitest#10992</a>
<a href="https://github.com/vitest-dev/vitest/commit/5146df80b"><!-- raw
HTML omitted -->(5146d)<!-- raw HTML omitted --></a></li>
<li><strong>browser</strong>:
<ul>
<li>Encode iframeId in tester iframe URL [backport to v4]  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a>,
<strong>Pduhard</strong> and <strong>Claude Opus 4.8</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10955">vitest-dev/vitest#10955</a>
<a href="https://github.com/vitest-dev/vitest/commit/10b2cd201"><!-- raw
HTML omitted -->(10b2c)<!-- raw HTML omitted --></a></li>
<li>Trigger playwright/chromium gc on lower disk availability [backport
to v4]  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>OpenCode</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10951">vitest-dev/vitest#10951</a>
<a href="https://github.com/vitest-dev/vitest/commit/9851dbc41"><!-- raw
HTML omitted -->(9851d)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>mocker</strong>:
<ul>
<li>Restrict redirect mocks to the fs allowlist [backport to v4]  -  by
<a href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a>
in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10974">vitest-dev/vitest#10974</a>
<a href="https://github.com/vitest-dev/vitest/commit/fe5a11d3c"><!-- raw
HTML omitted -->(fe5a1)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<h5>    <a
href="https://github.com/vitest-dev/vitest/compare/v4.1.10...v4.1.11">View
changes on GitHub</a></h5>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="9bd8d464e6"><code>9bd8d46</code></a>
chore: release v4.1.11 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10995">#10995</a>)</li>
<li><a
href="9851dbc41c"><code>9851dbc</code></a>
fix(browser): trigger playwright/chromium gc on lower disk availability
[back...</li>
<li>See full diff in <a
href="https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 15:16:02 -07:00
Dotta 5f87090894
Make managed Cloud OAuth handoffs invisible (#12790)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Apps let people give agents governed access to external providers
> - Paperclip Cloud brokers shared provider authorization for managed
stacks
> - The managed flow sent the browser through a confirmation page after
the tenant had already prepared sign-in
> - A lost confirmation response could also show an expired-session
error before the provider page opened
> - This pull request adds an opaque handoff contract and one shared
tenant coordinator
> - The benefit is a direct and recoverable transition from Paperclip to
every Cloud-brokered provider

## Linked Issues or Issue Description

**What happened?**

A managed Paperclip Cloud connection opened the Cloud confirmation
route. A response-loss race could show an expired-session error while
the authorization still continued.

**Expected behavior**

The current Paperclip loading state must stay visible while the tenant
exchanges an opaque session. The browser must then open the provider
directly. Self-hosted and direct OAuth must keep their existing
behavior.

**Steps to reproduce**

1. Open Apps on a Paperclip Cloud stack.
2. Start a managed provider connection.
3. Select Continue to sign in.
4. Observe that the browser visits the Cloud confirmation route before
it reaches the provider.

**Paperclip version or commit**

`b872cd3d1b404bdaff70af493a2973ceb7e5d6ec`

**Deployment mode**

Paperclip Cloud hosted stack.

No related open issue or pull request was found in the repository
search.

## What Changed

- Add a backward-compatible opaque Cloud handoff to the shared OAuth
start contract.
- Validate the Cloud descriptor on the server and expose no
browser-selected endpoint.
- Exchange managed handoffs through one fixed same-origin route in every
Apps OAuth launcher.
- Keep dialog popups reserved before asynchronous work and retain the
tenant loading state.
- Add recent-login resume storage, bounded retry behavior, terminal
tenant errors, tests, and Storybook states.

## Verification

- `pnpm check:token-gates`
- `pnpm -r typecheck`
- Focused connector and UI suites: 184 passed and 202 skipped.
- `pnpm build`
- `pnpm build-storybook`
- The full local suite reached one unrelated macOS path-alias failure.
The untouched test expected `/var/...` and received the equivalent
`/private/var/...`. The same test reproduces in isolation.

## Risks

- A malformed managed descriptor now fails closed in Paperclip instead
of opening a URL.
- A legacy Cloud deployment can omit the descriptor. Paperclip then uses
the existing validated confirmation URL.
- Direct provider OAuth and self-hosted flows do not receive a handoff
and remain unchanged.
- Rollback is a normal revert of this commit because the contract is
optional and backward compatible.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with GPT-5.6, reasoning mode, tool use, code execution,
and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 16:33:13 -05:00
Nicky Leach 66ea41812d
test(server): make the instance settings route suite deterministic under CPU contention (#12789)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip tests its server routes with mocked services and database
calls
> - The instance settings route suite reset and reloaded its module
graph before each test
> - Two concurrent module imports could bind a route to the real service
module under CPU contention
> - The task-drain overlap test also relied on a fixed delay and
operating system request order
> - This pull request loads the mocked graph once and waits for real
events that prove request order
> - The benefit is a deterministic 48-test suite with no production code
change

## Linked Issues or Issue Description

**What happened?**

The instance settings route suite failed intermittently under CPU
contention. A request that expected a 200 or 403 response sometimes
received 500. The failing test changed between runs.

**Expected behavior**

The suite must use the configured service mocks for every test and must
produce the expected response on every run.

**Steps to reproduce**

1. Run `npx vitest run
server/src/__tests__/instance-settings-routes.test.ts` many times in
parallel on a busy host.
2. Compare the result with the same command on the base branch.
3. Observe intermittent 500 responses on the base branch and stable
results on this branch.

**Paperclip version or commit**

Commit `02ae87010e621cf46bfbdf0d48b6f73887448a83`.

**Deployment mode**

Local dev (`pnpm dev`). The change affects tests only.

**Installation method**

Built from source.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Database mode**

Not database-related. The test uses a mocked database layer.

**Access context**

Not applicable.

**Node.js version**

The CI environment runs the repository-supported Node.js version.

**Operating system**

Linux in continuous integration.

**Relevant logs or output**

The base branch reproduced `expected 500 to be 200` and `expected 500 to
be 403` under parallel contention.

**Relevant config (if applicable)**

Not applicable.

**Additional context**

The branch loads the mocked module graph once per suite, restores mock
behavior before each test, waits for the real transaction events, and
sends the DELETE request after the POST holds the transition queue.

## What Changed

- Load the mocked instance settings module graph once for the suite.
- Restore each mock implementation before every test.
- Wait for two real transaction events instead of a fixed 30 millisecond
delay.
- Send the overlapping DELETE request after the POST proves that it
holds the transition queue.
- Keep the test count at 48 with no skipped tests.

## Verification

- Run `npx vitest run
server/src/__tests__/instance-settings-routes.test.ts`.
- Confirm that all 48 tests pass.
- Run the 20-way parallel contention differential.
- Confirm that the base arm passed 18 of 20 runs and reproduced two
failures.
- Confirm that the branch arm passed 20 of 20 runs, with 48 tests in
each run.
- Confirm that `git status --porcelain` is clean at the submitted
commit.

## Risks

Low risk. The change affects one test file and does not change
production code, route behavior, database schema, or public API
behavior.

## Model Used

OpenAI Codex, GPT-5, current deployment, tool use and code execution
enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 14:23:07 -07:00
Nicky Leach b872cd3d1b
test(server): select exposure reservation host ports at run time (#12783)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server manages runtime exposure and host port leases for
workspace services
> - This test suite used fixed host port pairs inside the Linux
ephemeral port range
> - An unrelated short-lived socket could take one pair and cause a
false test failure
> - This pull request selects free host port pairs at run time and
starts above the low lease lane
> - The benefit is a more stable test suite with the same deterministic
allocator checks

## Linked Issues or Issue Description

**What happened?**

The runtime exposure reservation test suite used two fixed app and HMR
port pairs. These ports sit inside the Linux ephemeral port range. An
unrelated socket could use a pair during the test, and the guest bind
could fail with `EADDRINUSE`.

**Expected behavior**

The suite must select two free app and HMR port pairs before each test.
It must avoid the low lease lane that a live instance can own without a
listener.

**Steps to reproduce**

1. Run `npx vitest run
server/src/__tests__/workspace-runtime-exposure-reservation.test.ts`.
2. Start another process that briefly uses one fixed test port.
3. Observe that the guest bind can fail even when the allocator works
correctly.

**Paperclip version or commit**

`c982003e00f4e8a325bafec3af4ddb113c0c1f8a`

**Deployment mode**

Local dev test run.

**Installation method**

Built from source with pnpm.

**Agent adapter(s) involved**

Not adapter-specific. This change tests the runtime exposure allocator.

**Database mode**

Not database-related.

## What Changed

- Select two free app and HMR port pairs in `beforeEach`.
- Start the scan 500 ports above the runtime exposure range minimum.
- Keep the synthetic host stub limited to the selected pairs.
- Keep all seven test cases and the existing lifecycle coverage.

## Verification

- Run `npx vitest run
server/src/__tests__/workspace-runtime-exposure-reservation.test.ts`.
- Run `npx vitest run
server/src/services/workspace-runtime-exposure.test.ts`.
- Run `pnpm --filter @paperclipai/server exec tsc --noEmit`.
- Confirm the full CI suite reaches a terminal green state.

## Risks

Low risk. This change updates one test file and does not change
production code. A port can still become busy after discovery and before
the guest bind; the test documents this remaining race.

## Model Used

OpenAI GPT-5 (`gpt-5`), tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 13:18:37 -07:00
Nicky Leach 0cf06c8fa1
test(server): make secret write-serialization tests deterministic (#12781)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip stores and controls secrets through server services
> - The secret service tests check that concurrent writes use one lock
at a time
> - Fixed sleep times do not prove that a provider write started or
stayed queued
> - This pull request uses provider-write signals and measured waits to
test lock behavior
> - The benefit is stable test results and stronger detection of lock
failures

## Linked Issues or Issue Description

**What happened?**

The secret write-serialization tests used fixed 20 ms sleeps. The sleeps
sometimes ran before a provider write or after a queued write entered.
The tests then failed or missed a broken lock.

**Expected behavior**

The tests must wait for real provider-write events and must detect a
queued write that enters before the first write finishes.

**Steps to reproduce**

1. Run `npx vitest run server/src/__tests__/secrets-service.test.ts`.
2. Repeat the test file under sustained load.
3. Remove the write lock and run the concurrency tests.
4. Observe intermittent timing failures or missed lock failures.

**Paperclip version or commit**

`13bff0adee0216ee9ec67c843e9ead94aa788c68`

**Deployment mode**

Local dev (`pnpm dev`)

**Installation method**

Built from source (`pnpm dev` / `pnpm build`)

**Agent adapter(s) involved**

Not adapter-specific (core test)

**Database mode**

Not database-related

**Additional context**

This pull request changes tests only. It does not change production
code.

## What Changed

- Wait for a deferred signal when the first operation reaches its
provider write.
- Measure an uncontended provider-write duration and use a safety
multiple for the queued-write check.
- Release the test gate in a `finally` block so failed assertions do not
leave a write active.
- Throw when the measurement helper does not observe the provider write.

## Verification

- `npx tsc --noEmit -p server/tsconfig.json` reports no errors in the
changed file.
- `npx vitest run server/src/__tests__/secrets-service.test.ts` passes
90 of 90 tests.
- The engineer ran the test file five times under sustained load, and
all runs passed.
- Full CI must pass after this pull request starts.

## Risks

Low risk. The change affects test code only. The measured wait can
expose a real lock regression, but it does not change runtime behavior.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution. The exact context
window and reasoning mode are not exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 12:54:42 -07:00
Devin Foley db4eeb1688
fix(server): validate project goal ids exist and belong to the company (#12779)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Projects can link to goals, through the `goalIds` list or the legacy
`goalId` field. The project service writes those links on create and
update.
> - The service never checked the goal ids. A nonexistent id died at the
`projects.goal_id` foreign key as an opaque 500, and the caller got no
actionable feedback — observed live on 2026-09-03, where one caller
retried the same bad id four times.
> - The foreign key also only proves a goal exists, not who owns it. A
goal id from another company linked silently on a multi-company
instance.
> - This pull request asserts every resolved goal id exists under the
caller's company before any write, and rejects with a 422 that names the
unknown ids.
> - The benefit is a clear, actionable client error instead of a 500,
and no cross-company goal links.

## Linked Issues or Issue Description

**What happened?**

`POST /companies/:companyId/projects` with a `goalIds` entry that does
not exist fails with an internal error: `insert or update on table
"projects" violates foreign key constraint
"projects_goal_id_goals_id_fk"`. The caller sees a 500 and retries. A
goal id that exists but belongs to a different company is accepted and
linked.

**Expected behavior**

The request fails fast with a 422 that names the unknown goal id(s).
Goals from other companies are rejected the same way. Valid links behave
exactly as before.

**Steps to reproduce**

1. Create a company and no goals.
2. `POST /companies/:companyId/projects` with `{ "name": "Rocket",
"goalIds": ["<any-uuid>"] }`.
3. Before this change: 500 from the foreign key. After: 422 naming the
id.

**Deployment mode**

Any; observed on an authenticated public deployment.

## What Changed

- `assertGoalsBelongToCompany` in the project service: one query for the
resolved ids scoped to the company; unknown ids produce `unprocessable`
(422) with the ids in the message and details
- called on create (before the project row insert, so no partial writes)
and on update (scoped to the existing project's company); both `goalIds`
and the legacy `goalId` field flow through the same resolution
- new embedded-Postgres test file: valid link, nonexistent id on create
with no partial insert, legacy field, another company's goal on create,
and a foreign-goal update that leaves existing links unchanged

## Verification

- `pnpm vitest run src/__tests__/project-goal-validation.test.ts` — 5
passed
- adjacent suites (`project-icon-persistence`,
`project-shortname-resolution`, `issue-goal-fallback`,
`project-goal-telemetry-routes`, `heartbeat-referenced-projects`,
`projects-list-archived-routes`) — 35 passed

## Risks

- Low risk. One extra indexed select per create/update that carries goal
ids. Requests that previously 500ed now 422; requests that silently
linked a foreign goal now fail — both are corrections, not regressions.
- Existing rows with foreign links (written before this check) are
untouched; only new writes validate.

## Model Used

Claude Fable 5 (claude-fable-5) via Claude Code, extended thinking with
tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found for goal-id validation)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (doc
comments; no user-facing docs affected)
- [x] I have considered and documented any risks above
2026-09-03 12:17:27 -07:00
Devin Foley 2177b85eb5
fix(server): retry cloud-tenant auth sync once on a dropped DB connection (#12773)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed-cloud deployments authenticate tenant requests through
trusted headers. The middleware syncs the tenant's user, company, and
membership rows on the way through.
> - Pooled Postgres endpoints sometimes close an established connection
under an in-flight query (pooler recycle, compute suspend). The driver
reconnects on the next query, but the statement on the wire fails.
> - In this path a single dropped statement fails the whole request with
a 500. This happened live on 2026-09-03: the idempotent company
bootstrap insert died with `write CONNECTION_CLOSED`.
> - This pull request retries the actor resolution exactly once when the
error chain carries a postgres.js closed-connection code. The sync is
idempotent end to end, so the replay is safe.
> - The benefit is that a routine pooler blip no longer fails an
authenticated request on the entry path.

## Linked Issues or Issue Description

**What happened?**

A cloud tenant request hit the trusted-header authentication middleware
while the pooled Postgres endpoint closed the connection mid-query. The
insert failed with `write CONNECTION_CLOSED <host>:5432` wrapped in a
`Failed query: insert into "companies" …` error, and the request failed.

**Expected behavior**

The driver reconnects on the next query, and every statement in the
tenant sync is idempotent (upserts, on-conflict inserts, deletes; the
write debounce records only after the full sync succeeds). One
in-request retry should absorb the blip and serve the request.
Non-transient failures must keep failing fast.

**Steps to reproduce**

1. Run an authenticated public deployment against a pooled Postgres
endpoint.
2. Have the pooler close the connection while the middleware's tenant
sync insert is on the wire.
3. Before this change the request fails with a 500; after it the retry
serves the request.

**Deployment mode**

Authenticated public (managed cloud), external pooled PostgreSQL.

## What Changed

- `resolveCloudTenantActor` now delegates to the (unchanged) resolution
body through `retryOnTransientDbConnectionError`, which retries exactly
once on a transient closed-connection failure
- `isTransientDbConnectionError` walks the error `cause` chain (drizzle
wraps the driver error) for the postgres.js codes `CONNECTION_CLOSED`,
`CONNECTION_ENDED`, `CONNECTION_DESTROYED`; both helpers are exported
for tests
- New unit test file `cloud-tenant-transient-db-retry.test.ts`:
detection matrix (including a `23505` staying non-transient),
retry-once-then-succeed, no-retry on non-transient,
propagate-on-second-failure

## Verification

- `pnpm vitest run
src/__tests__/cloud-tenant-transient-db-retry.test.ts` — 5 passed
- `pnpm vitest run
src/__tests__/cloud-tenant-company-provisioning.test.ts` — 7 passed
against embedded Postgres, driving the real resolution path through the
new wrapper

## Risks

- Low risk. The retry is bounded to one attempt, gated on three explicit
driver codes, and wraps an operation that is already idempotent by
design. Every other failure propagates unchanged.
- A genuinely down database now fails after two attempts instead of one
— a few milliseconds of added latency on an already-failing request.

## Model Used

Claude Fable 5 (claude-fable-5) via Claude Code, extended thinking with
tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found for connection-retry work in this path)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (doc
comments; no user-facing docs affected)
- [x] I have considered and documented any risks above
2026-09-03 12:16:55 -07:00
Dotta 9dd6526b47
fix(security): harden privileged server boundaries (#12776)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server controls secrets, host files, outbound requests, and
workspace commands
> - A red-team review found cases where restricted callers could cross
these trust boundaries
> - These cases could expose credentials or let untrusted input reach
privileged resources
> - This pull request applies least-privilege checks at each affected
server boundary
> - The benefit is safer agent execution without changing the
private-instance bootstrap contract

## Linked Issues or Issue Description

**What happened?**

Several server paths used authorization, redaction, or content-delivery
rules that were too broad. Restricted agent keys could obtain
company-level operational data. Some adapter and instruction paths could
reach server-owned network or file resources without the required owner
approval.

**Expected behavior**

Paperclip must redact credential values, enforce restricted-key scopes,
guard outbound network access, prevent same-origin script execution, and
reserve host-level file and command controls for authorized operators.

**Steps to reproduce**

1. Configure an authenticated development instance at the parent commit.
2. Exercise the affected APIs with a restricted agent key or a
non-instance-admin company user.
3. Observe that the parent commit returns privileged data or accepts a
privileged operation.
4. Repeat on this branch and observe a redacted response, a safe
download, or an HTTP 403 response.

**Paperclip version or commit**

The findings reproduce from commit `39898ab22` and are fixed by this
pull request.

**Deployment mode**

Authenticated self-hosted server and local development modes.

**Installation method**

Built from source with pnpm.

## What Changed

- Redact generic secret `value` and `token` fields recursively in
structured logs.
- Classify exact and separator-suffixed `KEY` environment names as
secrets in company exports.
- Limit restricted self-identity responses and protect company run, log,
and secret catalog APIs.
- Route HTTP adapter requests through DNS-pinned SSRF protection with
exact private-origin allowlisting.
- Download HTML, SVG, and other script-capable assets with `nosniff` and
a sandbox CSP.
- Require instance-admin access for external instruction roots and
exports that read them.
- Block agent-authenticated host command persistence across supported
workspace runtime shapes.
- Apply the central runtime-management decision before workspace command
controls.
- Keep the documented first-user instance-admin claim contract
unchanged.
- Add regression tests and server-owner configuration documentation.

## Verification

- `pnpm -r typecheck` passes.
- The Node 24 remediation suite passes with 365 tests. It skips 25
environment-gated tests.
- `pnpm build` passes under Node 24.
- `git diff --check` passes.
- The full local runner reaches known macOS-only general-server harness
failures before the serialized route lane. The Linux PR matrix is the
authoritative full-suite gate.

## Risks

- Restricted agent keys now receive HTTP 403 responses from company-wide
run, log, and secret catalog endpoints.
- Script-capable assets now download instead of rendering inline.
- External instruction roots now require instance-admin access.
- Private HTTP adapter endpoints now require an exact origin in
`PAPERCLIP_HTTP_ADAPTER_PRIVATE_ENDPOINT_ALLOWLIST`.
- Public HTTP adapter endpoints remain enabled. Redirects and metadata
or link-local targets remain blocked.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5. The exact serving snapshot and context-window size
are not exposed. The model used tool-enabled reasoning, repository
access, code execution, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 14:15:32 -05:00
Paperclip DevOps Engineer 31a63638ac
fix(agents): redact plaintext env values in agent read and mutation responses (#9860)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents are configured through `adapterConfig`, whose `env` block
holds the credentials an agent needs to talk to its provider (API keys,
tokens, and similar)
> - Those bindings come in several shapes: a legacy bare string, `{
type: "plain", value }`, and the indirection forms `{ type: "secret_ref"
}` / `{ type: "user_secret_ref" }`
> - Every endpoint that serializes an agent returned `adapterConfig` as
stored, so every `plain` binding was returned verbatim in the API
response
> - That means any caller able to read an agent — including the agent
itself via `GET /api/agents/me` — received live credentials in
plaintext, and those values then propagate into client state, logs, and
network traces
> - The exposure spans three response families that share no common
serializer: the single-agent detail reads, the company agent-list read,
and the create/update/lifecycle routes that echo the stored row straight
back
> - This pull request routes all three through one presenter that
redacts `plain` env bindings, so the leak is closed server-side and
cannot be bypassed by the caller
> - The benefit is that agent credentials stop appearing in API
responses while `secret_ref` indirection continues to work unchanged

## Linked Issues or Issue Description

No public upstream issue exists for this, so the underlying bug is
described inline below following
`.github/ISSUE_TEMPLATE/bug_report.yml`.

**What happened?**

Every endpoint that serializes an agent returned the full plaintext
value of each `adapterConfig.env` entry whose `type` was `"plain"` (and
each legacy bare-string binding). Any actor authorized to read an agent
received that agent's live credentials in the response body. Three
distinct response families were affected:

- **Single-agent reads** — `GET /api/agents/{id}` and `GET
/api/agents/me`, via `buildAgentDetail`.
- **Company agent list** — `GET /api/companies/{companyId}/agents`,
which serializes rows directly and therefore does not inherit any fix
applied to `buildAgentDetail`. Callers that pass the configuration-read
check received unredacted rows for every agent in the company in a
single request, making this the broadest of the three.
- **Mutation responses** — agent create, `PATCH /api/agents/{id}`, and
the `pause` / `resume` / `clear-error` / `approve` / `terminate` routes,
each of which echoes the stored row back to the caller.

**Expected behavior**

Read endpoints should never emit stored plaintext credentials. `plain`
bindings should be replaced with a redaction sentinel before
serialization, while `secret_ref` and `user_secret_ref` bindings — which
contain no secret material — pass through untouched.

**Steps to reproduce**

1. Configure an agent with an `adapterConfig.env` entry such as `{
"OPENAI_API_KEY": { "type": "plain", "value": "sk-example" } }`.
2. Call `GET /api/agents/{id}` (or authenticate as that agent and call
`GET /api/agents/me`).
3. Observe `sk-example` returned verbatim in the response body.
4. Call `GET /api/companies/{companyId}/agents` as a
configuration-reading caller and observe `sk-example` returned verbatim
for that agent alongside every other agent's credentials.
5. Call `PATCH /api/agents/{id}` with any unrelated field (for example
`{ "title": "Renamed" }`) and observe `sk-example` returned verbatim in
the mutation response.

**Paperclip version or commit**

Reproduced on `master` at `f12bb27b`.

**Deployment mode**

Self-hosted / local development server.

## What Changed

- `server/src/redaction.ts`: adds `redactAgentAdapterConfig`, which
rewrites every bare-string or `{ type: "plain", value }` env binding to
`{ type: "plain", value: "***REDACTED***" }` and passes `secret_ref` /
`user_secret_ref` bindings through unchanged. Reuses the existing
`REDACTED_EVENT_VALUE` and `isSecretRefBinding` /
`isUserSecretRefBinding` / `isPlainBinding` helpers — no new
dependencies.
- `server/src/redaction.ts`: `env` is destructured out and sanitized
only by `redactAgentEnvBinding`, while the remaining adapter keys go
through `redactEventPayload`. Previously the already-redacted `env` was
passed back through `sanitizeRecord`, so each binding was processed
twice — safe only because the sentinel is a fixed point of that second
pass. The two paths are now disjoint, making the invariant structural
rather than coincidental.
- `server/src/routes/agents.ts`: `buildAgentDetail` applies
`redactAgentAdapterConfig` before serialization, so `GET
/api/agents/{id}` and `GET /api/agents/me` both redact at the response
layer. Restricted views inherit the same protection.
- `server/src/routes/agents.ts`: adds `redactAgentRowForResponse`, the
single presenter for every response that emits a raw agent row, and
applies it to the company agent-list route and to the create / update /
pause / resume / clear-error / approve / terminate routes. It composes
with `redactForRestrictedAgentView` rather than replacing it: that
helper is an authorization filter (blank the whole config for low-trust
actors), this one is secret hygiene (mask values for every actor scope),
and the two invariants stay independent. `buildAgentDetail` now
delegates to the same presenter instead of inlining the call.
- `server/src/routes/agents.ts`: adds `restoreRedactedAgentEnv` on the
PATCH path so a client that round-trips a redacted detail response back
through `PATCH /api/agents/{id}` does not zero out stored values —
redacted-sentinel entries matching an existing key are restored from
storage.

## Verification

- `pnpm --filter @paperclipai/server exec tsc --noEmit` — clean.
- `pnpm --filter @paperclipai/server exec vitest run
agent-permissions-routes.test.ts` — 57 tests pass.
- Adjacent suites (`redaction`, `agent-adapter-validation-routes`,
`agent-cross-tenant-authz-routes`, `agents-pending-approval-config`,
`agents-service-secret-bindings`, `built-in-agent-routes`,
`plugin-managed-agents`, `agent-skills-routes`) — 8 files, 72 tests
pass, no regressions.
- Both new route tests were confirmed to **fail** with the route changes
reverted and pass with them applied, so they genuinely pin the behaviour
rather than passing incidentally.

Tests added:

- `server/src/__tests__/redaction.test.ts`: covers legacy-string, `{
type: "plain" }`, `secret_ref`, and `user_secret_ref` bindings,
asserting the plaintext value never appears in the serialized result;
plus coverage that non-env adapter keys are still sanitized, that env
binding shapes survive intact, and that configs with no `env` block are
handled.
- `server/src/__tests__/agent-permissions-routes.test.ts`: `GET
/api/agents/{id}` asserts redaction rather than plaintext passthrough;
new `GET /api/agents/me` redaction test across the same binding shapes;
new test asserting the `PATCH` round-trip preserves stored values; new
test asserting the board `GET /api/companies/{companyId}/agents`
response redacts every binding shape; new test asserting a mutation
response redacts rather than echoing the stored plaintext.

No real secret values appear in any test, fixture, or commit message.

## Risks

- **Behavioral change for API consumers.** Any client that read a
plaintext credential out of an agent detail, agent-list, or mutation
response will now receive `***REDACTED***`. This is the intended
security fix, but it is a breaking change for such consumers, which must
move to `secret_ref` indirection.
- **Mutation responses are redacted too.** Callers that previously
relied on a create or update response to echo back the credential they
had just written must now read it from their own request. This is
consistent with the `restoreRedactedAgentEnv` round-trip path, which
already assumes the client holds a redacted copy.
- **Round-trip data loss, mitigated.** A client that GETs an agent and
PATCHes the object straight back would otherwise persist the sentinel
over the real value. `restoreRedactedAgentEnv` restores redacted entries
from storage; the round-trip is covered by a regression test. A PATCH
that *intentionally* sets a value literally equal to the sentinel is not
distinguishable and would be treated as "unchanged" — an acceptable
trade-off given the sentinel is not a plausible credential.
- **No migration.** Stored data is untouched; redaction happens purely
at serialization time, so the change is fully reversible by revert.
- **Overlap with existing PRs** — see the duplicate-search note below.
Maintainers may prefer to consolidate rather than merge this in
isolation.

## Model Used

Claude Opus 4.8 (`claude-opus-4-8`), extended thinking enabled, with
tool use and local test execution.

## Duplicate search

Searching open and closed PRs for prior art surfaced several overlapping
efforts against the same defect. Linking them for maintainer triage — I
am not claiming this PR supersedes them, and consolidation may well be
preferable:

- #9823 — `fix(security): redact adapterConfig secrets on all agent read
endpoints` (closest overlap)
- #8779 — `fix(server): redact agent config secrets in read and mutation
responses`
- #8330 — `fix(server): redact adapterConfig.env for cross-actor agent
reads`
- #4856 — `fix(server): redact adapter env secrets in agent API
responses`
- #4763 — `fix(server): redact adapter_config secrets in agent detail
responses`
- #1839 — `fix: redact secret env vars from agent API responses`
- #4967 — `fix(routines): redact adapterConfig.env in GET
/api/routines/{id}` (same class, routines surface)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I searched the GitHub PR list (open and closed) for similar or
duplicate PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
Paperclip ticket id — **not met**: the branch and title carry an
internal ticket id. Renaming the branch would invalidate this PR; happy
to reopen from a clean branch if maintainers prefer.
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
user-facing docs affected)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
the one P2 (env entries processed twice) is addressed above
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Matthew Glover <5413384+glovario@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
2026-09-03 14:13:26 -05:00
dependabot[bot] 6826452856
chore(deps): bump sharp from 0.35.3 to 0.35.4 (#12563)
Bumps [sharp](https://github.com/lovell/sharp) from 0.35.3 to 0.35.4.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/lovell/sharp/releases">sharp's
releases</a>.</em></p>
<blockquote>
<h2>v0.35.4</h2>
<p><a
href="https://github.com/lovell/sharp-libvips/releases/tag/v1.3.3">https://github.com/lovell/sharp-libvips/releases/tag/v1.3.3</a></p>
<ul>
<li>
<p>Bound resize dimensions to coordinate limit.</p>
</li>
<li>
<p>Bound composite left and top to coordinate limit.
<a href="https://redirect.github.com/lovell/sharp/pull/4564">#4564</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Round palette bit depth up for png and gif colours.
<a href="https://redirect.github.com/lovell/sharp/pull/4569">#4569</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Ensure tiff.subifd input option is used.
<a href="https://redirect.github.com/lovell/sharp/pull/4572">#4572</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Ensure <code>info.pages</code> is correct when limiting input page
range.
<a href="https://redirect.github.com/lovell/sharp/pull/4578">#4578</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Improve support for input Streams finishing before output is
requested.
<a href="https://redirect.github.com/lovell/sharp/pull/4584">#4584</a>
<a href="https://github.com/Jaybhade"><code>@​Jaybhade</code></a></p>
</li>
</ul>
<h2>v0.35.4-rc.0</h2>
<ul>
<li>
<p>Upgrade to libvips v8.18.6 for upstream bug fixes.</p>
</li>
<li>
<p>Bound resize dimensions to coordinate limit.</p>
</li>
<li>
<p>Bound composite left and top to coordinate limit.
<a href="https://redirect.github.com/lovell/sharp/pull/4564">#4564</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Round palette bit depth up for png and gif colours.
<a href="https://redirect.github.com/lovell/sharp/pull/4569">#4569</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Ensure tiff.subifd input option is used.
<a href="https://redirect.github.com/lovell/sharp/pull/4572">#4572</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Ensure <code>info.pages</code> is correct when limiting input page
range.
<a href="https://redirect.github.com/lovell/sharp/pull/4578">#4578</a>
<a
href="https://github.com/metsw24-max"><code>@​metsw24-max</code></a></p>
</li>
<li>
<p>Improve support for input Streams finishing before output is
requested.
<a href="https://redirect.github.com/lovell/sharp/pull/4584">#4584</a>
<a href="https://github.com/Jaybhade"><code>@​Jaybhade</code></a></p>
</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="7f1a0a22cc"><code>7f1a0a2</code></a>
Release v0.35.4</li>
<li><a
href="f927818924"><code>f927818</code></a>
Upgrade to sharp-libvips v1.3.3</li>
<li><a
href="e80209240d"><code>e802092</code></a>
Prerelease v0.35.4-rc.0</li>
<li><a
href="e13eb2f97a"><code>e13eb2f</code></a>
CI: Fix wasm32 build (<a
href="https://redirect.github.com/lovell/sharp/issues/4589">#4589</a>)</li>
<li><a
href="a82a0b3d58"><code>a82a0b3</code></a>
Upgrade to libvips v8.18.6</li>
<li><a
href="8044fe43e3"><code>8044fe4</code></a>
Bound resize dimensions to coordinate limit</li>
<li><a
href="147f8591a1"><code>147f859</code></a>
Docs: changelog entries for <a
href="https://redirect.github.com/lovell/sharp/issues/4578">#4578</a> <a
href="https://redirect.github.com/lovell/sharp/issues/4584">#4584</a></li>
<li><a
href="ee5bfb853d"><code>ee5bfb8</code></a>
Tests: use yauzl directly rather than via extract-zip wrapper</li>
<li><a
href="7a7788928f"><code>7a77889</code></a>
Bump uraimo/run-on-arch-action from 3.1.0 to 3.2.0 (<a
href="https://redirect.github.com/lovell/sharp/issues/4588">#4588</a>)</li>
<li><a
href="ea5bef24c1"><code>ea5bef2</code></a>
Improve support for input Streams finishing before output is requested
(<a
href="https://redirect.github.com/lovell/sharp/issues/4584">#4584</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/lovell/sharp/compare/v0.35.3...v0.35.4">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 10:16:58 -07:00
Devin Foley 174e35a144
fix(server): stop paging Sentry for supervised boot races in managed cloud (#12772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server refuses to boot when its database is not migrated, or
when an authenticated public deployment has no `DATABASE_URL`. These
refusals are deliberate and correct.
> - In managed cloud, a supervisor creates each stack, migrates its
fresh database, applies configuration, and restarts the app. The app
container often boots before those steps finish.
> - Each early boot hits one of the two refusals, exits, and captures
the refusal to Sentry. One fleet build batch produces hundreds of
identical expected events. Real errors get buried.
> - This pull request classifies exactly those two refusals as expected
transients when `PAPERCLIP_CLOUD_API_ORIGIN` marks a supervised
deployment, and skips only the Sentry capture for them.
> - The benefit is a clean error signal: expected provisioning noise
stops, and every real failure still reports.

## Linked Issues or Issue Description

**What happened?**

A managed-cloud stack boots its app container before the supervisor
migrates the empty database or finishes applying configuration. The
container refuses to start, crash-loops briefly, and converges after the
supervisor restarts it. Every refused boot sends an error event to
Sentry. A batch of new stacks produces hundreds of these expected
events.

**Expected behavior**

The refusal logs and exits nonzero, so the supervisor can act. Sentry
receives no event for an expected provisioning transient. Sentry still
receives events for real failures: schema drift, malformed
configuration, and every refusal outside managed cloud.

**Steps to reproduce**

1. Set `PAPERCLIP_MIGRATION_AUTO_APPLY=false`,
`PAPERCLIP_MIGRATION_PROMPT=never`, `SENTRY_DSN`, and
`PAPERCLIP_CLOUD_API_ORIGIN`.
2. Point `DATABASE_URL` at an empty database and start the server.
3. The server refuses to start. Before this change it also captures the
refusal to Sentry on every boot.

**Deployment mode**

Authenticated public (managed cloud).

## What Changed

- New `server/src/startup-refusals.ts`: a `StartupRefusalError` class
for refusals whose remedy belongs to the deployment supervisor,
`migrationRefusalError()` to classify a pending-migrations refusal (zero
applied migrations = never migrated = supervised transient; any applied
history = drift = plain always-reported `Error`), and
`shouldReportStartupFailure()` for the capture decision.
- `server/src/index.ts`: the pending-migrations refusal uses the
classifier; the missing-`DATABASE_URL` refusal under the
authenticated-public contract becomes a `StartupRefusalError` (the
malformed-URL refusal stays a plain `Error`); the startup crash handler
consults `shouldReportStartupFailure()` before `captureException`.
Logging and the nonzero exit are unchanged.
- New `server/src/__tests__/startup-refusals.test.ts` covering the
classification and decision matrix, including the unchanged self-hosted
paths.

## Verification

- `pnpm vitest run src/__tests__/startup-refusals.test.ts` — 7 passed.
- Review the decision matrix in the test file: refusals report when
`PAPERCLIP_CLOUD_API_ORIGIN` is absent or blank; non-refusal errors and
non-`Error` throwables always report; drift always reports.

## Risks

- Low risk. The change only skips a Sentry capture in one narrow,
marker-gated case. Boot behavior, logging, and the exit code do not
change.
- Self-hosted deployments do not set `PAPERCLIP_CLOUD_API_ORIGIN`, so
their reporting is unchanged, and the tests pin that.
- A supervised deployment with a genuinely stuck migration runner loses
per-boot Sentry events for that stack. The supervisor's own health
checks and monitoring own that signal, and the container logs still
carry the refusal.

## Model Used

Claude Fable 5 (claude-fable-5) via Claude Code, extended thinking with
tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found for startup Sentry suppression)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(module doc comment; no user-facing docs affected)
- [x] I have considered and documented any risks above
2026-09-03 10:02:42 -07:00
Dotta 0798c77fde
Secure Cloud canonical runtime identity (#12766)
Accept and persist Cloud-signed canonical runtime identity before activation, then route absolute self-URLs through the durable runtime identity provider.

Co-Authored-By: Codex <codex@openai.com>
2026-09-03 12:01:47 -05:00
dependabot[bot] 2e8521e57c
chore(deps): bump better-auth from 1.7.0 to 1.7.2 (#12565)
Bumps
[better-auth](https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth)
from 1.7.0 to 1.7.2.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/better-auth/better-auth/releases">better-auth's
releases</a>.</em></p>
<blockquote>
<h2>v1.7.2</h2>
<h2><code>better-auth</code></h2>
<h3>Bug Fixes</h3>
<ul>
<li>Fixed permanent user bans to clear expiration dates from previous
temporary bans. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10823">#10823</a>)</li>
<li>Fixed client types with more plugins being assignable to types
declaring fewer plugins. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10907">#10907</a>)</li>
<li>Added warnings for invalid signed session data in the cookie cache.
(<a
href="https://redirect.github.com/better-auth/better-auth/pull/10934">#10934</a>)</li>
<li>Fixed disabled MyISAM indexes from satisfying migration index
checks. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10877">#10877</a>)</li>
<li>Fixed programmatic migrations on Cloudflare D1 while preserving
existing-index validation. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10875">#10875</a>)</li>
<li>Allowed <code>~</code> in relative callback URLs validated by
trusted-origin checks. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10041">#10041</a>)</li>
<li>Improved validation of relative callback and redirect URLs with
paths, queries, and fragments. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10979">#10979</a>)</li>
<li>Allowed same-origin form submissions with <code>Referrer-Policy:
no-referrer</code> while continuing to reject untrusted origins. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10959">#10959</a>)</li>
<li>Improved <code>getTestInstance</code> performance with a faster
default password hasher. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10879">#10879</a>)</li>
<li>Standardized built-in placeholder emails to the namespaced
<code>{identifier}@{namespace}.placeholder.invalid</code> format. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10982">#10982</a>)</li>
</ul>
<p>For detailed changes, see <a
href="c50200bfc7/packages/better-auth/CHANGELOG.md"><code>CHANGELOG</code></a></p>
<h2><code>@better-auth/core</code></h2>
<h3>Bug Fixes</h3>
<ul>
<li>Fixed async context loss in Cloudflare Workers bundles with multiple
runtime conditions. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10855">#10855</a>)</li>
<li>Fixed auth request logs to respect the configured logger, log level,
and disabled setting. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10939">#10939</a>)</li>
<li>Improved validation of relative callback and redirect URLs with
paths, queries, and fragments. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10979">#10979</a>)</li>
<li>Standardized built-in placeholder emails to the namespaced
<code>{identifier}@{namespace}.placeholder.invalid</code> format. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10982">#10982</a>)</li>
<li>Added synchronous and optional access to the current auth endpoint
context. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10938">#10938</a>)</li>
</ul>
<p>For detailed changes, see <a
href="c50200bfc7/packages/core/CHANGELOG.md"><code>CHANGELOG</code></a></p>
<h2><code>@better-auth/oauth-provider</code></h2>
<h3>Bug Fixes</h3>
<ul>
<li>Fixed Client ID Metadata Document registration when clients share at
least one supported grant with the server. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/11010">#11010</a>)</li>
<li>Improved validation of relative callback and redirect URLs with
paths, queries, and fragments. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10979">#10979</a>)</li>
<li>Fixed relative redirect URLs containing fragments. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10983">#10983</a>)</li>
</ul>
<p>For detailed changes, see <a
href="c50200bfc7/packages/oauth-provider/CHANGELOG.md"><code>CHANGELOG</code></a></p>
<h2><code>@better-auth/drizzle-adapter</code></h2>
<h3>Bug Fixes</h3>
<ul>
<li>Fixed one-to-one Drizzle relations when <code>usePlural</code> is
enabled. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10941">#10941</a>)</li>
<li>Added validation for missing Drizzle schema fields in compound
<code>where</code> clauses. (<a
href="https://redirect.github.com/better-auth/better-auth/pull/10859">#10859</a>)</li>
</ul>
<p>For detailed changes, see <a
href="c50200bfc7/packages/drizzle-adapter/CHANGELOG.md"><code>CHANGELOG</code></a></p>
<h2><code>@better-auth/kysely-adapter</code></h2>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/better-auth/better-auth/blob/main/packages/better-auth/CHANGELOG.md">better-auth's
changelog</a>.</em></p>
<blockquote>
<h2>1.7.2</h2>
<h3>Patch Changes</h3>
<ul>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10875">#10875</a>
<a
href="d5d889bfd8"><code>d5d889b</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Fix programmatic migrations failing on Cloudflare D1 while preserving
existing-index validation across supported databases.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10982">#10982</a>
<a
href="b4ad5a110c"><code>b4ad5a1</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Built-in placeholder emails now consistently use the namespaced
<code>{identifier}@{namespace}.placeholder.invalid</code> format.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10934">#10934</a>
<a
href="c7a5c1a7ed"><code>c7a5c1a</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Cookie-cache reads now warn when signed session data is invalid
instead of silently appearing as a signed-out session.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10879">#10879</a>
<a
href="78f0c3922c"><code>78f0c39</code></a>
Thanks <a
href="https://github.com/apps/starslingdev"><code>@​starslingdev</code></a>!
- Test suites using <code>getTestInstance</code> now run faster because
the shared fixture avoids production password-hashing costs by default.
Custom <code>emailAndPassword.password</code> implementations continue
to take precedence.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10823">#10823</a>
<a
href="ce8a3ab544"><code>ce8a3ab</code></a>
Thanks <a href="https://github.com/sosyz"><code>@​sosyz</code></a>! -
Ensure permanently banning a user clears any expiration from a previous
temporary ban.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10907">#10907</a>
<a
href="a021eafaf2"><code>a021eaf</code></a>
Thanks <a href="https://github.com/heliohm"><code>@​heliohm</code></a>!
- A client created with more plugins is again assignable to a client
type declaring fewer plugins, as in 1.6.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10959">#10959</a>
<a
href="c8dcfa57e1"><code>c8dcfa5</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Allow same-origin form submissions from pages using
<code>Referrer-Policy: no-referrer</code> while continuing to reject
untrusted request origins.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10979">#10979</a>
<a
href="fced1a5d36"><code>fced1a5</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Allow relative callback and redirect URLs to use standard path, query,
and fragment syntax while preserving open-redirect protections.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10041">#10041</a>
<a
href="f6891a2d2d"><code>f6891a2</code></a>
Thanks <a
href="https://github.com/GautamBytes"><code>@​GautamBytes</code></a>! -
Allow <code>~</code> in relative callback URLs validated by trusted
origin checks.</p>
</li>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10877">#10877</a>
<a
href="649818a296"><code>649818a</code></a>
Thanks <a href="https://github.com/bytaesu"><code>@​bytaesu</code></a>!
- Prevent disabled MyISAM indexes from satisfying migration index
checks.</p>
</li>
<li>
<p>Updated dependencies [<a
href="557e19bfad"><code>557e19b</code></a>,
<a
href="64da15b0b1"><code>64da15b</code></a>,
<a
href="d5d889bfd8"><code>d5d889b</code></a>,
<a
href="b4ad5a110c"><code>b4ad5a1</code></a>,
<a
href="ea77118d4e"><code>ea77118</code></a>,
<a
href="5aea9f7728"><code>5aea9f7</code></a>,
<a
href="fced1a5d36"><code>fced1a5</code></a>,
<a
href="e1d40116e2"><code>e1d4011</code></a>]:</p>
<ul>
<li><code>@​better-auth/core</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/kysely-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/drizzle-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/memory-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/mongo-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/prisma-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
<li><code>@​better-auth/telemetry</code><a
href="https://github.com/1"><code>@​1</code></a>.7.2</li>
</ul>
</li>
</ul>
<h2>1.7.1</h2>
<h3>Patch Changes</h3>
<ul>
<li>
<p><a
href="https://redirect.github.com/better-auth/better-auth/pull/10863">#10863</a>
<a
href="845bbd1de6"><code>845bbd1</code></a>
Thanks <a
href="https://github.com/gustavovalverde"><code>@​gustavovalverde</code></a>!
- <code>auth migrate</code> no longer attempts to add a required column
with no default value to a table that already has rows. It stops with an
error naming the column and the backfill to run first. Previously the
generated statement failed on SQLite, Postgres, and SQL Server; on MySQL
it filled the new column with an empty string for every existing row and
reported success. If <code>auth migrate</code> already ran against a
MySQL database on 1.7, run the check in the upgrade guide's account
identity section.</p>
<p><code>getMigrations</code> throws the new
<code>UnsafeMigrationError</code> (exported from
<code>better-auth/db/migration</code>) for this refusal, so callers can
distinguish it from other migration errors such as an index-definition
conflict.</p>
<p><code>auth generate</code> still emits the statements for external
migration tooling, with a comment banner naming any column that needs a
manual backfill first.</p>
<p>A required field whose database column is still nullable logs a
warning instead of blocking the migration.</p>
<p>A CLI command that fails now prints its error and exits with a
non-zero code instead of an unhandled promise rejection.</p>
</li>
<li>
<p>Updated dependencies []:</p>
<ul>
<li><code>@​better-auth/core</code><a
href="https://github.com/1"><code>@​1</code></a>.7.1</li>
<li><code>@​better-auth/drizzle-adapter</code><a
href="https://github.com/1"><code>@​1</code></a>.7.1</li>
</ul>
</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="ba12fcdfa7"><code>ba12fcd</code></a>
chore: release v1.7.2 (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10870">#10870</a>)</li>
<li><a
href="79904f0be8"><code>79904f0</code></a>
fix(origin-check): support fragments in relative redirect URLs (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10983">#10983</a>)</li>
<li><a
href="c8dcfa57e1"><code>c8dcfa5</code></a>
fix(origin-check): validate null origins using fetch metadata (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10959">#10959</a>)</li>
<li><a
href="e1d40116e2"><code>e1d4011</code></a>
fix(logger): respect configured logger in auth request context (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10939">#10939</a>)</li>
<li><a
href="557e19bfad"><code>557e19b</code></a>
refactor(context): clarify auth endpoint context access (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10938">#10938</a>)</li>
<li><a
href="b4ad5a110c"><code>b4ad5a1</code></a>
refactor: centralize placeholder email generation (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10982">#10982</a>)</li>
<li><a
href="fced1a5d36"><code>fced1a5</code></a>
fix(origin-check): improve relative callback URL validation (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10979">#10979</a>)</li>
<li><a
href="f6891a2d2d"><code>f6891a2</code></a>
fix(origin-check): allow tilde in relative callback URLs (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10041">#10041</a>)</li>
<li><a
href="ce8a3ab544"><code>ce8a3ab</code></a>
fix(admin): ban without a duration should clear the previous expiration
(<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10823">#10823</a>)</li>
<li><a
href="a021eafaf2"><code>a021eaf</code></a>
fix(client): a client with more plugins fits a narrower client type
again (<a
href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/1">#1</a>...</li>
<li>Additional commits viewable in <a
href="https://github.com/better-auth/better-auth/commits/v1.7.2/packages/better-auth">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 09:46:13 -07:00
dependabot[bot] eb7b4d1371
chore(deps): bump @aws-sdk/client-s3 from 3.1115.0 to 3.1120.0 (#12567)
Bumps
[@aws-sdk/client-s3](https://github.com/aws/aws-sdk-js-v3/tree/HEAD/clients/client-s3)
from 3.1115.0 to 3.1120.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/releases">@​aws-sdk/client-s3's
releases</a>.</em></p>
<blockquote>
<h2>v3.1120.0</h2>
<h4>3.1120.0(2026-08-27)</h4>
<h5>Documentation Changes</h5>
<ul>
<li><strong>client-opensearch:</strong> Updating SDK and CLI
documentation for AttachDataSource API. (<a
href="d696fe7602">d696fe76</a>)</li>
</ul>
<h5>New Features</h5>
<ul>
<li><strong>client-lambda-microvms:</strong> Added
InsufficientCapacityException to RunMicrovm for capacity-related
failures. Added lifecycle status field (AVAILABLE, DEPRECATED) to
ListManagedMicrovmImageVersions. Added ConflictException to
CreateMicrovmAuthToken and CreateMicrovmShellAuthToken for unregistered
MicroVMs. (<a
href="72a8ff8092">72a8ff80</a>)</li>
<li><strong>client-codedeploy:</strong> Added a deploymentMode parameter
to CreateDeployment. Set it to RESTART to restart an EC2 and on-premises
fleet, using the last successful revision, honoring Deployment
Configuration. (<a
href="78d4f9640b">78d4f964</a>)</li>
<li><strong>client-cloudwatch-logs:</strong> Added resultCount to
QueryStatistics in GetQueryResults. This field returns the total number
of output rows in the final result set, helping customers
programmatically determine whether a query produced results after all
operations including post-aggregation filters. (<a
href="0e4d242b71">0e4d242b</a>)</li>
<li><strong>client-datazone:</strong> Add cascadeDelete to DeleteDomain.
When specified, DataZone recursively deletes all projects, environments,
subscriptions, and their underlying AWS resources before removing the
domain. Deletion progress is reported via deleteProgress and resource
failures via failureReasons on GetDomain. (<a
href="3a74dc4b94">3a74dc4b</a>)</li>
<li><strong>client-rds:</strong> Adding support for the full snapshot
size, in bytes, of DB instance snapshots. (<a
href="ab2f66f5f5">ab2f66f5</a>)</li>
<li><strong>client-ec2:</strong> EC2 allows AMI owners to define
compatible instance types on their AMIs, blocking RunInstances calls
automatically for launches on non-permitted instance types. (<a
href="311b3b26db">311b3b26</a>)</li>
<li><strong>client-cognito-identity-provider:</strong> Adds the
AdminDeleteSoftwareToken API operation, enabling administrators to
remove a user's registered TOTP (software token) MFA configuration from
a user pool. (<a
href="f661bebc4d">f661bebc</a>)</li>
</ul>
<hr />
<p>For list of updated packages, view
<strong>updated-packages.md</strong> in
<strong>assets-3.1120.0.zip</strong></p>
<h2>v3.1119.0</h2>
<h4>3.1119.0(2026-08-26)</h4>
<h5>Chores</h5>
<ul>
<li><strong>codegen:</strong> smithy-aws-typescript-codegen 0.53.0 (<a
href="https://redirect.github.com/aws/aws-sdk-js-v3/pull/8276">#8276</a>)
(<a
href="dffb383bdc">dffb383b</a>)</li>
</ul>
<h5>New Features</h5>
<ul>
<li><strong>client-sagemaker:</strong> Amazon SageMaker AI now supports
ml.g7 instances for model optimization. You can now run model
optimization jobs on ml.g7 instances, in supported AWS Regions. (<a
href="6d5e106634">6d5e1066</a>)</li>
<li><strong>client-devops-agent:</strong> AWS DevOps Agent now supports
trigger filter groups for Release Readiness Review, letting you control
when the capability auto-triggers based on webhook events and target
branches. (<a
href="bc3d53d550">bc3d53d5</a>)</li>
<li><strong>client-license-manager-user-subscriptions:</strong> Released
support for License Expiry field in ListProductSubscriptions API (<a
href="454d7f7ffb">454d7f7f</a>)</li>
<li><strong>client-ec2:</strong> Adds deleting state to possible VPC
States. (<a
href="43091d55b3">43091d55</a>)</li>
<li><strong>client-network-firewall:</strong> Adding new status enum for
Firewalls. (<a
href="4cb21cb3b8">4cb21cb3</a>)</li>
</ul>
<hr />
<p>For list of updated packages, view
<strong>updated-packages.md</strong> in
<strong>assets-3.1119.0.zip</strong></p>
<h2>v3.1118.0</h2>
<h4>3.1118.0(2026-08-25)</h4>
<h5>Documentation Changes</h5>
<ul>
<li><strong>client-marketplace-metering:</strong> Updated documentation
to clarify duplicate-billing prevention and BatchMeterUsage retry
guidance (<a
href="322310259e">32231025</a>)</li>
</ul>
<h5>New Features</h5>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/aws/aws-sdk-js-v3/blob/main/clients/client-s3/CHANGELOG.md">@​aws-sdk/client-s3's
changelog</a>.</em></p>
<blockquote>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1119.0...v3.1120.0">3.1120.0</a>
(2026-08-27)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1118.0...v3.1119.0">3.1119.0</a>
(2026-08-26)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1117.0...v3.1118.0">3.1118.0</a>
(2026-08-25)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1116.0...v3.1117.0">3.1117.0</a>
(2026-08-24)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
<h1><a
href="https://github.com/aws/aws-sdk-js-v3/compare/v3.1115.0...v3.1116.0">3.1116.0</a>
(2026-08-21)</h1>
<p><strong>Note:</strong> Version bump only for package
<code>@​aws-sdk/client-s3</code></p>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="d6be6f8dd3"><code>d6be6f8</code></a>
Publish v3.1120.0</li>
<li><a
href="ba4e4498a7"><code>ba4e449</code></a>
Publish v3.1119.0</li>
<li><a
href="c65dd6533d"><code>c65dd65</code></a>
Publish v3.1118.0</li>
<li><a
href="78b069ac77"><code>78b069a</code></a>
Publish v3.1117.0</li>
<li><a
href="d760a00859"><code>d760a00</code></a>
Publish v3.1116.0</li>
<li><a
href="8369ada75d"><code>8369ada</code></a>
chore(codegen): update to sync with the latest smithy-ts (<a
href="https://github.com/aws/aws-sdk-js-v3/tree/HEAD/clients/client-s3/issues/8272">#8272</a>)</li>
<li>See full diff in <a
href="https://github.com/aws/aws-sdk-js-v3/commits/v3.1120.0/clients/client-s3">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 08:43:04 -07:00
Nicky Leach 1d493eb62a
test(heartbeat): drain in-flight runs before native-isolation TRUNCATE (#12751)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip runs agent heartbeats and stores their run state in a
database
> - The direct-adapter native-isolation tests start heartbeat runs and
then clear database state
> - A terminal run status does not prove that its background database
work has stopped
> - The teardown can then deadlock with a live run during PostgreSQL
`TRUNCATE`
> - This pull request drains active runs before teardown and adds a
guard for queued or running runs
> - The benefit is stable test teardown without a production code change

## Linked Issues or Issue Description

This change fixes an intermittent test deadlock in the direct-adapter
native-isolation suite.

**What happened?**

The test teardown could run PostgreSQL `TRUNCATE` while a heartbeat
execution still held a write transaction. PostgreSQL then returned error
`40P01` during some test runs.

**Expected behavior**

The test teardown must wait until all heartbeat executions finish before
it clears the test database.

**Steps to reproduce**

1. Run
`server/src/__tests__/heartbeat-direct-adapter-native-isolation.test.ts`
repeatedly.
2. Run the suite against PostgreSQL-backed native isolation.
3. Observe intermittent deadlock error `40P01` during teardown.

**Paperclip version or commit**

Commit `57515726d3ef45a07df9b5ee2dfaf7d108556478`.

**Deployment mode**

Built from source with the native-isolation test suite.

**Agent adapter(s) involved**

Not adapter-specific. The test covers the direct adapter path.

**Database mode**

External PostgreSQL used by the native-isolation test suite.

**Additional context**

Related prior attempt:
[#12715](https://github.com/paperclipai/paperclip/pull/12715). This pull
request starts from current `master` and does not depend on that pull
request.

## What Changed

- Drain active heartbeat run executions before `afterEach` runs
`TRUNCATE`.
- Assert that no heartbeat run remains `queued` or `running` before
teardown.
- Drain active executions before `afterAll` removes the temporary
database.
- Create one shared `heartbeatService` instance in `beforeAll` so the
drain tracks the test runs.

## Verification

- Run
`server/src/__tests__/heartbeat-direct-adapter-native-isolation.test.ts`
20 times. All 20 runs pass.
- Run the target suite with
`server/src/__tests__/native-run-finalizer.test.ts`. Both files pass
with 19 tests.
- Run `tsc --noEmit`. The branch adds no new error compared with
`master`.
- Run the pull request checks after GitHub starts them.

## Risks

Low risk. The change affects one test file and no production code. The
added drain can expose an incomplete test run before teardown, which is
the intended guard.

## Model Used

OpenAI GPT-5. Exact runtime model ID: GPT-5. The context window is not
exposed to this agent. The model used tool calls and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 06:38:37 -07:00
scotttong 597fd63b61
feat(ui): add streamlined navigation foundation (#12746) 2026-09-02 23:55:43 -07:00
Nicky Leach 9064cfd09e
feat(codex-local): give each Codex account its own home and path secret (#12709)
## Thinking Path

> - Paperclip is the control plane for companies that use AI agents for
work
> - Local adapters connect Paperclip agents to provider command line
tools
> - The Codex adapter stores login data in a shared company home
> - A shared home cannot keep credentials for more than one Codex
account
> - This pull request gives each account a safe home and a matching
company secret
> - The benefit is that one company can use multiple Codex accounts at
the same time

## Linked Issues or Issue Description

**Problem or motivation**

A company can hold only one Codex subscription credential because device
login uses one shared home. A second account cannot log in without
replacing or conflicting with the first credential.

**Proposed solution**

This change validates the vendor account identifier, stores each
credential in its own home, and creates a company secret that points to
that home. Repeat login calls return success when the matching secret
already exists.

**Roadmap alignment**

The change supports the roadmap goal for centrally managed secrets with
scoped access and audited resolution.

**Additional context**

The security review returned approve with no blocking finding. The
branch adds shared account-handle validation and tests for device login
and the Codex local adapter.

## What Changed

- Add strict allowlist validation for Codex account handles.
- Store each Codex account credential in a separate home under the Codex
cache root.
- Verify that the resolved account home stays inside the cache root.
- Create the `CODEX_HOME_<handle>` company secret for each account.
- Keep repeat and concurrent login calls safe and idempotent.
- Add shared helper and route, adapter, and validation tests.

## Verification

- `pnpm --filter @paperclipai/adapter-codex-local test` passes with 343
tests.
- `pnpm --filter @paperclipai/server test
src/__tests__/agent-device-login-routes.test.ts` passes with 25 tests.
- The adapter suite passes with 23 tests.
- The shared package and Codex adapter typechecks pass.
- Continuous integration must pass on every check before merge.

## Risks

The account handle becomes part of a directory path and secret name. The
strict allowlist and root containment check reduce path traversal risk.
Existing single-account homes remain unchanged unless a new device login
creates an account-specific home.

## Model Used

OpenAI GPT-5 (exact runtime model ID: gpt-5), with tool use and code
execution. The runtime context window is not exposed in this run.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-02 14:46:53 -07:00
Michael Nguyen dfdfc8664e
feat(claude-local): add Claude Fable 5.1 support (#12730)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Claude local adapter lets operators select a Claude model for an
agent.
> - Claude Fable 5.1 was absent from the adapter model lists.
> - The adapter runtime also used a Claude Code build that rejected
Fable 5.1.
> - This pull request adds the direct Anthropic ID and the AWS Bedrock
inference profile ID.
> - It also updates the Claude ACP runtime and keeps the Paperclip usage
and isolation patches.
> - The benefit is that operators can select and run Claude Fable 5.1
through the Claude adapter.

## Linked Issues or Issue Description

Refs #8810. That issue covers related model ID handling. This change
does not change provider-prefixed model IDs.

**Agent or provider**

Claude Code through the built-in `claude_local` adapter. The requested
model is Claude Fable 5.1.

**Why this adapter is useful**

Operators can use Fable 5.1 without entering an undocumented model ID.
The configured model also reaches both supported Claude execution lanes.

**How the agent is invoked**

The CLI lane sends `--model claude-fable-5-1`. The ACP lane sends
`ANTHROPIC_MODEL=claude-fable-5-1` to
`@agentclientprotocol/claude-agent-acp`.

**Are you willing to implement it?**

Yes. This pull request includes the implementation and tests.

**Additional context**

Claude Code 2.1.232 rejected Fable 5.1 and required version 2.1.251 or
newer. ACP package 0.73.0 includes Claude Code 2.1.257. The update keeps
Paperclip's usage metadata and isolated-context behavior.

## What Changed

- Added `claude-fable-5-1` to the direct Claude fallback list.
- Added `us.anthropic.claude-fable-5-1` to the AWS Bedrock list.
- Kept the existing default model at the first position in each list.
- Updated the Claude ACP dependency from 0.70 to 0.73.
- Carried the Paperclip usage and isolated-context changes into the 0.73
patch.
- Added a Claude Code 2.1.251 minimum-version preflight for Fable 5.1
when using the standard `claude` executable, surfaced in both adapter
Test and execution. Explicit custom wrappers retain their existing
compatibility contract.
- Kept local adapter Tests from executing caller-selected binaries: when
runtime `PATH` selects a different Claude executable than the trusted
probe, the Test warns and defers the authoritative version check to
execution instead of approving or rejecting the alternate installation.
- Added tests for model listing, discovery deduplication, Bedrock
filtering, model pass-through in both execution lanes, old-CLI rejection
before launch, custom-wrapper compatibility, and local runtime-PATH
mismatch handling.

## Verification

- `pnpm --filter @paperclipai/adapter-claude-local typecheck`
- `pnpm exec vitest run
packages/adapters/claude-local/src/server/execute.remote.test.ts
packages/adapters/claude-local/src/server/test.remote.test.ts
packages/adapters/claude-local/src/server/test.probe.test.ts
packages/adapters/claude-local/src/server/acp.test.ts
server/src/__tests__/adapter-models.test.ts` (72 tests passed)
- `node --test scripts/acpx-patch-packaging.test.mjs` (13 tests passed)
- `pnpm -r typecheck`
- `pnpm build`
- A local Paperclip agent run completed with `usageJson.model` set to
`claude-fable-5-1` through ACP 0.73.0 and its bundled Claude Code
2.1.257.
- `pnpm test:run` completed 5,638 passing tests and 24 skipped tests. It
also found 24 failures in unrelated workspace-runtime,
path-canonicalization, and runtime-exposure tests on macOS with Node 26.
These failures do not touch this diff. Clean pull request CI is the
final full-suite gate.

## Risks

- The ACP dependency update can change Claude runtime behavior outside
model selection. Focused ACP tests, the full typecheck, the production
build, and a real local Fable run reduce this risk.
- The 0.73 patch must stay aligned with the installed ACP version.
Dependency-resolution CI verifies the manifest and patch pair.
- Fable 5.1 adds a short `claude --version` preflight to standard
CLI-lane Tests and runs. The result is intentionally not cached so an
in-place Claude Code upgrade takes effect without restarting Paperclip.
Explicit custom wrappers are not version-probed because their output and
compatibility contract can differ from the standard executable.
- Local Tests preserve the existing deny-by-default probe boundary and
do not execute a binary selected by caller-controlled `PATH`. A
mismatched runtime binary produces an explicit warning without blocking
an otherwise valid setup; execution independently validates the actual
runtime-selected CLI before launch.
- The AWS Bedrock identifier differs from earlier IDs because Fable 5.1
has no `-v1` suffix. The model-list test locks this exact value.
- There is no schema change or migration.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

Provider: OpenAI. Model: GPT-5 Codex. The host did not expose a more
specific model ID or context-window size. Capabilities used: agentic
reasoning, repository editing, shell execution, web research, and local
runtime verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-02 14:32:01 -07:00
Dotta 0f94521017
fix(runner): restore local session and task integrity (#12721)
## Thinking Path

> - Paperclip is the control plane for agents that perform work.
> - Paperclip Runner connects durable provider sessions to individual
task runs through PRP.
> - Provider continuity and per-run authority are different lifetimes.
> - The existing implementation mixed those lifetimes and lost event
metadata between provider frames, runnerd, persistence, API
sanitization, and the task thread.
> - That caused failed continuation, missing progress and Plans,
duplicate replies, hidden failures, and unsafe recovery.
> - This repair gives every heartbeat fresh authority, preserves
qualified provider-session continuity, and restores one lossless
presentation path without changing direct adapters.

## Linked Issues or Issue Description

**What happened?**

A second native heartbeat could reuse tickets, leases, command receipts,
sequence state, and run identity from the first heartbeat. Provider
phase and item identity could be lost before the UI read them. Redaction
could corrupt protocol discriminators while still missing malformed
credential tails. The task thread could fold progress into the final
response, hide failures, or show more than one final answer. Native
Codex also exposed approval modes that do not yet have a durable
approval bridge.

**Expected behavior**

Each heartbeat uses a new PRP authority epoch. Codex and OpenCode
preserve exact qualified provider sessions; ACPX emits an explicit
continuity event when its qualified process-replacement policy is used.
Every accepted provider event is presented, classified as internal, or
surfaced as unsupported. The task page shows chronological progress,
reasoning summaries, activity, Plans, interactions, terminal failures,
and exactly one final reply. Direct adapters retain their existing path.

**Steps to reproduce**

1. Enable the unified experimental Paperclip Runner setting.
2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex
agent.
3. Run response, Plan, structured-question/resume, restart,
cancellation, and failure scenarios.
4. Reload the task while active, waiting, failed, and settled.
5. On the old implementation, observe stale run authority, missing
classifications, incomplete output, or duplicated/folded replies.

**Paperclip version or commit**

The repair is based directly on `master` at
`87d05e194b643810d16d20612115acd01d735d43`.

**Deployment mode**

Local development with the embedded database.

Related work: Refs #12616, #12646, #12666, #12685, and #12700.

## What Changed

- Rotates PRP control-plane, outbox, ticket, lease, command, receipt,
and sequence authority for each heartbeat while carrying forward only a
validated provider-session identity.
- Reads `control-plane-state.json`, validates both durable schemas and
lifecycle values, resumes coherent current runs, archives qualified
settled authority, and quarantines malformed or mismatched scoped state
without moving ambiguous live legacy state.
- Preserves Codex provider phase and stable item identities so
commentary remains progress and only `final_answer` becomes final.
- Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning
lifecycle mapping.
- Makes ACPX normalization lossless for visible reasoning, tool
lifecycle metadata, stable bounded identities, Plan revisions,
structured requests, failures, and qualified process replacement. Only
the compatible terminal assistant message is promoted as final.
- Applies schema-aware redaction before generic JWT-shaped detection and
scans every diagnostic string leaf. Malformed raw/escaped quoted
credential tails are redacted in both server and durable Rust state.
- Restores snapshot-style chronological task presentation, expandable
tool activity, inline Plan cards, visible waiting/resume/cancel/failure
states, and exactly one final answer.
- Makes `never` the only qualified native Codex permission mode and
rejects unsupported persisted native modes with remediation. OpenCode
and ACPX policies remain intact.
- Keeps the unified experimental Runner setting as the only enablement
flag. Onboarding and direct Codex, Claude, and OpenCode stay on their
legacy execution/finalization paths.
- Adds cross-language goldens, authority/recovery/fault coverage, exact
response/count assertions, and native plus legacy acceptance scenarios.

## Verification

- Pull-request GitHub Actions run Rust formatting/tests, TypeScript
checks, server/UI tests, builds, protocol drift checks, browser E2E, and
security scans.
- A separate workflow-only validation ref is pinned directly on this PR
head and runs the 35-cell paid local matrix: three core scenarios plus
structured-question resume and restart/resume for native Codex, native
OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode.
Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315
- Acceptance requires exact single visible replies, monotonic sequences,
matching envelope discriminators, one semantic terminal, one run
terminal, no unresolved interaction, no duplicate mutation, no secret
leakage, provider continuity, and zero native rows for direct adapters.
- Per maintainer direction, tests are running in GitHub Actions rather
than on the slower local host. Only formatters and static diff checks
were run locally.

## Risks

- Recovery from old or partial filesystem state is sensitive. The repair
fails closed, preserves active or unverifiable authority, and
quarantines only state whose scoped ownership is safe to move.
- Provider event formats can change. Closed validators and boundary
goldens turn new or malformed events into visible diagnostics instead of
silent drops.
- Shared task presentation could affect direct adapters. Runtime-fact
gating plus the direct-adapter matrix protect the existing path.
- Managed and remote providers are not qualified here. Shared code
continues to compile and fail safely, but live qualification is
deferred.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5. The exact deployed snapshot and
context-window size are not exposed to this task. It used agentic
reasoning, repository inspection, code editing, Git, parallel subagents,
and GitHub Actions.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] The paid local-provider matrix is green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 16:11:26 -05:00
Dotta 87d05e194b
feat(work-products): add rich cards and run artifact inventory (#12717)
## Thinking Path

> - Paperclip is the control plane for AI-agent companies.
> - Agent outputs must remain visible after a run and easy to inspect
from a task.
> - The thread and artifact inventory need one consistent rich-card
vocabulary.
> - Run uploads also need durable artifact registration and
producing-run context.
> - Reviewers need deterministic examples for each rich-card kind and
state.
> - This pull request adds the shared presentation, registration,
inventory, and Storybook review coverage.
> - The benefit is a complete output path that reviewers can inspect
without seeded data.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This change improves work-product presentation in task threads and the
task Artifacts tab.

**Subsystem affected**

The change affects shared work-product contracts, the runner diff path,
server attachment and work-product services, GitHub metadata refresh,
the React board UI, and Storybook.

**Current behavior**

The thread used generic cards. Some files uploaded by a run existed only
as message attachments. The Artifacts tab showed a flat list without run
context or filters. Storybook showed only one resting card per kind.

**Proposed behavior**

The thread uses rich cards for supported work-product types. Each
run-produced file registers one attachment-backed artifact work product.
The Artifacts tab groups outputs by run and supports filters. Storybook
shows every kind and requested state, PR lifecycle states, stats
variants, truncation, mobile layout, and message-tail media.

**Reason and benefit**

Users can identify outputs quickly. Reviewers can inspect all card
permutations without creating task data.

**Breaking changes**

None. The metadata fields and automatic artifact registration are
additive. Existing attachments and work products keep their current
behavior.

## What Changed

- Added a shared rich work-product card with kind-specific content and a
compact inventory variant.
- Added pull-request and commit diff metadata plus bounded GitHub state
refresh.
- Added media strips and typed file chips to message-tail attachments.
- Registered each run-produced attachment as an artifact work product in
the same server transaction.
- Grouped task artifacts by run with agent and timestamp headings.
- Added type and run filters, image thumbnails, compact cards, and a
company Artifacts link.
- Added a Storybook kind-by-state matrix with stats variants for all
eight visual kinds.
- Added PR open, draft, merged, and closed examples, long-title
truncation, an exact 375-pixel viewport, and message-tail overflow
coverage.
- Closed reconciled runtime work products when the linked runtime stops
or disappears, so the card shows `Stopped` instead of `Unhealthy`.

### Screenshots

Before: one resting card per kind.

![Previous rich-card
inventory](https://pages.paperclip.ing/rich-work-product-storybook-20260902/before-inventory.png)

After: the kind and state matrix.

![Rich-card kind and state
matrix](https://pages.paperclip.ing/rich-work-product-storybook-20260902/after-kind-state-matrix.png)

After: message-tail media at 375 pixels.

![Message-tail thumbnails and typed
chips](https://pages.paperclip.ing/rich-work-product-storybook-20260902/after-message-tail.png)

[Open the Storybook evidence
viewer](https://pages.paperclip.ing/rich-work-product-storybook-20260902/).

The earlier artifact inventory comparison remains available in the
[artifact inventory
viewer](https://pages.paperclip.ing/rich-artifacts-inventory-proof-20260902/).

## Verification

- `pnpm --filter @paperclipai/ui typecheck` passed.
- `pnpm check:token-gates` passed.
- `pnpm build-storybook` passed.
- `pnpm exec vitest run
server/src/__tests__/work-product-runtime-reconciliation.test.ts` passed
with 5 tests.
- Chromium visual checks passed at desktop and 375-pixel widths.
- All 30 latest-head GitHub checks passed. One unrelated annotation test
was flaky and passed on its single retry.
- Greptile passed at 5/5 with zero unresolved threads.

## Risks

- Low risk. The Storybook change adds review fixtures only. The runtime
fix changes read-time reconciliation without database writes.
- The matrix is intentionally large so every permutation stays visible
in one review surface.

> I checked `ROADMAP.md`. This work does not duplicate planned core
work.

## Model Used

- OpenAI Codex with GPT-5 and GPT-5.6-sol across this pull request.
Reasoning, tool use, and code execution were enabled. The context-window
size is not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My public branch name describes the change and contains no
internal task id
- [x] I have run tests locally and the changed-path tests pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-02 15:27:54 -05:00
Dotta 8c3b8c432a
Simplify app connections and enable managed Google access (#12728)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps subsystem gives humans and agents governed access to
external tools.
> - The current connection flow hides Apps behind an experimental gate
and repeats setup text.
> - Google sharing choices and generic MCP permissions do not use one
consistent opening model.
> - Self-hosted installs also need a safe default origin for managed
OAuth without a manual config file.
> - This pull request makes Apps available, simplifies connection setup,
and applies one governed permissions model.
> - The benefit is a shorter connection flow that works on a clean
self-hosted install.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the Apps connection setup flow, managed Google connection
flow, generic MCP connection flow, navigation, and runtime origin
discovery.

**Subsystem affected**

Cross-cutting. This changes `ui/`, `server/`, `packages/shared/`,
connector documentation, and browser tests.

**Current behavior**

Apps require an experimental switch. Setup pages repeat titles and
explanatory copy. Connection names require manual input. Google
credential sharing does not always offer both personal and organization
access. Generic MCP providers do not start with the same permission
choices. Managed OAuth needs a public URL setting even when the request
already has a safe HTTPS origin.

**Proposed behavior**

Apps are available by default. Setup asks only for required permissions
and sharing choices. Paperclip creates conflict-free connection names.
Google apps and generic MCP providers use the same human and agent
access model. Managed OAuth derives a validated same-origin HTTPS URL
when no explicit public URL is set.

**Reason and benefit**

A clean self-hosted install can connect a managed Google app without
hidden setup. Humans can share a service account with their
organization. The shorter flow reduces duplicated choices and setup
errors.

**Breaking changes**

The Apps experimental switch is removed. Existing connection APIs remain
compatible. New connections can receive a numeric suffix when a name
already exists.

No duplicate or related public issue was found.

## What Changed

- Removed the Apps experimental gate and the breadcrumb that leaves the
Apps section.
- Simplified all connection setup pages and moved optional provider
requirements into one small link.
- Added consistent human and agent access choices for Google apps,
Zapier, and generic MCP connections.
- Added organization sharing to Google Workspace credentials while
keeping personal access available.
- Generated connection names automatically and resolved name conflicts
with numeric suffixes.
- Derived a validated public HTTPS origin from the request for
config-free managed OAuth.
- Updated connector contracts, tests, browser coverage, and authoring
documentation.

## Verification

- `pnpm check:token-gates`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts` (273 passed)
- Targeted UI/service regression suite (308 passed)
- Six targeted Playwright connection journeys on a fresh onboarding
instance (6 passed)
- Fresh-install browser proof through Tailscale HTTPS: enrolled with
Paperclip Cloud, connected managed Google Drive, and completed a real
read operation.
- [Exact-head CI
run](https://github.com/paperclipai/paperclip/actions/runs/33669760711):
all 23 matrix jobs passed, including build, typecheck, server,
serialized, canary, and all browser shards.
- Greptile 5/5 on `0ae2a859f269984ee950d0af231a5b09a06f3dfd`, with no
unresolved review threads.

## Risks

Apps are now visible to all operators. The removed experimental flag no
longer hides unfinished app definitions. Managed Google availability
still depends on the Cloud profile rollout and active instance
enrollment. Automatic conflict handling changes only the display name of
a newly conflicting connection.

> I checked [`ROADMAP.md`](ROADMAP.md). MCP Tool Gateway and Apps are
shipped. Connected Apps is planned, and this change improves the
existing shipped connection flow.

## Model Used

OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-02 14:05:53 -05:00
Dotta fdf8c8464d
feat(runner): add managed provider backends (#12699)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner provides durable, provider-neutral agent
execution.
> - The current stack supports qualified local providers but omits the
managed provider paths from the integration branch.
> - Claude Managed Agents and AWS AgentCore need explicit profile
qualification, durable recovery, usage accounting, and cleanup controls.
> - This pull request adds those managed backends as the third part of
the Runner parity stack.
> - The benefit is managed execution without weakening the default-off
Runner rollout gate.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: Runner, server orchestration, database profiles, CLI, and
adapter configuration UI.

**Problem or motivation**

The current Runner stack cannot select or execute the managed Claude
Agents API or AWS Bedrock AgentCore Harness backends. It also lacks
qualified profile storage and recovery checks for those remote
resources.

**Proposed solution**

Add qualified managed and remote profiles, API and CLI management, exact
provider selection, durable lifecycle handling, cumulative usage
accounting, bounded cleanup, and retention acknowledgement. Keep
`enableNativeRunner` default-off.

**Alternatives considered**

A direct copy of the old integration branch was rejected because its
provider contracts, model values, credential flow, and migration history
no longer match the current base. A single large parity pull request was
also rejected because stacked review keeps each subsystem bounded.

**Roadmap alignment**

This continues the existing Runner architecture and rollout work. It
does not introduce a separate execution system.

**Additional context**

This pull request is based on the merged #12691 and #12685 stack. It
also closes the delayed security-review findings reported on #12691 by
binding qualified ACPX and OpenCode launch artifacts to the bytes
actually executed. A GitHub search for managed agent, AgentCore, and
Claude managed work found no duplicate public issue or pull request.

## What Changed

- Add Claude Managed Agents and AWS AgentCore provider executors to
runnerd.
- Add qualified managed and remote profile storage, routes, OpenAPI
contracts, CLI commands, and migration 0237.
- Validate profile ownership, enabled state, exact qualified revision,
model, agent version, and secret binding before persistence and
recovery.
- Persist durable provider session and owned skill state for
restart-safe cleanup.
- Reconcile uncertain create responses and delete remote sessions before
owned skills.
- Track cumulative provider usage and enforce positive session spend
caps.
- Recover interrupted AgentCore usage at the next turn boundary by
charging the prior invocation ceiling exactly once; keep the session
gated until an explicit monotonic budget raise.
- Isolate AgentCore AWS configuration from host profiles and
credential-process/SSO configuration while preserving workload identity.
- Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact
paths; remove the ambient executable override.
- Snapshot and content-verify ACPX and OpenCode commands, scripts, and
provider executables before launch. Linux executes sealed inherited
descriptors; macOS uses authenticated private snapshots with retry-safe
rematerialization at the spawn boundary.
- Persist canonical ACPX and OpenCode launch-profile digests, reject
drift across fresh recovery, and make recovery failures sticky.
- Close and journal unsafe ACPX active-turn recovery before any provider
bootstrap or reconnect.
- Add managed provider fields to the Runner configuration UI and
permission projection.
- Preserve the default-off `enableNativeRunner` experimental flag.

## Verification

- `pnpm -r typecheck`
- `pnpm build`
- Focused managed server, database, CLI, Runner TypeScript, Rust,
Claude, AgentCore, ACPX, OpenCode, process-supervisor, and
durable-recovery tests passed.
- `cargo test -p paperclip-runner-core --lib --locked` (160 tests)
- `cargo check --workspace --all-targets --locked`
- Native Codex integration tests passed (60 tests); native provider
tests passed (7 tests); server native-runtime tests passed (87 tests).
- Verified-launch replacement, nested-spawn retry, exact-version,
profile-drift, sticky-failure, and no-bootstrap active-recovery tests
passed.
- `git diff --check`
- The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust
workspace lockfile adds the approved `rustix` dependency used for safe
descriptor handling while `#![forbid(unsafe_code)]` remains enabled.

## Risks

- The provider APIs can change while they are in beta. Exact
qualification and fail-closed recovery checks limit drift.
- Remote cleanup can fail after a partial create. Durable ownership
inventories and retry-safe deletion preserve recovery state.
- Migration 0237 adds profile tables. The generated migration and
snapshot pass the repository migration checks.
- Managed execution can incur provider cost. Positive default spend caps
and explicit retention acknowledgement limit accidental use.
- An interrupted AgentCore invocation without final metadata is
conservatively charged to its active session ceiling. This can overstate
cost, but cannot undercount it; later work requires an explicit budget
increase.
- Linux qualified launches use sealed memory descriptors. macOS lacks
executable-descriptor APIs, so the runner uses owner-only private
snapshots and minimizes linked-path lifetime; hostile same-UID processes
remain outside the documented local-host trust boundary.
- The global Runner feature remains default-off.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with tool use, code execution, and subagent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 00:48:30 -05:00
Dotta 84bedd4ca1
feat(runner): activate qualified OpenCode and ACPX providers (#12691)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner is the experimental native runtime for governed
agent work.
> - The runtime contracts already describe Codex, OpenCode, and ACPX
providers.
> - The merged control plane still rejected OpenCode and ACPX for new
runner agents.
> - Runnerd also selected only the Codex provider implementation.
> - This pull request activates the qualified OpenCode and ACPX paths
from the form to runnerd.
> - The benefit is one durable runner path with provider-specific
permissions and recovery.

## Linked Issues or Issue Description

Refs #12685

**Subsystem affected**

This change affects the runner package, server orchestration, adapter
configuration, and UI configuration.

**Problem or motivation**

Paperclip Runner stores provider contracts for OpenCode and ACPX. New
agents cannot select those providers. Runnerd cannot execute those
stored provider descriptors. The UI also shows only Codex.

**Proposed solution**

Accept the qualified OpenCode 1.18.17 profile and the fixed ACPX Claude
and Codex profiles. Route them through runnerd. Keep provider selection,
model selection, permissions, credentials, events, and recovery inside
closed provider-specific boundaries.

**Alternatives considered**

One option was to keep the contracts dormant. That option leaves stored
configuration and runtime behavior out of sync. Another option was to
enable every ACPX agent. That option is not safe because Pi does not yet
have the same verified launch path.

**Roadmap alignment**

This change supports the completed cloud and sandbox agent milestone. It
also supports self-healing runs and governed agent execution. It does
not add a new roadmap surface.

## What Changed

- Add one server profile resolver for Codex, OpenCode, and qualified
ACPX descriptors.
- Keep `adapterConfig` as the provider and permission authority for
fresh runs.
- Add Paperclip Runner provider, ACPX agent, and provider-specific
permission controls to the UI.
- Reset the model to a compatible qualified value when the provider
changes.
- Route Codex, OpenCode, and ACPX through the durable runnerd provider
selector.
- Add a durable ACPX executor with bounded state, recovery, events, tool
receipts, and identity checks.
- Remove Codex labels from OpenCode events, results, evidence, and
recovery diagnostics.
- Pass only provider-specific credential names to child processes.
- Keep ACPX Pi unavailable and reject it before process launch.
- Keep the existing Paperclip Runner experimental flag unchanged.

## Verification

- `pnpm exec vitest run
packages/paperclip-runner/src/backends/native-backend-factory.test.ts
packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts
packages/adapters/codex-local/src/ui/build-config.test.ts
ui/src/adapters/codex-local/config-fields.test.tsx
server/src/__tests__/adapter-registry.test.ts
server/src/__tests__/adapter-routes.test.ts
server/src/__tests__/agent-adapter-validation-routes.test.ts
server/src/__tests__/company-portability.test.ts
server/src/services/native-runtime/runtime-mode.test.ts
server/src/services/native-runtime/native-session-executor.test.ts
server/src/services/heartbeat-runner-provider-config.test.ts`
- The focused TypeScript, server, and UI suites passed 274 tests.
- `cargo test -p paperclip-runner-core --test native_provider_backend`
- The executable native provider integration suite passed 4 tests.
- `cargo test -p paperclip-runner-core --lib`
- The Rust unit suite passed 91 tests.
- `pnpm -r typecheck`
- `pnpm check:token-gates`
- `pnpm build`
- `git diff --check codex/runner-parity-task-runtime...HEAD`

## Risks

- This changes provider process selection and durable recovery. The
experimental flag still gates every fresh Paperclip Runner run.
- OpenCode requires a model in `provider/model` form and stays pinned to
version 1.18.17.
- ACPX accepts only exact Claude and Codex profile versions and models.
Pi stays unavailable.
- ACPX steering stays unavailable and reports that limit through the
driver capabilities.
- Child processes receive explicit environment allowlists. They do not
inherit the full server environment.
- This pull request has no database migration.

## Model Used

OpenAI Codex, GPT-5, with tool use, code execution, and subagent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 21:54:30 -05:00
Dotta 72b9f92d76
fix(runner): restore task runtime parity (#12685)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task view shows a running agent and lets an operator guide that
agent.
> - The merged runner stack lost parts of the accepted task experience.
> - Native event errors could hide current reasoning from the operator.
> - Queued message steering had no server route on `master`.
> - This pull request restores the task-runtime behavior and keeps the
runner experimental gate.
> - The benefit is a visible and steerable native run with durable
fallback behavior.

## Linked Issues or Issue Description

**What happened?**

The task view could stop showing current runner reasoning. The steering
action also failed because the server route was absent. Runner
instruction files were not declared as supported.

**Expected behavior**

The task view must show current provider activity. It must use the live
log when durable native events are empty or unavailable. The operator
must be able to steer a queued message into the active native turn.

**Steps to reproduce**

1. Enable the Paperclip Runner experimental setting.
2. Start a native runner task.
3. Open the task view while the run emits reasoning.
4. Queue a message and select the steering action.

**Paperclip version or commit**

The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`.

**Additional context**

Related closed work: Refs #12592.

## What Changed

- Restored the queued-comment steering route for active native sessions.
- Added durable and queue-bound steering acknowledgements for safe
retries.
- Restored runner instruction bundle support.
- Added live-log fallback when native events are empty or unavailable.
- Restored the compact live reasoning ticker in the task view.
- Added a visible temporary-unavailable state when both activity sources
fail.
- Kept the unified Paperclip Runner experimental gate unchanged.

## Verification

- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm -r typecheck`
- `pnpm check:token-gates`
- `pnpm build`
- Seven focused test files passed with 140 tests.
- The final steering regression file passed with 12 tests.
- The broad local test run reached unrelated workspace, port, and shared
database failures. The changed-area tests remained green.

## Risks

- The steering route changes queue and run records in one transaction.
Tests cover stale targets, unavailable sessions, lost responses, and
wrong-queue acknowledgements.
- Native events remain the primary transcript source. The live log is
used only when event data is absent or its poll fails.
- The experimental gate still hides and rejects the runner when the
setting is off.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with tool use, code execution, and subagent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documentation change is required for this regression repair
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 21:17:54 -05:00
Nicky Leach b4f302d040
feat(grok-local): copy a refreshed sandbox credential back to the host (#12696)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip runs agents through provider-specific adapters in local
and remote environments
> - A remote Grok run can refresh its credential inside its sandbox
> - The host copy can become stale when teardown discards that refreshed
credential
> - This pull request copies the refreshed credential back through a
locked, fail-closed teardown path
> - The benefit is that later Grok runs can use the refreshed host
credential without another login

## Linked Issues or Issue Description

Refs: #12618

**Agent or provider**

Grok local adapter.

**Why this adapter is useful**

A remote Grok run can refresh its access token during a run. Copying the
refreshed credential back to the host keeps later runs ready to use.

**How the agent is invoked**

Paperclip invokes the Grok local adapter through its remote subscription
run path. The adapter stages the company Grok home as a sandbox asset.

The change adds a copy-out step on the teardown path.

## What Changed

- `grok-auth-merge-decision.cjs` adds a host predicate in its own
process. It compares the whole `<issuer>::<uuid>` identity key of the
two files. It reads `expires_at` as an ISO-8601 string, an epoch-seconds
number, or an epoch-milliseconds number. It exits 10 to use the source,
20 to keep the destination, 21 when the expiry shape is unreadable, and
22 when the source expiry sits more than 400 days after the host clock.
It fails closed in every unclear case: an unusable side, a different
identity, an absent expiry, a tie, an unreadable expiry, and an
implausible expiry all keep the destination.
- `grok-auth-merge-decision.ts` adds a wrapper that runs the predicate
and maps the exit code to a typed result.
- `grok-auth-copyback.ts` adds `copyBackGrokAuth({ hostHomeDir,
readSandboxAuth, log, env })`. It locks on `hostHomeDir` with
`withDirectoryMergeLock`, stages the sandbox bytes into a private `0600`
temporary file, runs the predicate, and installs the file with an atomic
rename in the same directory. It keeps no backup of the displaced
credential. It leaves no temporary file on the success path, the keep
path, or an error path. On an error it logs the `errno` code only, then
re-throws.
- `execute.ts` adds a `restore` callback to the Grok `home` asset. The
callback takes the destination from
`resolveManagedGrokHomeDir(process.env, agent.companyId)`, never from
`env.GROK_HOME`. A copy-out failure does not fail the run.
- `package.json` updates the `build` script to copy
`grok-auth-merge-decision.cjs` into `dist/server/`, because `tsc` does
not copy a `.cjs` file.

**The credential shape this predicate reads**

A redacted sample of a real vendor credential answered four structural
questions. The answers hold no credential bytes, no account identifier,
no file path, and no timestamp value.

1. `expires_at` is present.
2. `expires_at` sits inside the value object, under the
`<issuer>::<uuid>` key. It is not a top-level field.
3. `expires_at` is an ISO-8601 string. It carries UTC time with a
trailing `Z` and six fractional-second digits.
4. A normal run rewrites `auth.json`. The value object carries a
`refresh_token` next to `expires_at`, so the client refreshes the access
token and rewrites the file.

## Verification

- [x] `pnpm vitest run packages/adapters/grok-local` — 114 tests in 12
files pass.
- [x] `pnpm --filter @paperclipai/adapter-grok-local typecheck` — clean.
- [x] `pnpm --filter @paperclipai/adapter-grok-local build` — succeeds,
and `dist/server/grok-auth-merge-decision.cjs` exists after the build.
- [x] Continuous integration is green on every check.

## Risks

The predicate keeps the host credential when identity, expiry, file
access, or freshness data is unclear. The copy-out path can log an error
and leave the run successful when it cannot install the refreshed
credential. The atomic rename and directory lock protect the host file
from partial writes and concurrent copy-out actions.

## Model Used

OpenAI GPT-5, current deployment. The exact runtime version and context
window are not exposed to this agent. The model used tool calls and code
inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 18:27:00 -07:00
Dotta 8f9f850c20
fix: limit plan-to-auto transition to plan confirmation (#12695)
<!-- This pull request uses ASD-STE100 Simplified Technical English. -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The issue thread controls plan review and agent work modes.
> - A user can accept a full plan or confirm a smaller checkbox action.
> - Only full plan acceptance must start automatic agent work.
> - The current transition did not check the interaction kind.
> - This pull request limits the transition to an accepted plan
confirmation.
> - The benefit is a safe and clear start of agent work after plan
approval.

## Linked Issues or Issue Description

**What happened?**

An accepted confirmation that targeted a plan could change an issue from
planning mode to standard mode. This included a checkbox confirmation. A
checkbox action is not approval of the full plan.

**Expected behavior**

Only acceptance of a current full-plan confirmation starts automatic
agent work. Other interaction kinds and rejected confirmations keep the
current work mode.

**Steps to reproduce**

1. Put an issue in planning mode.
2. Create a checkbox confirmation that targets the current plan
revision.
3. Accept the checkbox confirmation.
4. Observe that the issue enters standard mode before this fix.

**Paperclip version or commit**

The problem was present on `master` before this change.

**Deployment mode**

The problem is in the core server logic and is not deployment-specific.

## What Changed

- Require a full `request_confirmation` interaction before plan
acceptance starts automatic work.
- Add service tests for acceptance, rejection, stale interaction kinds,
and unchanged standard-mode behavior.
- Check the route activity log for the planning-to-standard mode change.
- Document the plan acceptance transition in the V1 contract.

## Verification

- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts
server/src/__tests__/issue-thread-interaction-routes.test.ts` passes 140
tests.
- `pnpm -r typecheck` passes.
- `pnpm build` passes.
- `pnpm test:run` was also started. Unrelated workspace-runtime tests
failed because fixed local runtime ports were occupied or offset on the
shared host. The same failures reproduce alone. The changed test files
pass alone.

## Risks

- Risk is low. The change adds one interaction-kind guard to the
existing transition.
- A full accepted plan confirmation still changes planning mode to
standard mode and an eligible review issue to todo in one transaction.
- Checkbox confirmations, questions, rejection, and standard-mode issues
keep their previous behavior.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5, reasoning, tool use, and code execution. The
runtime does not expose the exact model suffix or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 17:18:42 -05:00
Dotta 4b6de5327e
Remove cheap model profiles (#12683)
## Thinking Path

> - Paperclip manages agents that use different model providers and
adapters.
> - Paperclip must keep agent execution rules clear and predictable.
> - The cheap-model profile added a second execution mode across
adapters, task recovery, APIs, and the UI.
> - That mode increased configuration and recovery complexity.
> - This pull request removes the cheap-model profile as a product
feature.
> - The benefit is one model-selection path for normal work and recovery
work.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This change simplifies model selection across agent configuration, task
execution, recovery, and adapter capabilities.

**Current behavior**

Paperclip exposes cheap-model profiles in adapter metadata, agent
runtime configuration, task overrides, recovery rules, APIs, and the
board UI. Recovery work can select a different model profile from the
agent's configured model.

**Proposed behavior**

Paperclip uses the agent's configured model for normal work and recovery
work. Status-only recovery stays limited to coordination work. The API
rejects legacy model-profile configuration. A migration removes stored
model-profile values from existing agent, issue, and historical revision
records.

**Reason and benefit**

One model path reduces configuration, API, UI, and recovery complexity.
It also prevents status recovery from becoming a separate product-level
model-routing feature.

**Breaking changes**

This change removes model-profile fields and adapter capability
metadata. Existing stored model-profile values are removed by an
idempotent migration. The validators reject new legacy profile values
with clear errors.

## What Changed

- Removed model-profile types, adapter capabilities, API fields, and
model selection logic.
- Removed cheap-model controls from agent and task UI surfaces.
- Kept status-only recovery limited to coordination context while normal
continuations use the configured agent model.
- Added an idempotent migration that removes stored model-profile values
from agents, issues, and configuration revisions without changing issue
update timestamps.
- Updated tests and product documentation for the single-model behavior.

## Verification

- `pnpm check:token-gates` passes.
- `pnpm -r typecheck` passes.
- `pnpm build` passes.
- `pnpm test:run` completed with 5,607 passing tests and 8
environment-sensitive failures in unrelated fixed-port and
database-deadlock suites. The same failures repeated in an isolated
rerun. CI is the final clean-room result.

## Risks

- This is an intentional breaking change for clients that send
model-profile fields.
- The migration changes legacy agent, issue, and configuration-revision
JSON. It is idempotent and preserves unrelated fields and issue update
timestamps.
- The change is cross-cutting because the removed feature existed in
adapters, shared contracts, the server, plugins, and the UI.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The
runtime did not expose the context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 14:57:38 -05:00
Dotta 1ab159d3a7
feat(apps): consolidate connector management (#12684)
Completes the post-managed-OAuth connector lifecycle, Paperclip Cloud provisioning defaults, governed test flows, and consolidated Apps UI.\n\nCo-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-01 14:55:35 -05:00
Dotta 141f202e40
Clean up experimental settings features (#12681)
## Thinking Path

> - Paperclip is the open source app that people use to manage AI agents
for work.
> - Instance settings control optional product features and developer
tools.
> - The experimental settings page mixed active experiments, internal
tools, and old recovery controls.
> - Some workspace links also used the selected company instead of the
workspace owner.
> - These problems made settings hard to scan and could send users to
the wrong company route.
> - This pull request removes old controls, groups developer settings,
and resolves workspace links from workspace data.
> - The benefit is a smaller settings surface and correct workspace
navigation.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the instance experimental settings page, task watchdog
controls, dependency wake recovery, and execution workspace routes.

**Current behavior**

The settings page shows old recovery controls and mixes product
experiments with internal developer settings. Task watchdogs require an
extra feature flag. Some direct workspace links use the current company
prefix instead of the company that owns the workspace.

**Proposed behavior**

Remove the old task recovery experiment and its unused API surface. Make
task watchdog controls available without the removed flag. Put worktree
execution and managed environment controls in the developer section.
Resolve direct workspace links from the workspace owner and reject a
company prefix that does not own the workspace.

**Reason and benefit**

The smaller settings page is easier to understand. The server keeps only
the dependency wake backstop that it still uses. Workspace links open
under the correct company route.

**Breaking changes**

This removes the experimental issue graph recovery preview and run
endpoints. It also removes the task watchdog feature flag. Task watchdog
data and dependency wake behavior remain available.

## What Changed

- Removed the old task watchdog and issue graph recovery feature flags.
- Removed the old issue graph recovery preview, run controls, API
contracts, and unused recovery implementation.
- Kept resolved dependency wakes as the scheduler backstop.
- Grouped product experiments and Paperclip developer settings on the
instance settings page.
- Made task watchdog controls available without an extra experimental
flag.
- Added owner-aware redirects and company checks for execution workspace
routes.
- Hid the false stopped-state badge while a workspace has no active
runtime state.
- Updated focused server and UI tests for the new behavior.

## Verification

- `pnpm check:token-gates`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm test:run` completed with 5,620 passing tests and four failures
in unchanged workspace runtime port tests. The same four failures repeat
when the two files run alone.
- The complete GitHub CI matrix passed, including all server, serialized
server, build, canary, and end-to-end jobs.

## Risks

- Clients that call the removed experimental recovery endpoints must
stop calling them.
- The route checks depend on workspace detail access. An unknown or
cross-company workspace returns the global not-found page.
- There are no database migrations, lockfile changes, workflow changes,
or design image changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5. The exact deployment ID and context window are
not exposed. Reasoning, tool use, and code execution were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 14:23:05 -05:00
Nicky Leach ed3559dd21
feat(server): split the Sentry DSN into front-end and backend variables (#12678)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip reports server and browser errors through optional Sentry
monitoring
> - One environment variable sends both error types to one Sentry
project
> - Operators need separate control for browser and server error data
> - This pull request adds specific variables and keeps the existing
variable as a fallback
> - The benefit is separate monitoring without breaking current
deployments

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The Sentry configuration for server and browser monitoring uses one
environment variable.

**Subsystem affected**

Cross-cutting (multiple of the above)

**Current behavior**

`SENTRY_DSN` supplies the server and browser clients. Both clients
therefore report to the same Sentry project.

**Proposed behavior**

`SENTRY_DSN_FRONTEND` supplies the browser client. `SENTRY_DSN_BACKEND`
supplies the server process. `SENTRY_DSN` remains a fallback for either
component.

**Reason and benefit**

Operators can send browser and server errors to separate Sentry
projects. Operators can also activate only one component.

**Breaking changes**

None. Existing deployments can continue to use `SENTRY_DSN`.

## What Changed

- Add `resolveSentryDsns(env)` and use it in the server and browser
configuration paths.
- Add precedence, empty-string, fallback, and route tests.
- Update the README, observability guide, and stale code comments.
- Log one warning when the server uses the legacy fallback without
exposing a DSN value.

## Verification

- `pnpm vitest run --project server sentry-dsn` — 8 tests pass.
- `pnpm vitest run --project server auth-routes` — 21 tests pass.
- The earlier run of the three targeted suites passed 40 tests.
- `tsc --noEmit` passes for the files in this diff.
- All required GitHub Actions checks pass, including the full
continuous-integration suite.

## Risks

The main risk is an incorrect environment variable precedence rule. Unit
tests cover specific values, empty strings, and legacy fallback
behavior. The existing `SENTRY_DSN` path remains compatible.

## Model Used

OpenAI Codex — GPT-5, current runtime, tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 11:02:04 -07:00
Dotta 86ebdf842e
fix(runner): keep agents running when app connections expire (#12670)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents can receive governed access to connected apps through the
runtime MCP gateway.
> - A connected app can become unavailable when its sign-in expires or
its health state needs attention.
> - The native runner treated that optional app state as a fatal runtime
setup error.
> - One unavailable app could therefore stop all unrelated agent work.
> - This pull request removes the fatal dependency and keeps the
available app assignment immutable.
> - The benefit is that an agent can continue its work while the stream
tells the user which app needs reconnection.

## Linked Issues or Issue Description

**What happened?**

An agent could not start a native run when one assigned app connection
was disabled, degraded, failed, or missing its secret. Runtime context
creation or MCP delivery threw an error before the agent could do
unrelated work.

**Expected behavior**

The run must continue without the unavailable app. Healthy assigned apps
must remain available. The stream must explain which app needs
reconnection. A changed assignment must not give a native run new access
after its immutable context is captured.

**Steps to reproduce**

1. Assign an MCP app connection to a Paperclip Runner agent.
2. Set the connection to a state that needs attention, such as
`degraded`.
3. Start a task run for that agent.
4. Observe that native runtime setup fails before the agent starts.

**Paperclip version or commit**

Reproduced from `ee2a19062`. The branch is rebased on `dda4dff64`.

**Deployment mode**

Local development from source with embedded Postgres.

No matching public issue or open pull request was found in the GitHub
search.

## What Changed

- Filter unavailable assigned app connections from the immutable native
runtime MCP snapshot.
- Keep healthy assigned connections and their tools in the snapshot.
- Replace the fatal native MCP availability check with an optional
stream warning callback.
- Withhold MCP delivery when the current assignment digest does not
match the captured native context.
- Prevent a warning delivery failure from stopping the agent run.
- Add regression tests for disabled, degraded, mixed healthy and
unavailable, and assignment-drift cases.

## Verification

- `pnpm exec vitest run
server/src/services/native-runtime/runtime-context.test.ts
server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts` passes with
8 tests.
- `pnpm -r typecheck` passes.
- `pnpm check:token-gates` passes.
- `pnpm build` passes.
- `pnpm test:run` was attempted. Unrelated workspace runtime and
port-exposure tests failed on this macOS host. The same files also
failed when run without the changed MCP tests. The changed MCP tests
remained green. Clean GitHub CI is the final full-suite check.

## Risks

- Low migration risk. This change has no schema or API contract
migration.
- An unavailable app is absent from the run MCP surface until it is
reconnected and a later run captures it again.
- Assignment drift fails closed. The agent keeps running, but the
changed gateway is not delivered.
- This pull request does not auto-block the issue before the agent
decides that the app is required. It emits reconnect guidance in the
stream. The existing connection-request interaction remains the path for
a required app.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, `gpt-5.6-sol`, with high reasoning, repository tools,
code execution, and browser automation. The runtime did not expose the
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 10:48:18 -05:00
Dotta 14c7efa068
fix(workspaces): enable UI hot reload by default (#12612)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed worktrees can run a Paperclip development server for each
task
> - The managed runtime used the built UI when its service did not set
the UI development middleware option
> - This made new UI source changes require a manual build instead of a
hot reload
> - The runtime must supply the development default while it must keep
an explicit operator choice
> - This pull request enables the UI development middleware for new
managed Paperclip development services
> - The benefit is that UI edits appear in the managed worktree browser
without a manual build

## Linked Issues or Issue Description

**What happened?**

A new managed Paperclip development worktree served the built UI by
default. An operator had to set `PAPERCLIP_UI_DEV_MIDDLEWARE=true`
before UI source changes could hot reload.

**Expected behavior**

New managed Paperclip development worktrees must enable the UI
development middleware by default. An explicit
`PAPERCLIP_UI_DEV_MIDDLEWARE=false` value must continue to disable it.

**Steps to reproduce**

1. Start a managed Paperclip development service without
`PAPERCLIP_UI_DEV_MIDDLEWARE`.
2. Open its UI.
3. Change a UI source file.
4. Observe that the browser does not receive the change until the UI is
built again.

**Paperclip version or commit**

This was reproduced on `317394456` from `master`.

**Deployment mode**

Local development with a managed worktree runtime.

## What Changed

- Set `PAPERCLIP_UI_DEV_MIDDLEWARE=true` for managed `paperclip-dev`
services when the service does not set a value.
- Keep explicit service values, including `false`.
- Add a regression test and document the default and the opt-out.

## Verification

- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/workspace-runtime.test.ts -t "enables UI dev middleware by
default"`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm test:run` completed with 5,397 passing tests. Four existing
runtime-port tests could not use ports `42000` and `52000` because a
live managed runtime owns those ports on this host. The new regression
test passed separately.

## Risks

- Risk is low. The change applies only to managed services named
`paperclip-dev`.
- A service can keep the built UI by setting
`PAPERCLIP_UI_DEV_MIDDLEWARE=false`.
- There is no database or API contract change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, `gpt-5.6-sol`, hosted Codex context window, high
reasoning, tool use, code execution, and multi-file repository editing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-01 10:06:53 -05:00
Dotta ee2a190626
Unify Paperclip Runner experimental controls (#12666)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner is an experimental execution adapter.
> - The adapter and its required sandbox ingress had separate settings.
> - A user could enable one setting and still have an unusable runner
configuration.
> - The runtime already makes one durable native or legacy decision for
each run.
> - This pull request uses that runtime decision for ingress
authorization.
> - The benefit is one clear opt-in with safe recovery for existing
native runs.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the experimental settings and transport authorization for
Paperclip Runner.

**Subsystem affected**

Cross-cutting. This change affects the React settings UI, shared
settings contracts, adapter utilities, and server runtime selection.

**Current behavior**

Settings shows separate Paperclip Runner and Runner Preview Ingress
controls. A user can enable the runner but leave required sandbox
ingress disabled.

**Proposed behavior**

Settings shows only Paperclip Runner. Its native runtime decision also
authorizes provider WebSocket ingress when the execution target requires
it. A persisted native run keeps its recovery transport after the
setting is disabled.

**Reason and benefit**

Paperclip Runner is one experimental capability. One opt-in removes an
invalid partial configuration and makes the rollout boundary easier to
understand.

**Breaking changes**

The Runner Preview Ingress card is removed. The old
`enableRunnerPreviewIngress` key remains accepted in stored settings and
managed configuration, but it has no server runtime effect. The public
adapter-utils input remains compatible through a deprecated alias.

**Additional context**

Refs: #12638, #12641, #12656.

## What Changed

- Removed the separate Runner Preview Ingress card from Experimental
Settings.
- Made resolved native runtime selection authorize required provider
ingress.
- Preserved ingress recovery for persisted native runs after the rollout
flag is disabled.
- Kept the old settings key and adapter-utils input as deprecated
compatibility contracts.
- Added focused UI, runtime policy, transport, stored-settings, and
managed-config regression tests.
- Updated deployment documentation and feature descriptions.

## Verification

- GitHub Actions will run typecheck, tests, build, policy, and browser
shards.
- Focused tests cover the single settings control, runtime
authorization, fail-closed transport selection, the deprecated public
input, and old managed configuration.
- No local tests were run, per the maintainer request to use GitHub
Actions for verification.
- `git diff --check` passes.

## Risks

Low to moderate risk. The effective ingress gate changes from a separate
stored flag to the resolved native run decision. Fresh runs still
require `enableNativeRunner`. Persisted native runs remain recoverable.
Legacy adapters never receive ingress authorization.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 09:21:23 -05:00
Dotta 1955b0e2d8
Gate Paperclip Runner setup behind an experimental flag (#12656)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent adapters control how Paperclip starts and resumes an agent
runtime.
> - Paperclip Runner is an experimental Rust runtime and must stay
opt-in.
> - The server already rejected new runner selections when the flag was
off.
> - Some setup and onboarding views did not enforce the same boundary.
> - This pull request exposes the existing flag and applies it to every
new setup path.
> - The benefit is a safe rollout with unchanged legacy onboarding and
recoverable existing native runs.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves experimental adapter selection in Settings, onboarding,
new-agent setup, invite setup, and company import.

**Subsystem affected**

Cross-cutting: the React UI and the server onboarding seed service.

**Current behavior**

The server defaulted Paperclip Runner to off, but Settings did not
expose the flag. First-run onboarding could show the runner after
opt-in. A direct new-agent URL and some setup pickers could also reveal
native runner configuration before the availability check completed.

**Proposed behavior**

Settings has a default-off Paperclip Runner toggle. Explicit agent
configuration shows the runner only after the server reports that the
flag is enabled. First-run and invite onboarding always use legacy
adapters. Existing native agents and runs remain readable and
recoverable.

**Reason and benefit**

This keeps the experimental runtime out of normal onboarding. It also
gives administrators one clear opt-in before users can create a native
runner agent.

**Breaking changes**

None. Legacy adapter selection and execution stay unchanged. Existing
native records remain available.

## What Changed

- Added the Paperclip Runner opt-in to Experimental Settings.
- Refreshed adapter availability after the setting changes.
- Kept UI and server-seeded onboarding on legacy adapters.
- Made native runner choices fail closed in new-agent, invite, and
import setup.
- Preserved edit and recovery behavior for existing native agents and
runs.
- Added focused regression tests for flag-off and flag-on behavior.

## Verification

- GitHub Actions will run the repository test, typecheck, build, and
policy gates.
- Focused tests cover Settings, onboarding, agent creation, invite
setup, import setup, and server-seeded onboarding.
- No local test suite was run, per the maintainer request to use GitHub
Actions for verification.
- `git diff --check` passes.

## Risks

Low risk. The change narrows new adapter selection only. The server
remains the final enforcement point. Existing native records do not
depend on the current flag value for read or recovery behavior.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5, with reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 05:57:40 -05:00
Dotta 1ed29abaa6
fix(runner): harden dormant provider boundaries (#12654)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner currently enables only the Codex production path.
> - The package also contains dormant OpenCode and ACPX provider
boundaries.
> - Dormant boundaries must still fail safe before later activation
work.
> - Provider children must not inherit unrelated server secrets or host
homes.
> - Permission defaults must require interaction instead of broad
automatic approval.
> - This pull request hardens those boundaries without activating them.
> - The benefit is a safer base for later provider-specific runnerd
work.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the inactive OpenCode and ACPX provider boundary in
Paperclip Runner.

**Subsystem affected**

The adapter permission contract, Runner provider environment, and native
execution input builder.

**Current behavior**

Dormant OpenCode code can inherit the full server environment. Its
default permission mode allows operations. ACPX also defaults to broad
approval. The provider guard can accept inherited object property names.

**Proposed behavior**

Use exact provider identifiers. Use interactive defaults. Allow only
required OpenCode environment keys. Reject invalid proxy permission
modes.

**Reason and benefit**

This reduces accidental authority and secret exposure before future
provider activation.

**Breaking changes**

No production provider is activated. Codex runtime selection and Codex
credential-home discovery do not change. Dormant OpenCode and ACPX
callers that omit permission modes now receive safer defaults.

## What Changed

- Change dormant OpenCode and ACPX permission defaults to interactive
modes.
- Reject prototype property names as provider identifiers.
- Default dormant ACPX input to the qualified Codex agent profile.
- Add an explicit OpenCode runner environment allowlist.
- Exclude host homes, server credentials, database values, and Node
injection options.
- Add a fail-closed OpenCode proxy permission parser.
- Add focused tests for defaults, filtering, and invalid values.

## Verification

GitHub Actions must run:

- Adapter utility tests.
- Paperclip Runner tests, type checks, and build.
- Server native runtime tests.
- Repository test, type-check, build, policy, and security gates.

No local test command was run. The repository owner requested
GitHub-only verification.

## Risks

Future OpenCode credential providers must add required variables to the
allowlist through review. The safer defaults can pause dormant internal
scenarios that relied on implicit broad approval. Production Codex
behavior is unchanged.

## Model Used

OpenAI Codex with the GPT-5 agent model. The work used high reasoning,
repository inspection, tool use, and parallel security review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 04:46:36 -05:00
Dotta 131f5c4065
feat(runner): add administration and observability (#12641)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Administrators need bounded controls for experimental native
execution.
> - The lower stack adds remote Codex execution and the task workspace.
> - Operators need to configure Codex safely and inspect provider
traces.
> - Unsupported providers must not appear as runnable choices.
> - This pull request adds Codex-only administration and observability.
> - The benefit is a default-off operational surface for production
diagnosis.

## Linked Issues or Issue Description

Refs #12640.
Refs #12616.
Refs #12352.

**Subsystem affected**

Agent configuration, instance experimental settings, run ledger,
provider trace inspector, and administrator actions.

**Problem or motivation**

The native runner lacks one safe operator surface for Codex permissions,
lifecycle, raw trace capture, and run inspection. The integration branch
also contains provider choices that the production backend cannot
execute yet.

**Proposed solution**

Expose only the qualified Codex controls. Keep Paperclip Developer Mode
and runner preview ingress off by default. Gate raw trace actions by
administrator access and existing trace authorization.

**Alternatives considered**

Exposing unfinished providers would create configurations that fail at
runtime. Always-on tracing would increase sensitive data and storage
risk.

**Roadmap alignment**

This work supports governed Cloud and Sandbox agents and production
diagnostics.

## Stack

- Base PR: #12640.
- Lower PRs: #12639 and #12638.
- This PR contains only its 54-file administration and observability
delta.
- This is the final feature PR in the Codex production stack.

## What Changed

- Added Codex-only Paperclip Runner permission and lifecycle controls.
- Added bounded warm idle configuration.
- Kept the provider field fixed to Codex.
- Added administrator-only one-run raw trace requests.
- Added a persistent future-run raw trace toggle.
- Added trace status, metadata, ledger, and canonical runner inspection.
- Added JSON-RPC request-origin grouping and finalization lineage.
- Restored the stateful PRP transcript parser and focused projection
tests required by trace inspection.
- Added default-off Paperclip Developer Mode.
- Added Honeycomb run links for authorized developer mode.
- Disabled the legacy operational skill for `paperclip_runner`.
- Did not expose OpenCode, ACPX, Pi, Claude Managed, or AWS runner
choices.
- Did not change migrations, workflows, dependencies, or
`pnpm-lock.yaml`.

## Verification

- GitHub Actions will run UI tests, server tests, repository typecheck,
build, browser tests, security, and policy gates.
- Tests cover Codex configuration defaults and bounds, administrator
trace actions, persistent settings, ledger inspection, trace lineage,
and Honeycomb links.
- Existing server trace authorization and retention tests remain the
backend authority.
- Local tests were not run. The requested verification policy uses
GitHub Actions for this series.
- `git diff --check runner/task-workspace-experience...HEAD` passes.
- The delta contains 54 files.

## Risks

- Raw provider traces can contain sensitive provider data.
- Existing server authorization controls access, reveal, download,
retention, and deletion.
- The UI gates trace actions by administrator access and developer mode.
- All new instance settings remain off by default.
- Fresh Paperclip Runner configuration remains Codex-only.
- Direct adapters and legacy task behavior do not change in this PR.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode,
repository tools, GitHub tools, and parallel code-audit agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 03:41:23 -05:00
Dotta 0a422fda52
feat(runner): add remote execution substrate (#12638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner gives native runs a durable and governed execution
path.
> - The current native path runs on the control-plane host.
> - Remote environments need an authenticated execution-target contract.
> - The contract must not change direct adapters or enable new runtimes
by default.
> - This pull request adds the remote execution substrate and Daytona
ingress.
> - The benefit is a bounded base for later remote runner transport
work.

## Linked Issues or Issue Description

Refs #12616.
Refs #12352.

**Subsystem affected**

Cross-cutting. This change touches runner transport, server
orchestration, plugin contracts, and shared settings.

**Problem or motivation**

Native execution cannot resolve an authenticated runner ingress through
a remote environment. The server also lacks one provider-neutral
contract for remote execution targets.

**Proposed solution**

Add a default-off runner preview ingress capability. Add
transport-neutral runner connectivity. Add remote execution target and
lifecycle handling. Add a Daytona ingress implementation with redacted
credentials.

**Alternatives considered**

A provider-specific server path would duplicate orchestration and
authorization. A public endpoint without an environment contract would
weaken the trust boundary.

**Roadmap alignment**

This work supports the Cloud and Sandbox agents milestone. It also
supports self-healing runs and governed tool access.

## What Changed

- Added execution-target traits for local, SSH, and sandbox
environments.
- Added plugin RPC contracts for runner ingress endpoints.
- Added authenticated Daytona preview ingress.
- Added transport-neutral PRP outbound connections.
- Added remote runner artifact verification and fail-closed provider
selection.
- Added bounded native session resume, cancellation, and lifecycle
recovery.
- Preserved Codex-only selection for fresh experimental runner starts.
- Preserved all direct adapter execution and finalization paths.
- Removed stale Pi provider-pack requirements that security review
rejected.
- Kept the rollout controls off by default.
- Did not change pnpm-lock.yaml, Cargo, database migrations, or GitHub
workflows.

## Verification

- GitHub Actions will run the repository test, typecheck, build,
security, and policy gates.
- Focused tests cover ingress validation, redaction, execution targets,
remote lifecycle, cancellation, resume, and legacy adapter selection.
- Local tests were not run. The requested verification policy uses
GitHub Actions for this series.
- `git diff --check origin/master...HEAD` passes.
- The diff contains 52 files.

## Risks

- Remote execution crosses a trust boundary.
- The implementation validates target capabilities, artifact digests,
provider-pack pins, and connection metadata.
- The feature remains default-off.
- Fresh native selection remains Codex-only.
- Existing direct adapters remain on the legacy path.
- This PR does not yet make remote Codex runnable. The next PR adds the
Rust WSS and TLS transport.

## Model Used

OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode,
repository tools, GitHub tools, and parallel code-audit agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 01:29:06 -05:00
Dotta 51ad751e0b
feat(runner): integrate Codex native execution (#12616)
## Thinking Path

> - Paperclip is the open source control plane for teams of AI agents.
> - Agent runs currently use direct adapters and their established
finalization paths.
> - The new runner package needs one production integration before it
can execute a real provider through the server.
> - That integration must not change direct adapters or expose
unsupported providers.
> - The rollout must also preserve native runs that were already
recorded when the feature flag changes.
> - This pull request adds a default-off, Codex-only native execution
path and its authority boundary.
> - The benefit is a recoverable production vertical slice with explicit
compatibility guards.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting server orchestration and adapter selection.

**Problem or motivation**

The runner package exists, but the server cannot yet start and recover a
governed Codex run through it. A careless integration could also route
existing direct adapters into the native runtime or lose cancellation
and finalization state.

**Proposed solution**

Add a hidden `paperclip_runner` adapter for Codex. Keep it behind the
default-off instance flag. Bind native execution, resume, cancellation,
semantic tool authority, and finalization to the recorded company,
issue, run, and coordinator identities. Leave every direct adapter on
its existing path.

**Alternatives considered**

A multi-provider launch was rejected because only Codex has the complete
production bridge in this series. Replacing direct adapter execution was
rejected because the runner remains experimental.

**Roadmap alignment**

This work supports governed tool access, action attribution, and
self-healing runs. It keeps the integration narrow and default-off.

## What Changed

- Add the Codex-only native session executor and persisted resumption
path.
- Add run-scoped semantic tool projection, authorization, receipts, and
idempotency.
- Add audited native cancellation with durable issue and coordinator
binding.
- Add result fencing so a recorded result cannot reacquire the provider
and run twice.
- Reject fresh runner starts when the rollout flag is off while
preserving recorded native recovery.
- Keep direct adapters outside native status, cancellation, record
creation, and finalization.
- Add focused conformance, recovery, cancellation, status, portability,
and compatibility coverage.

## Verification

- GitHub Actions is the authoritative test environment for this large
stack.
- The PR policy and lightweight stack checks run while this is a middle
PR.
- The full required suite runs when this PR becomes the lowest unmerged
or top PR.
- Greptile will review this exact delta after the branch is pushed.

## Risks

- The main risk is routing a legacy adapter into native execution.
Runtime selection and heartbeat tests cover that boundary.
- The next risk is stale or cross-company cancellation. Durable binding
checks and transactional audit persistence cover it.
- The adapter remains hidden and default-off. Only Codex is admitted.
- There are no database migration, lockfile, or GitHub workflow changes
in this PR.

## Stack

1. [Runner package, SDK, and developer
tools](https://github.com/paperclipai/paperclip/pull/12608)
2. This PR: Codex production server integration
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617)

## Model Used

OpenAI Codex with GPT-5, extended reasoning, repository tools, and
parallel review agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 22:51:17 -05:00