Commit Graph

4186 Commits

Author SHA1 Message Date
Dotta a015ac7a57
fix(ui): allow interrupting queued issue runs (#9725)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The issue detail chat lets operators send messages while an issue
run is live
> - Messages sent during a live run are shown as queued and can expose
an interrupt action
> - The UI only treated running runs as interruptible, even though
queued runs can also own the pending message
> - That mismatch hid the interrupt button after sending a message while
the agent run was still queued
> - This pull request treats queued and running issue runs as
interruptible and preserves the exact target run on optimistic and
persisted comments
> - The benefit is that operators can immediately interrupt the queued
run their message is waiting behind

## Linked Issues or Issue Description

- **What happened:** Sending a message while an issue-owned agent run
was in `queued` state showed the message as pending but did not make the
interrupt action available.
- **Expected behavior:** A message queued behind either a queued or
running issue run should retain that run as its interrupt target and
expose the interrupt control.
- **Steps to reproduce:** Open an in-progress issue with an issue-owned
run still queued, send a chat message, and inspect the queued message
actions.
- **Version/commit:** Reproduced against the pre-change `master` UI
behavior.
- **Deployment mode:** Paperclip board UI with a queued issue execution
run.

## What Changed

- Generalized issue-run resolution from running-only to
queued-or-running interruptible runs.
- Used the interruptible run consistently for optimistic queue metadata,
persisted comment decoration, cancel controls, and targeted
interruption.
- Added a regression test that sends a message behind a queued run and
verifies the exact run is cancelled.

## Verification

- `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx` — 44 tests
passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm check:token-gates` — repository-wide gate currently reports nine
pre-existing `#9627` comment literals; this patch adds no token literals
or gate violations.

## Risks

- Low risk: the behavior change is limited to selecting queued
issue-owned runs as valid interrupt targets in the existing chat flow.
- Cancellation remains targeted by run ID, and the regression test
verifies the queued run ID is preserved through optimistic and persisted
comment states.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI `gpt-5.4` via Codex CLI; context-window size is not exposed by
this runtime; reasoning-enabled with repository, shell, GitHub CLI, and
code-execution tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 16:55:51 -05:00
Dotta 59fb27ff79
feat(inbox): let agents safely tidy user inboxes (#9724)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies and their work
> - The inbox is a per-user attention view, so archiving an item must
not alter the underlying issue, assignment, or status
> - Agents can help responsible users tidy resolved work only when the
action is company-scoped, reversible, policy-controlled, and fully
attributable
> - The database and authorization foundations landed in #9654 and
#9658, but the end-to-end archive routes, audit details, agent workflow
guidance, and operator UI still need to ship together
> - Separate stacked PRs #9659 and #9661 made the complete behavior
harder to review and land as one coherent capability
> - This pull request consolidates the remaining server,
shared-contract, documentation, skill, and UI work on top of current
master
> - The benefit is a single reviewable change that lets agents safely
archive responsible-user inbox items and lets users control or undo that
behavior

## Linked Issues or Issue Description

### Subsystem affected

Cross-cutting inbox management across shared contracts, server
authorization/routes/services, shipped agent skills, and the board UI.

### Problem or motivation

Agents may complete work whose issue remains in the responsible user's
Mine inbox. Existing board-user archive behavior does not provide the
agent-facing policy endpoints, target resolution, heartbeat-run
attribution, typed denials, conservative workflow guidance, or UI needed
for safe agent-managed cleanup.

### Proposed solution

Allow authorized agents to archive or unarchive responsible-user inbox
items under the user's open, allowlist, or disabled policy; preserve
actor/agent/run attribution in issue detail and activity records; expose
policy controls and agent archive attribution in the UI; and document
conservative cleanup rules for agents and PR gardening.

### Alternatives considered

- Reuse generic issue mutation permissions: rejected because inbox state
belongs to a target user and requires user-scoped authorization.
- Automatically archive every completed or closed item: rejected because
completion signals can still require human review or a decision.
- Keep the backend and UI as separate stacked PRs: superseded by this
consolidated PR so the complete user-visible behavior can be reviewed
and verified together.

### Related work

- Builds on merged foundations #9654 and #9658.
- Supersedes the remaining stacked changes in #9659 and #9661.
- `ROADMAP.md` has no overlapping inbox archive or inbox authorization
initiative.

## What Changed

- Added shared inbox-agent policy types and validators plus
company-scoped self-service policy routes and OpenAPI coverage.
- Enabled agent archive/unarchive mutations with responsible-user
targeting, policy enforcement, typed failures, attribution, idempotency,
and detailed activity auditing.
- Returned agent archive attribution in issue detail and documented
reversible inbox cleanup semantics in the implementation spec and
Paperclip skill.
- Added conservative PR-gardening inbox tidy guidance that keeps GitHub
access read-only and avoids archiving work that still needs human
action.
- Added the Profile settings policy control and Issue Properties
attribution/unarchive UI with focused component coverage and narrow-pane
handling.

## Verification

- `pnpm exec vitest run
server/src/__tests__/inbox-archive-routes.test.ts
server/src/__tests__/inbox-agent-policy-routes.test.ts
server/src/__tests__/authorization-service.test.ts
server/src/__tests__/openapi-routes.test.ts
ui/src/components/InboxAgentPolicyControl.test.tsx
ui/src/components/IssueProperties.test.tsx` — 110 passed.
- `pnpm --filter @paperclipai/db exec vitest run
src/inbox-archive-agent-policies-migration.test.ts` — 1 passed.
- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 9 passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm check:token-gates` — changed files are clean; the
repository-wide command currently reports nine unrelated pre-existing
`#9627` literals outside this PR's diff.

## Risks

- Agent inbox mutations broaden an existing endpoint path, so
authorization and target resolution must remain fail-closed; focused
route and authorization tests cover allowed and denied paths.
- Archive state affects only the responsible user's inbox presentation
and remains reversible; it does not mutate issue status, assignment, or
visibility.
- The UI policy defaults to the existing open behavior, while allowlist
and disabled modes can reduce agent access.
- This PR intentionally builds on #9654 and #9658 and contains no new
migration number or modification to an already-applied migration.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.4, medium reasoning, repository tool use,
shell execution, code review, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change
(`feat/inbox-agent-archive-complete`) and contains no internal Paperclip
ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 16:49:18 -05:00
Dotta 89af58b6dc
Clarify task-level model overrides in issue properties (#9710)
## Thinking Path

> - Paperclip is the control plane where operators configure and inspect
AI-agent work.
> - An issue can override the assigned agent's primary model for that
task.
> - The issue-properties UI labeled that per-task setting as `Custom ·
<model>`, which could be read as a property of the model rather than a
replacement of the agent default.
> - Operators need the UI to distinguish the agent's primary model from
an issue-specific override at both the collapsed summary and selection
point.
> - This pull request renames the lane presentation to `Override` and
adds concise provenance text without changing the stored lane value or
adapter configuration behavior.
> - The benefit is that operators can immediately understand which model
will run and why it differs from the agent default.

## Linked Issues or Issue Description

No matching public GitHub issue was found. Related implementation work:
#9700 moved Codex ACPX model configuration to startup and is already on
`master`; #9355 also addresses ACP session-config rejection behavior.
Those PRs concern execution behavior, while this PR is limited to
clarifying the issue-level override UI.

Bug report details:

- **What happened?** The issue-properties Model row displayed `Custom ·
<model>` for a task-level
`assigneeAdapterOverrides.adapterConfig.model`. Operators could misread
`Custom` as describing the model itself and could not see that the value
replaced the agent's primary model for this issue.
- **Expected behavior:** The UI should explicitly identify a task-level
model as an override and explain that it replaces the agent's primary
model for the issue.
- **Steps to reproduce:** Configure an agent with a primary model, set a
different model on an issue, and inspect the Model row and model picker
in Issue Properties.
- **Paperclip version or commit:** Reproduced on the pre-change branch
derived from current `master`.
- **Deployment mode:** Local development or any deployment using the
board UI.
- **Installation method:** Built from source.
- **Agent adapters involved:** Adapter-agnostic; any adapter exposing
model selection.
- **Database mode:** Not database-related.
- **Access context:** Board operator viewing issue properties.
- **Relevant logs or output:** Not applicable; this is a presentation
ambiguity.
- **Relevant config:** An issue-level
`assigneeAdapterOverrides.adapterConfig.model` differing from the
assigned agent's primary model.
- **Additional context:** The internal lane identifier remains `custom`;
only user-facing copy and explanatory text change.
- **Privacy checklist:** No private instance links, internal ticket IDs,
secrets, usernames, or local paths are included.

## What Changed

- Renamed the collapsed issue model label from `Custom · <model>` to
`Override · <model>` and added a provenance tooltip.
- Renamed the model-picker lane from `Custom` to `Override` and added
explanatory subtext at the selection point.
- Updated the Issue Properties component test to assert the new label.
- Added an isolated Storybook fixture for the task model override state
so visual review has a stable target.

## Verification

- `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` — 43
tests passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `git diff --check public-gh/master...HEAD` — passed.
- `pnpm check:token-gates` — the changed files are clean, but the
repo-wide command currently reports nine unrelated `#9627` comment
references already present on `master` as color literals.

## Risks

- Low risk: this changes display strings and explanatory copy only;
override storage, lane identifiers, API contracts, and execution
behavior are unchanged.
- The global token-gate false positive may also appear in CI until the
unrelated `#9627` references on `master` are allowlisted or the scanner
ignores issue-number comments.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex coding agent assisted this PR preparation using medium
reasoning, repository-aware shell execution, Git/GitHub tooling, and
local test/typecheck execution. The hosted runtime did not expose an
exact model ID or context-window size to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 15:53:03 -05:00
Dotta 52aea90263
feat: organize skills with nested folders and My Skills (#9633)
## Thinking Path

> - Paperclip is the open source control plane people use to organize
and govern AI-agent companies
> - Company skills are durable resources that users browse, import,
assign, and maintain over time
> - A flat skill list plus tags does not provide a stable location or
hierarchy for personal, company, project-imported, and bundled skills
> - Folder paths need to be canonical, company-scoped, safe to move, and
preserved across re-imports without changing skill IDs
> - The `/skills` UI also needs traversal, breadcrumbs, move/create
flows, and a dedicated My Skills namespace that work on desktop and
mobile
> - This pull request adds the folder data model and APIs, reserved-root
lifecycle, project import behavior, and the folder-first skills
experience
> - The benefit is a predictable filesystem-like organization model
while tags remain available for cross-cutting classification

## Linked Issues or Issue Description

Refs #9619 — the reviewed folder foundation was intentionally closed and
folded into this combined feature PR.
Refs #9026 — earlier flat-folder attempt superseded by this integrated
implementation.
Refs #3281 — related skill organization proposal; this PR uses canonical
persisted folders rather than deriving groups from skill keys, and does
not add hidden-skill behavior.

**Feature request**

- **Problem:** Skills currently lack a canonical hierarchical location,
making personal skills, project imports, bundled skills, and
company-authored skills difficult to traverse and manage at scale.
- **Proposed behavior:** Add nested company-scoped folders with stable
paths, reserved My/Projects/Bundled roots, subtree queries, safe
move/create operations, and a folder-first `/skills` library UI.
- **Import behavior:** New project scans file skills under
`projects/<project-slug>`; later imports update content without
overriding a user-selected folder.
- **Alternatives considered:** Tags alone remain useful for
cross-cutting classification, but they do not provide canonical
location, nesting, reserved namespaces, or stable import placement.
- **Roadmap alignment:** Extends the completed Skills Manager and
Scheduled Routines capabilities without duplicating an active roadmap
item.

## What Changed

- Adds `folders` persistence for routine and skill folders, nested
canonical paths, parent/slug/system-key fields, migration backfills, and
reapply-safe migrations `0174`–`0175` after current master migrations.
- Adds company-scoped folder CRUD, cycle/depth/namespace validation,
reserved My/Projects/Bundled lifecycle, item moves, subtree filtering,
and folder paths on skill results.
- Preserves project-import placement: first import files into the
project folder, while re-import keeps user-owned placement and stable
skill IDs.
- Adds the `/skills` folder tree rail, tags facet, breadcrumbs,
subfolder browser, move/new-folder dialog, canonical detail location,
inline tag editing, and folder-aware Studio creation.
- Keeps bundled skills read-only even when their source metadata is
incomplete by detecting the reserved Bundled folder and hiding
selection/move actions.
- Extends routine folder UI and OpenAPI coverage, and adds regression
tests across migrations, services, routes, tree helpers, pages, and
Studio creation.

## Verification

- `pnpm exec vitest run
packages/db/src/nested-skill-folders-migration.test.ts
server/src/__tests__/folders-routes.test.ts
server/src/__tests__/folders-service.test.ts
server/src/__tests__/company-skills-service.test.ts
server/src/__tests__/routines-service.test.ts
ui/src/components/folders/FolderControls.test.tsx
ui/src/components/folders/SkillFolderTree.test.tsx
ui/src/components/folders/skill-folder-tree.test.ts
ui/src/pages/CompanySkills.test.tsx ui/src/pages/Routines.test.tsx
ui/src/pages/SkillStudio.test.tsx
ui/src/lib/company-skill-routes.test.ts ui/src/lib/skill-create.test.ts`
— 13 files, 192 tests passed.
- `pnpm exec vitest run ui/src/pages/CompanySkills.test.tsx
ui/src/components/folders/SkillFolderTree.test.tsx` — 2 files, 20 tests
passed after preserving the existing PR's bundled-skill fixes.
- `pnpm -r typecheck` — passed for all workspace packages.
- `pnpm test:run` — passed in an isolated CI-like environment with
inherited Paperclip runtime identity and static AWS credential variables
removed.
- `pnpm build` — production build passed for all workspace packages.
- Greptile iteration 2 — 5/5 confidence with zero unresolved threads on
commit `ff2d67aa71`.
- Latest-head GitHub checks — all success, neutral, or skipped; PR is
mergeable with a clean merge state.
- `pnpm check:token-gates` — reports nine existing `#9627` comment false
positives already present on `master`; this PR introduces no new token
violation.

## Risks

- **Migration/backfill:** `0174` creates the foundation and `0175` adds
nested/reserved semantics. Both are ordered after current master
migration `0173`, are covered by numbering/safety checks, and are
designed to be reapply-safe.
- **Reserved namespaces:** My, Projects, and Bundled roots are
service-managed. Regression coverage prevents namespace squatting,
cross-company folder use, bundled writes, cycles, and excessive depth.
- **Behavioral change:** Project scans choose a project folder only on
initial creation; existing skills deliberately retain their current
folder during refresh.
- **UI scope:** The folder rail applies to the Installed library;
Catalog retains the discovery-oriented category sidebar.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI `gpt-5.5` in Codex CLI, medium reasoning mode; runtime did not
expose a context-window value. Used repository/file tools, terminal
execution, Git/GitHub operations, test execution, and code editing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 15:50:45 -05:00
Nicky Leach 1d2b6af5ac
fix(tests): stabilize heartbeat cleanup for tsx update (#9573)
## Thinking Path

> - Paperclip is an open-source platform for orchestrating AI agents,
built on an embedded-Postgres server running a heartbeat loop to advance
agent work.
> - The server test suite exercises heartbeat liveness escalation and
retry scheduling logic against a real embedded database; tests create
and tear down full database state across every case.
> - Dependabot PR #9480 bumps `tsx` from 4.22.4 to 4.23.1. The new
version exposed two fragile teardown patterns in the heartbeat tests
that caused failures.
> - The first problem: `TRUNCATE TABLE "companies" CASCADE` in the
liveness-escalation teardown clashes with FK constraints when child
tables (e.g. `heartbeat_run_events`, `issue_tree_hold_members`) hold
rows that tsx 4.23.1's changed execution order materialises before the
CASCADE runs.
> - The second problem: the retry-scheduling test duplicated a 10-line
delete block inline at two mid-test reset points; one copy deleted
`heartbeat_run_events` after `heartbeat_runs` (wrong FK order) and
`activityLog` was deleted twice.
> - A third concern was identified during review: several `GET
/tool-connections/:connectionId` routes called `assertCompanyAccess`
before checking whether the actor has access at all, leaking 403
(existence oracle) instead of 404. This is fixed in this PR.
> - This PR updates the three `tsx` version pins to `^4.23.1`, replaces
the TRUNCATE with explicit child-to-parent deletes, centralises the
retry cleanup into a shared `cleanupRetryFixture()` helper, and adds
`hasCompanyAccess` pre-checks before the four affected
`assertCompanyAccess` calls in `tool-access.ts`.
> - The benefit is CI green on tsx 4.23.1, cleaner non-duplicated
teardown code across both test files, and no cross-tenant existence
leakage on tool-connection routes.

## Linked Issues or Issue Description

Refs #9480 (`tsx` 4.22.4 → 4.23.1 dependabot bump whose CI failures this
fixes)

## What Changed

- **cli/package.json**, **packages/db/package.json**,
**server/package.json**: bump `tsx` dev-dependency range from `^4.22.4`
to `^4.23.1` so package manifests agree with the lockfile update landing
in #9480. `pnpm-lock.yaml` is left untouched — GitHub Actions owns
lockfile regeneration.
- **heartbeat-issue-liveness-escalation.test.ts**: replace `TRUNCATE
TABLE "companies" CASCADE` with explicit FK-ordered deletes. The new
chain adds `heartbeatRunEvents`, `issueTreeHoldMembers`,
`agentRuntimeState`, and `companySkills` before their respective parent
tables.
- **heartbeat-retry-scheduling.test.ts**: extract the repeated teardown
block into a `cleanupRetryFixture()` helper; call it from `afterEach`
and the two mid-test resets; fix `heartbeatRunEvents` deleted before
`heartbeatRuns` (parent-child FK order); remove the duplicate
`activityLog` delete.
- **server/src/routes/tool-access.ts**: add `hasCompanyAccess`
pre-checks before `assertCompanyAccess` on four `GET
/tool-connections/:connectionId` and `GET
/tool-profiles/:profileId/new-tools` routes. Returns 404 instead of 403
when the actor cannot access the resource, closing the cross-tenant
existence oracle.

## Verification

```sh
# Focused test run (49 tests, all pass)
pnpm exec vitest run \
  server/src/__tests__/heartbeat-retry-scheduling.test.ts \
  server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts

pnpm --filter @paperclipai/server typecheck   # pass
pnpm -r typecheck                              # pass
pnpm build                                     # pass
```

Full `pnpm test:run` was also attempted: server suite (242 files, 2 243
tests) and UI suite (310 files, 2 536 tests) both passed. A backup-dir
assertion in `src/__tests__/onboard.test.ts` failed but is unrelated to
this diff — it expects a temp `PAPERCLIP_HOME` but receives the global
instance path.

## Risks

Low risk. Changes are limited to test teardown logic, dev-dependency
version pins, and existence-oracle guard additions on read-only
tool-connection routes. No new business logic or production data paths
are introduced.

## Model Used

Claude Sonnet 4.6 (`claude-sonnet-4-6`, 200 k context, tool use, agentic
coding)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 14:02:14 -05:00
dependabot[bot] db61cc97d3
build(deps-dev): bump tsx from 4.22.4 to 4.23.1 (#9480)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.22.4 to 4.23.1.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/privatenumber/tsx/releases">tsx's
releases</a>.</em></p>
<blockquote>
<h2>v4.23.1</h2>
<h2><a
href="https://github.com/privatenumber/tsx/compare/v4.23.0...v4.23.1">4.23.1</a>
(2026-07-13)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>support tsImport after global preload (<a
href="8d4ffc24f3">8d4ffc2</a>)</li>
<li><strong>watch:</strong> avoid clearing piped output (<a
href="95d0672e02">95d0672</a>)</li>
<li><strong>watch:</strong> treat script and dependency paths literally
(<a
href="79fddde523">79fddde</a>)</li>
</ul>
<h3>Performance Improvements</h3>
<ul>
<li>index transform cache lazily (<a
href="e818ad6081">e818ad6</a>)</li>
<li>load esbuild lazily in CLI (<a
href="d0679381b6">d067938</a>)</li>
<li>map Node TypeScript formats directly (<a
href="cdcc6232a3">cdcc623</a>)</li>
<li>use sync module hooks on Node v22.22.3+ (<a
href="f8992f1a50">f8992f1</a>)</li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.23.1"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
<h2>v4.23.0</h2>
<h1><a
href="https://github.com/privatenumber/tsx/compare/v4.22.5...v4.23.0">4.23.0</a>
(2026-07-03)</h1>
<h3>Bug Fixes</h3>
<ul>
<li>avoid redundant filesystem probes during module resolution (<a
href="257bbbb7eb">257bbbb</a>),
closes <a
href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a></li>
</ul>
<h3>Features</h3>
<ul>
<li>add multi-scenario startup benchmark suite (<a
href="c178197b10">c178197</a>),
closes <a
href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a>
<a
href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a>
<a href="https://github.com/hi/issues/signal">hi#signal</a> <a
href="https://redirect.github.com/privatenumber/tsx/issues/145">privatenumber/tsx#145</a>
<a
href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a></li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.23.0"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
<h2>v4.22.5</h2>
<h2><a
href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.22.5">4.22.5</a>
(2026-07-02)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>isolate hook state per async module.register() registration (<a
href="a305f365f0">a305f36</a>)</li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.22.5"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="79fddde523"><code>79fddde</code></a>
fix(watch): treat script and dependency paths literally</li>
<li><a
href="e818ad6081"><code>e818ad6</code></a>
perf: index transform cache lazily</li>
<li><a
href="cdcc6232a3"><code>cdcc623</code></a>
perf: map Node TypeScript formats directly</li>
<li><a
href="d0679381b6"><code>d067938</code></a>
perf: load esbuild lazily in CLI</li>
<li><a
href="95d0672e02"><code>95d0672</code></a>
fix(watch): avoid clearing piped output</li>
<li><a
href="6fd4607e8a"><code>6fd4607</code></a>
docs: add per-page metadata</li>
<li><a
href="f4176d8c63"><code>f4176d8</code></a>
docs: generate sitemap</li>
<li><a
href="8d4ffc24f3"><code>8d4ffc2</code></a>
fix: support tsImport after global preload</li>
<li><a
href="f0e89b244c"><code>f0e89b2</code></a>
docs: document Node's public type-stripping API vs internal loader
path</li>
<li><a
href="f8992f1a50"><code>f8992f1</code></a>
perf: use sync module hooks on Node v22.22.3+</li>
<li>Additional commits viewable in <a
href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.23.1">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-16 13:59:26 -05:00
dependabot[bot] c4da1c925f
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.52.0 to 0.59.0 (#9484)
Bumps
[@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp)
from 0.52.0 to 0.59.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@​agentclientprotocol/claude-agent-acp's
releases</a>.</em></p>
<blockquote>
<h2>v0.59.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.1...v0.59.0">0.59.0</a>
(2026-07-13)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> bump nanoid from 3.3.15 to 3.3.16 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>)
(<a
href="e67dacdcac">e67dacd</a>)</li>
<li><strong>deps:</strong> bump
<code>@​anthropic-ai/claude-agent-sdk</code> to 0.3.207 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>)
(<a
href="c7f5b8fe76">c7f5b8f</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Add subagent parent tool use attribution (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>)
(<a
href="9cd48c597a">9cd48c5</a>)</li>
<li>forward result text when the turn emitted no assistant message (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>)
(<a
href="61ae8609d4">61ae860</a>)</li>
<li>hold a turn open while its background subagents are still live (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>)
(<a
href="7a70f82739">7a70f82</a>)</li>
<li>Refine tool calls from streamed input (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>)
(<a
href="2c19974d5b">2c19974</a>)</li>
<li>Seed context window from SDK usage report (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>)
(<a
href="3ba2d367a2">3ba2d36</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/596">#596</a></li>
<li>Skip synthetic login messages on replay (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>)
(<a
href="c7dff3cf7e">c7dff3c</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/863">#863</a></li>
</ul>
<h2>v0.58.1</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.0...v0.58.1">0.58.1</a>
(2026-07-09)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>Use valid npm version for publish (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/855">#855</a>)
(<a
href="8b366d8e9a">8b366d8</a>)</li>
</ul>
<h2>v0.58.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.57.0...v0.58.0">0.58.0</a>
(2026-07-09)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> update to
<code>@​anthropic-ai/claude-agent-sdk</code> 0.3.205 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/854">#854</a>)
(<a
href="f664ced76d">f664ced</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Preserve live model on resumed sessions (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/848">#848</a>)
(<a
href="f3d8ae3eb3">f3d8ae3</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/845">#845</a></li>
<li>Report usage for cancelled active turns (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/846">#846</a>)
(<a
href="b03318f59c">b03318f</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/844">#844</a></li>
<li>tolerate missing text in streamed thinking chunks (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/852">#852</a>)
(<a
href="e944ceddea">e944ced</a>)</li>
<li>Use SDK guards for elicitation validation (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/850">#850</a>)
(<a
href="32b93501d7">32b9350</a>)</li>
</ul>
<h2>v0.57.0</h2>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.56.0...v0.57.0">0.57.0</a>
(2026-07-07)</h2>
<h3>Features</h3>
<ul>
<li>Update <code>@​anthropic-ai/claude-agent-sdk</code> to 0.3.202 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/843">#843</a>)
(<a
href="1612c07895">1612c07</a>)</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@​agentclientprotocol/claude-agent-acp's
changelog</a>.</em></p>
<blockquote>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.1...v0.59.0">0.59.0</a>
(2026-07-13)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps-dev:</strong> bump nanoid from 3.3.15 to 3.3.16 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>)
(<a
href="e67dacdcac">e67dacd</a>)</li>
<li><strong>deps:</strong> bump
<code>@​anthropic-ai/claude-agent-sdk</code> to 0.3.207 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>)
(<a
href="c7f5b8fe76">c7f5b8f</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Add subagent parent tool use attribution (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>)
(<a
href="9cd48c597a">9cd48c5</a>)</li>
<li>forward result text when the turn emitted no assistant message (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>)
(<a
href="61ae8609d4">61ae860</a>)</li>
<li>hold a turn open while its background subagents are still live (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>)
(<a
href="7a70f82739">7a70f82</a>)</li>
<li>Refine tool calls from streamed input (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>)
(<a
href="2c19974d5b">2c19974</a>)</li>
<li>Seed context window from SDK usage report (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>)
(<a
href="3ba2d367a2">3ba2d36</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/596">#596</a></li>
<li>Skip synthetic login messages on replay (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>)
(<a
href="c7dff3cf7e">c7dff3c</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/863">#863</a></li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.0...v0.58.1">0.58.1</a>
(2026-07-09)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>Use valid npm version for publish (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/855">#855</a>)
(<a
href="8b366d8e9a">8b366d8</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.57.0...v0.58.0">0.58.0</a>
(2026-07-09)</h2>
<h3>Features</h3>
<ul>
<li><strong>deps:</strong> update to
<code>@​anthropic-ai/claude-agent-sdk</code> 0.3.205 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/854">#854</a>)
(<a
href="f664ced76d">f664ced</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li>Preserve live model on resumed sessions (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/848">#848</a>)
(<a
href="f3d8ae3eb3">f3d8ae3</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/845">#845</a></li>
<li>Report usage for cancelled active turns (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/846">#846</a>)
(<a
href="b03318f59c">b03318f</a>),
closes <a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/844">#844</a></li>
<li>tolerate missing text in streamed thinking chunks (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/852">#852</a>)
(<a
href="e944ceddea">e944ced</a>)</li>
<li>Use SDK guards for elicitation validation (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/850">#850</a>)
(<a
href="32b93501d7">32b9350</a>)</li>
</ul>
<h2><a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.56.0...v0.57.0">0.57.0</a>
(2026-07-07)</h2>
<h3>Features</h3>
<ul>
<li>Update <code>@​anthropic-ai/claude-agent-sdk</code> to 0.3.202 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/843">#843</a>)
(<a
href="1612c07895">1612c07</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="30b7c06f76"><code>30b7c06</code></a>
chore(main): release 0.59.0 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/860">#860</a>)</li>
<li><a
href="c7f5b8fe76"><code>c7f5b8f</code></a>
feat(deps): bump <code>@​anthropic-ai/claude-agent-sdk</code> to 0.3.207
(<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>)</li>
<li><a
href="7a70f82739"><code>7a70f82</code></a>
fix: hold a turn open while its background subagents are still live (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>)</li>
<li><a
href="e67dacdcac"><code>e67dacd</code></a>
feat(deps-dev): bump nanoid from 3.3.15 to 3.3.16 (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>)</li>
<li><a
href="61ae8609d4"><code>61ae860</code></a>
fix: forward result text when the turn emitted no assistant message (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>)</li>
<li><a
href="2c19974d5b"><code>2c19974</code></a>
fix: Refine tool calls from streamed input (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>)</li>
<li><a
href="c7dff3cf7e"><code>c7dff3c</code></a>
fix: Skip synthetic login messages on replay (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>)</li>
<li><a
href="3ba2d367a2"><code>3ba2d36</code></a>
fix: Seed context window from SDK usage report (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>)</li>
<li><a
href="f4e750decf"><code>f4e750d</code></a>
feat(deps): bump the minor group with 11 updates (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/862">#862</a>)</li>
<li><a
href="9cd48c597a"><code>9cd48c5</code></a>
fix: Add subagent parent tool use attribution (<a
href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.52.0...v0.59.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-16 13:17:13 -05:00
LeonSGP d62c5adb7a
fix(ui): stop recreating markdown mention observers (#3809)
Fixes #3759

## Thinking Path

> - Paperclip orchestrates AI agents and issue workflows, so the comment
composer has to stay stable under normal typing.
> - The affected subsystem is the shared markdown comment editor used
across issue and workflow surfaces.
> - That editor decorates mention links after Lexical updates the
editable DOM.
> - The current mention decoration effect recreates its
`MutationObserver` whenever `value` changes, which happens on every
keystroke.
> - That observer also reacts to the DOM mutations produced by mention
decoration itself, creating unnecessary observer churn in Chrome.
> - This pull request keeps one observer instance alive, batches
decoration work into `requestAnimationFrame`, and disconnects the
observer while decoration writes run.
> - The benefit is lower observer churn while preserving the existing
mention chip behavior.

## What Changed

- Removed `value` from the mention-decoration observer effect dependency
list so the observer is not recreated on every external value update.
- Batched mention decoration with `requestAnimationFrame` and
temporarily disconnected the observer while DOM decorations are applied
to avoid self-triggered feedback loops.
- Added a regression test that verifies external value changes do not
recreate the mention decoration observer.

## Verification

- `pnpm install --frozen-lockfile`
- `pnpm --filter @paperclipai/ui exec tsc --noEmit`
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/MarkdownEditor.test.tsx` currently fails before test
collection on the existing repo baseline with `TypeError: undefined is
not an object (evaluating 'z.string')` from
`packages/shared/src/adapter-type.ts`

## Risks

- Low risk. The change is scoped to the mention decoration observer
lifecycle and keeps the existing decoration logic intact.
- The main behavior change is deferring decoration to the next animation
frame instead of running immediately on every observed mutation.

## Model Used

- OpenAI Codex, GPT-5-based coding agent with local tool use in the
Codex CLI environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] If this change affects the UI, I have included before/after
screenshots
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-16 10:44:17 -07:00
Dotta 8b04147ca4
perf(ui): stop polling the event-sourced company live-runs list (#9701)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its web UI keeps live views fresh with a live-events websocket plus
React Query, and #9627 began replacing polling with event-sourcing
(pushing data into the cache)
> - After the churn fixes (#9624) and #9627 shipped, a long-lived tab's
memory footprint was still climbing, so I re-profiled with the Chrome
DevTools MCP
> - The dominant remaining churn is React Query re-arming a polled
query's `refetchInterval` timer on **every** observer notification — and
our live-event handlers `setQueryData(liveRuns(companyId))` on nearly
every event, so each pushed update re-arms every `liveRuns` observer's
timer (the sidebar is always mounted)
> - #9627 event-sourced the live-runs data but left the now-redundant
`refetchInterval` in place, so we did half the fix — the poll is pure
waste and the thing re-arming timers
> - This pull request removes `refetchInterval` from the event-sourced
company live-runs queries so the frequent cache writes have no timer to
re-arm
> - The benefit is that the steady-state timer churn on the
most-observed resource collapses, so the off-heap footprint stops
climbing

## Linked Issues or Issue Description

No public GitHub issue exists; describing inline per CONTRIBUTING.md →
"Link Issues or Describe Them In-PR", following the bug report template.
Continues #9569 / #9624 / #9627.

**What happened?**

With the earlier fixes deployed, a browser tab left open on the app kept
growing its memory footprint. MCP profiling showed ~100+ `setInterval`
create/clear cycles per 5s on an aged tab (vs ~16 fresh), all from React
Query's refetch-interval timers being re-armed on every `setQueryData`
to the frequently-written `liveRuns(companyId)` query.

**Expected behavior**

A resource whose data is pushed (event-sourced) should not also poll;
cache writes should not repeatedly re-arm interval timers. Idle tabs
should hold a bounded footprint.

**Steps to reproduce**

Open a tab with agents streaming, leave it open, and instrument
`setInterval`/`clearInterval`: the churn rate climbs and traces to
`QueryObserver.updateTimers` (`refetchInterval`) for `liveRuns`,
re-armed by every live-event cache write.

**Paperclip version or commit**

Branch `perf/drop-live-runs-refetch-interval`, off `master` (after
#9627).

**Deployment mode**

Local dev (`pnpm dev`), web UI. Core UI live-updates plumbing; not
adapter-specific.

## What Changed

- Set `refetchInterval: false` on every site that polls the **plain**
`queryKeys.liveRuns(companyId)` query (event-sourced by #9627):
`Sidebar`, `SidebarAgents`, `Issues`, `Inbox`, `IssueDetail`
(companyLiveRuns), `ProjectDetail` (×2), `Routines`,
`ExecutionWorkspaceDetail`, `bridge-init`.
- Removed the now-unused `useVisibilityRefetchInterval` interval
vars/imports in `Issues`, `Inbox`, `IssueDetail`.
- **Left variant-key sites polling on purpose** — `Agents` page
(`[...liveRuns, "agents-page"]`) and `ActiveAgentsPanel` (`[...liveRuns,
scope, …]`) are NOT event-sourced by #9627 (different exact cache key),
so dropping their poll would make them stale. Those are a later phase.
- Freshness for the converted queries now comes from event-sourcing
(#9627) + its reconnect reconcile; the initial mount fetch and cross-tab
publish (`usePublishSharedQueryData`) still happen.

## Verification

- MCP profiling identified the churn: the single churning callback is
React Query's `refetchInterval` timer, re-armed by
`setQueryData(liveRuns)` on live events.
- `vitest`: all affected suites pass (`Sidebar`, `SidebarAgents`,
`Issues`, `Inbox`, `IssueDetail`, `ProjectDetail`, `Routines`,
`ExecutionWorkspaceDetail`, `LiveUpdatesProvider`) — 153 tests.
- Updated two `SidebarAgents` linger-window tests: they advanced fake
timers to the *exact* linger-expiry boundary and had relied on
poll-induced re-renders to flush. The linger self-schedules its own
`setTimeout`, so the tests now cross the boundary with a small margin +
an explicit flush (no product change).
- `tsc -b` clean.
- End-to-end footprint reduction should be re-measured against a rebuilt
bundle with the same instrumentation.

## Risks

Low, client-only.
- `liveRuns(companyId)` freshness now depends entirely on event-sourcing
+ reconnect reconcile (both from #9627). If an event path is missed, the
reconnect handler refetches once; durable replay is a planned later
phase.
- Variant-key run lists (Agents page, ActiveAgentsPanel) are unchanged
and still poll, so they don't regress.
- Issue-scoped run queries (`issues.liveRuns/activeRun/runs`) are
**not** touched here — they aren't event-sourced yet and are a separate
phase.

## Model Used

- **Provider:** Anthropic, via the Claude Code CLI.
- **Model:** Claude Opus 4.8 (`claude-opus-4-8`).
- **Reasoning mode:** Extended thinking enabled.
- **Capabilities used:** tool use (shell, file editing), and the Chrome
DevTools MCP to re-profile the live instance and pinpoint the
`refetchInterval` timer churn.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work (a perf/plumbing change)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (continues #9627; no duplicates)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered and documented any risks above
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no user-facing docs; rationale documented inline)
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-16 11:42:04 -05:00
Dotta 5a5c918705
fix(acpx): configure Codex models at startup (#9700)
Move Codex ACPX model, reasoning effort, and fast-mode settings into CODEX_CONFIG startup config so arbitrary model IDs avoid ACP session picker validation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-07-16 11:41:48 -05:00
Dotta 9a92124c63
fix(codex-local): raise default output-inactivity timeout to 30m (#9699)
## Thinking Path

> - Paperclip is the open source control plane for managing AI-agent
companies.
> - Agent adapters are responsible for launching, observing, and
terminating provider processes safely.
> - The local Codex adapter uses an output-inactivity monitor to stop
genuinely hung child processes.
> - The existing seven-minute default also stopped legitimate
long-running tasks that emitted no output while tests or remote checks
were still running.
> - Explicit per-agent timeout overrides already provide configuration
flexibility, so the smallest safe correction is to increase only the
default window.
> - This pull request raises the default to 30 minutes, retains the
existing termination behavior, and adds a regression assertion for the
new value.
> - The benefit is fewer unnecessary Codex restarts while still bounding
genuinely silent processes well below the platform safety limit.

## Linked Issues or Issue Description

**Pre-submission checklist**
- Searched open and closed GitHub issues and pull requests; no duplicate
fix exists. The original inactivity monitor was introduced in #5017.
- Reproduces on the current `master` implementation.
- The termination originates in Paperclip's Codex adapter inactivity
monitor rather than the model provider or local configuration.

**What happened?**
Long-running `codex_local` tasks were terminated after seven minutes
without stdout or stderr, even when the child process was still
performing legitimate work such as a quiet test suite or waiting for
remote checks.

**Expected behavior**
The default inactivity window should tolerate common long-running quiet
tasks while continuing to terminate processes that remain silent for an
extended period.

**Steps to reproduce**
1. Start a `codex_local` run with the default
`outputInactivityTimeoutMs` configuration.
2. Have the child process perform legitimate work without emitting
stdout or stderr for more than seven minutes.
3. Observe the adapter terminate the process at the old default
threshold.

**Version / deployment**
Current `master`, built from source in local/self-hosted deployments,
using the Codex adapter. This behavior is not database- or
access-context-specific.

## What Changed

- Raise `DEFAULT_CODEX_OUTPUT_INACTIVITY_TIMEOUT_MS` from seven minutes
to 30 minutes.
- Update the adapter configuration documentation to state the new
default.
- Add a focused regression assertion that pins the default to 30
minutes.

## Verification

- `pnpm vitest run
packages/adapters/codex-local/src/server/output-inactivity-monitor.test.ts`
— 15 tests passed.
- `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed.

## Risks

- Low risk: only the fallback default changes; explicit positive timeout
values and `null` disablement retain their existing behavior.
- A genuinely silent Codex process now remains alive up to 23 minutes
longer before the same SIGTERM/SIGKILL cleanup path runs.
- No schema, API, migration, UI, or workflow changes are included.

> This is a focused bug fix and does not overlap with planned core
feature work in `ROADMAP.md`.

## Model Used

- Anthropic Claude Fable 5 via the `claude_local` adapter produced the
initial investigation and implementation with repository/tool access.
- OpenAI Codex via the Codex CLI prepared the PR, added the focused
regression assertion, and ran verification; the runtime did not expose a
more specific model ID or context-window value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] I have used a public-friendly branch name without internal tracker
identifiers (execution-workspace exception: this branch is
runtime-managed and cannot be renamed)
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

The source branch is execution-workspace managed and cannot be renamed
during this run; the PR title and body intentionally contain no internal
tracker references.

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 11:41:29 -05:00
Dotta 6ec059ab4e
fix(server): suppress stale handoff alarms during live continuation (#9695)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The control plane records a successful-run handoff when productive
work ends without a durable next-step disposition
> - That handoff state was derived only from the latest activity event,
without checking whether a corrective run or wake was currently alive
> - As a result, actively progressing issues could still show a
high-severity missing-disposition alarm and blocked-inbox row
> - The same stale required event could also remain indefinitely when a
later successful run correctly skipped recovery because another valid
continuation path already existed
> - This pull request makes the derived state liveness-aware, suppresses
attention only while the live path exists, and resolves stale required
events on valid-path skips
> - The benefit is that productive work stays calm while genuine stalls
still resurface automatically when liveness disappears

## Linked Issues or Issue Description

- **Bug:** An issue whose latest successful-run handoff event is
`required` continues to report a missing disposition even while a
heartbeat run, scheduled retry, or queued/deferred/claimed wake is
actively targeting that issue.
- **Expected behavior:** The API should expose current continuation
liveness, the blocked inbox should suppress the alarm only while that
path remains live, and a later successful run that skips recovery
because a valid path exists should durably resolve the stale event.
- **Related but distinct:** #9370 changes disposition freshness at
detection time; #8748 adds an explicit policy opt-out. This PR preserves
detection/escalation policy and fixes read-time/current-liveness state.

## What Changed

- Extended `SuccessfulRunHandoffState` with `hasLiveContinuation` and
optional `liveRunId` evidence.
- Added bounded liveness hydration for required handoff states using
active heartbeat-run and wake-request signals.
- Suppressed `missing_disposition` blocked-inbox rows only while a run,
scheduled retry, or live wake targets the issue.
- Added durable `issue.successful_run_handoff_resolved` logging when
handoff detection skips because another valid continuation path owns the
next action.
- Added focused regressions for live/absent derived state, self-healing
attention suppression, valid-path skip classification, and
resolved-event logging.
- Updated UI normalization and fixtures for the shared contract without
changing rendering behavior.

## Verification

- `pnpm --filter @paperclipai/shared typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm vitest run
server/src/services/recovery/successful-run-handoff.test.ts
server/src/__tests__/issue-list-assignee-filter-routes.test.ts
server/src/__tests__/issue-blocker-attention.test.ts` — 56 passed
- `pnpm vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts -t "queues one
finish-handoff wake when a successful run leaves in-progress work
without a next action"` — 1 passed
- `git diff --check`

## Risks

- Low risk: no schema or migration changes, and detection, bounded
correction attempts, and escalation behavior are unchanged.
- Liveness lookups are limited to issues whose latest handoff state is
`required`; blocked-inbox suppression reuses rows already loaded by that
query path.
- Suppression is read-time and self-healing: when the run or wake stops,
the alarm returns on the next fetch.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using `gpt-5.4`, tool-enabled software-engineering
workflow with repository, shell, test, Git, GitHub, and Paperclip
control-plane access. Context-window size is not exposed by this
runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 10:49:47 -05:00
Dotta a04a77c9d3
feat(authz): govern agent inbox archive access (#9658)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies and their work
> - The inbox subsystem must let agents act for a responsible user
without silently granting access to every company user's tasks
> - Existing authorization had no inbox-specific action, target-user
scope, or per-user agent policy
> - Inbox archive data also needs company-safe ownership and replay-safe
schema changes before API mutations can rely on it
> - This pull request adds the database policy foundation and a
fail-closed `inbox:manage` authorization decision
> - The benefit is a least-privilege core for later inbox archive
endpoints, including explicit cross-user grants and low-trust denial

## Linked Issues or Issue Description

### Subsystem affected

Cross-cutting (`packages/db`, `packages/shared`, and `server`).

### Problem or motivation

Agents need to manage inbox state for the user responsible for their
run, but the control plane lacks an inbox-specific permission model and
user-targeted grant scope. A generic mutation path would risk cross-user
access or inconsistent policy enforcement.

### Proposed solution

Add inbox archive ownership and per-user agent policies, introduce
`inbox:manage`, and evaluate responsible-user defaults,
disabled/allowlist policies, active membership, low-trust presets, and
scoped cross-user grants in one authorization decision.

### Alternatives considered

Reusing generic issue mutation permissions was rejected because it
cannot express user-targeted inbox scope. Requiring grants for all
self-user access was rejected because it would make the responsible-user
path closed by default instead of using the requested per-user policy
model.

### Roadmap alignment

`ROADMAP.md` contains no overlapping inbox archive or inbox
authorization item; this is incremental control-plane authorization
work.

### Additional context

This PR provides the authorization and schema foundation. Route and UI
behavior can build on this decision without duplicating access-control
rules.

## What Changed

- Builds on the merged migration `0172_inbox_archive_agent_policies`
(#9654) for company/user-scoped inbox archives and per-user agent policy
rows.
- Added replay-safe migration `0173_inbox_policy_agent_cleanup` with a
GIN allowlist index and GIN-backed database cleanup that removes deleted
agent IDs from policy allowlists.
- Added Drizzle schema exports for inbox agent policies and
responsible-user ownership on inbox archives.
- Added the shared `inbox:manage` permission key and `scope.userIds`
evaluation for user-targeted grants.
- Added fail-closed inbox authorization for unresolved targets, inactive
memberships, low-trust agents, disabled policies, allowlist misses, and
ungranted cross-user access.
- Added migration replay coverage and the full inbox authorization
decision matrix.

## Verification

- `pnpm exec vitest run
packages/db/src/inbox-archive-agent-policies-migration.test.ts
server/src/__tests__/authorization-service.test.ts` — 50 tests passed.
- `pnpm --filter @paperclipai/db typecheck`
- `pnpm --filter @paperclipai/shared typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `git diff --check origin/master...HEAD`

## Risks

- The merged `0172` migration changed inbox archive uniqueness from
agent-owned to responsible-user-owned rows; `0173` is additive (index +
cleanup trigger) and idempotent, and replay coverage verifies both
remain safe for databases that already applied an earlier form.
- `scope.userIds` uses the existing JSON grant-scope parser, so
malformed privileged grant payloads continue to fail through the shared
parsing behavior rather than a dedicated schema.
- Cross-user grants intentionally act as board-admin overrides;
responsible-user default access remains bounded by disabled and
allowlist policies.
- The authorization action is not yet wired to public mutation routes,
limiting immediate behavioral impact while establishing the contract
those routes must use.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with `gpt-5.6-sol`, high reasoning effort, CLI tool use,
code execution, GitHub CLI, and Paperclip control-plane integration.
Context window size is not exposed by the configured adapter.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 09:51:48 -05:00
Dotta c65ab09d9f
fix(recovery): wait for provider quota resets (#9635)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies and keep assigned work moving safely.
> - The recovery subsystem decides whether a failed agent run should
retry, wait, block for configuration, or escalate to another owner.
> - Provider usage-limit failures currently arrive as generic
`adapter_failed` results, so stranded-work reconciliation can create
takeover recovery even when the provider states that capacity will reset
later.
> - Credential and model lookup failures are also configuration
problems, not evidence that another agent should take over the task.
> - This pull request classifies those failure families at recovery time
and persists the classification on the run.
> - Quota failures now schedule a monitor for the original assignee at
the parsed reset time, or after a bounded default backoff when no reset
time is available.
> - The benefit is that transient provider capacity waits no longer wake
recovery owners, while configuration failures stop with an actionable
classification.

## Linked Issues or Issue Description

No public GitHub issue exists for this exact change.

**What happened?** When an assigned issue's latest run failed with a
provider usage-limit message such as "try again at 12:00 AM (UTC),"
recovery treated the run as generic `adapter_failed` work and could
create a takeover action. Missing credentials and `model_not_found`
failures followed the same generic path.

**Expected behavior:** Provider quota failures should keep the original
assignee and schedule a monitor for the reset time, without creating
recovery work or immediately waking another owner. Missing credentials
and model lookup failures should be classified as
`configuration_incomplete` and blocked with the configuration fix
recorded.

**Steps to reproduce:**
1. Assign and start an issue for an agent.
2. Record a failed heartbeat run with `errorCode: adapter_failed` and a
provider quota/reset message.
3. Run stranded assigned-issue reconciliation.
4. Observe that the old behavior routes the issue through generic
recovery instead of waiting for provider capacity.

Reproduced on `master` at `9af96461d`. This is a core recovery bug, not
adapter-specific, and applies to built-from-source deployments with
either embedded PGlite or Postgres.

Related work checked: #9288 adds adapter-side Claude provider-limit
classification; #5392 suppresses some recovery creation for quota-class
errors; #9634 is a broader recovery-routing change with overlapping
provider-quota behavior. This PR is the narrow recovery-service fix with
focused parsed-reset, fallback-backoff, zero-takeover, and
configuration-failure coverage.

## What Changed

- Added conservative recovery-time classification for provider quota,
missing-credential, and model-not-found adapter failures.
- Parsed provider reset timestamps with a default one-hour backoff when
no usable reset time is present.
- Persisted `provider_quota` or `configuration_incomplete` metadata on
the failed heartbeat run.
- Scheduled quota monitors for the active issue owner, including the
current review participant, without creating recovery actions or
enqueueing takeover wakes.
- Routed configuration failures to blocked recovery with actionable
evidence instead of a takeover.
- Added unit and embedded-database regression coverage for
parsed/fallback quota timing, zero CTO/recovery wake behavior, and
configuration classification.

## Verification

- `pnpm exec vitest run
server/src/services/recovery/provider-failure-classification.test.ts
server/src/__tests__/issue-recovery-actions.test.ts
server/src/__tests__/issue-monitor-scheduler.test.ts
server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts` — 4
files passed, 66 tests passed.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.

## Risks

- Recovery behavior changes for text-matched adapter failures; matching
is intentionally conservative, and unmatched failures retain the
existing generic recovery path.
- Provider reset strings do not always include a date or timezone;
parsing chooses the next future matching time and falls back to a
one-hour wait when the timestamp is unusable.
- This overlaps the provider-quota portion of broader recovery-routing
PR #9634, so only one implementation should land if both remain open.
- No schema, migration, API contract, or UI changes are included. No
documentation update is needed because this corrects internal recovery
behavior without changing operator commands or configuration.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with model `gpt-5.4`, medium reasoning, tool use, and
code execution. The runtime does not expose its configured
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 09:24:38 -05:00
Dotta 85404b46c5
fix(server): throttle serial recovery repeats (#9651)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI-agent companies.
> - Its recovery services create productivity reviews and liveness
escalations when work stops making progress.
> - Existing uniqueness guards prevent concurrent duplicates, but
terminal recovery tasks can still be recreated serially without enough
time for conditions to change.
> - That creates noisy review churn for persistently stalled issues and
immediate liveness re-escalation after a recovery task closes.
> - This pull request adds bounded, configurable cooldown and no-action
suppression behavior to those two recovery paths.
> - The benefit is quieter recovery automation that still resumes
automatically after source activity or cooldown expiry.

## Linked Issues or Issue Description

### Pre-submission checklist

- [x] I have searched existing open and closed issues and this is not a
duplicate.
- [x] I can reproduce this behavior on `master`.
- [x] I have confirmed the behavior originates in Paperclip core
recovery orchestration, not an adapter, provider, or local
configuration.

### What happened?

Recovery reconciliation can serially recreate equivalent system-origin
tasks after previous tasks become terminal. Productivity reviews allowed
multiple creations for the same source issue within a rolling day, and a
closed liveness escalation could be recreated immediately for the same
incident or recovery leaf.

### Expected behavior

Productivity review creation should be limited to once per rolling 24
hours, repeated completed reviews that produced no source action should
eventually suppress further creation until activity resumes, and
recently terminal liveness escalations should receive a short cooldown
before recreation.

### Steps to reproduce

1. Create a stalled assigned issue that meets productivity-review
eligibility.
2. Complete repeated productivity-review tasks without adding
source-issue activity, then reconcile again within 24 hours.
3. Create and close a liveness escalation for a blocked issue graph,
then immediately reconcile the same graph.
4. Observe that equivalent system tasks can be recreated serially
without a meaningful state change.

### Paperclip version or commit

`5588ddf68175eea448f9d19677b97d7393c38c3d` (`master` when reproduced)

### Deployment mode

Local dev (`pnpm dev`)

### Installation method

Built from source (`pnpm dev` / `pnpm build`)

### Agent adapter(s) involved

- [x] Not adapter-specific (core bug)

### Database mode

Embedded PGlite (default — `DATABASE_URL` unset)

### Access context

Unclear / not applicable

### Node.js version

Current repository-supported Node.js runtime.

### Operating system

Linux development environment.

### Relevant logs or output

No error is emitted; the bug is repeated task creation visible in
persisted issue history.

### Relevant config (if applicable)

No special configuration is required.

### Additional context

The concurrent/open-task uniqueness guards work as designed; this change
targets serial repeats after matching tasks become terminal.

### Privacy checklist

- [x] I have reviewed all pasted output for PII and redacted where
necessary.

## What Changed

- Tightened the productivity-review creation cap to one review per
source issue in a rolling 24-hour window.
- Added configurable suppression after three consecutive completed
reviews with no source-issue activity, with automatic reset when source
activity occurs.
- Added a configurable one-hour default cooldown for matching terminal
liveness escalations.
- Exposed the liveness reconciliation clock/cooldown inputs for
deterministic orchestration tests.
- Added focused tests for daily enforcement, no-action suppression and
reset, and cooldown expiry.

## Verification

- `pnpm exec vitest run
server/src/__tests__/productivity-review-service.test.ts
server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts` — 36
tests passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- Confirm the focused tests demonstrate creation after source activity
and after the liveness cooldown expires.

## Risks

- Low-to-moderate behavioral risk: recovery tasks intentionally appear
less often, so overly aggressive thresholds could delay intervention for
a persistently stalled issue.
- Thresholds are configurable through reconciliation inputs, and source
activity resets productivity-review suppression.
- No database migration, public API change, telemetry contract change,
or UI behavior change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.3-Codex, with reasoning, repository/terminal
tool use, code execution, and test execution. The runtime did not expose
a reliable context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 09:22:50 -05:00
Dotta da549123cc
feat(db): add inbox archive agent policies (#9654)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The inbox tracks issue visibility separately for each responsible
user
> - Existing inbox archive records only identify the user whose inbox
changed, not the actor that made the change
> - Agent-managed inbox cleanup needs durable agent and heartbeat-run
attribution for auditability
> - Users also need a minimal core policy that can remain open, restrict
access to an allowlist, or disable agent inbox management
> - This pull request adds those database contracts without changing
routes or services
> - The benefit is a company-scoped, auditable foundation for later
agent inbox-management APIs

## Linked Issues or Issue Description

### Subsystem affected

`packages/db` — Drizzle schema and migrations.

### Problem or motivation

Agent workflows such as PR gardening can reduce inbox noise after work
is complete, but existing inbox archive records only identify the user
whose inbox changed. They cannot retain which agent acted or which
heartbeat run performed the action, and there is no user-specific policy
controlling whether agents may manage that inbox.

### Proposed solution

Add database-only contracts for agent-managed inbox archiving: actor
type, agent ID, and heartbeat-run attribution on archive rows, plus one
company-scoped policy per user with `open`, `allowlist`, or `disabled`
mode. Existing archive rows retain `user` attribution by default. Route,
service, and UI behavior are intentionally deferred.

### Alternatives considered

- Store attribution only in activity logs: rejected because archive
state needs durable, directly queryable attribution.
- Add richer per-agent rules immediately: rejected because advanced
policy rules belong in a later or enterprise layer; the core table stays
deliberately minimal.
- Implement APIs in the same change: rejected to keep this
migration-focused PR independently reviewable and safe to deploy.

### Roadmap alignment

`ROADMAP.md` does not currently list this capability. This PR
establishes only the database foundation and does not overlap a listed
roadmap item.

### Additional context

Expected behavior is that legacy archive rows upgrade without backfill,
new archive rows can reference an agent and heartbeat run, and each
company/user pair has at most one agent policy.

## What Changed

- Added actor type, agent ID, and heartbeat run ID attribution columns
to `issue_inbox_archives` with enum checks and `SET NULL` foreign keys.
- Added the company-scoped `user_inbox_agent_policies` table with mode
validation, JSONB agent allowlists, timestamps, and unique company/user
ownership.
- Added migration `0172_inbox_archive_agent_policies.sql` using
idempotent DDL for existing installations.
- Added an embedded PostgreSQL migration test covering legacy rows,
attribution round-trips, policy JSON, and policy uniqueness.

## Verification

- `pnpm --filter @paperclipai/db exec vitest run
src/inbox-archive-agent-policies-migration.test.ts`
- `pnpm --filter @paperclipai/db typecheck`

## Risks

- Low migration risk: adding the non-null actor type uses a constant
`user` default so legacy rows upgrade without a data backfill.
- Agent and run deletions clear attribution foreign keys by design; the
actor type remains available for audit interpretation.
- No API behavior changes are included, so the new contracts remain
unused until follow-up service work lands.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.4 with reasoning, repository tools, shell
execution, and test execution. The runtime did not expose a
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 08:48:39 -05:00
Dotta 054b076b58
fix(ui): keep inbox hover and j/k keyboard selection in sync across list reshapes (#9680)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Inbox is the primary triage surface; it supports both mouse
hover and `j`/`k` keyboard navigation over a flat, expandable list of
rows
> - Hover and keyboard selection are meant to be a single shared
"cursor": hovering a row and then pressing `j`/`k` should continue from
the hovered row, not jump elsewhere
> - A prior hover-perf rewrite moved the hovered row into a numeric
`hoveredIndexRef` and nulled it whenever the list reshaped; because the
inbox polls constantly, any refresh between hovering and pressing a key
dropped the hovered index
> - With the hovered index dropped, the next keypress fell back to
selection index `0`, stranding the cursor at the top of the list instead
of continuing from the hovered row
> - This pull request tracks the hovered row's stable identity (a nav
key) alongside the numeric index and re-anchors it by key across
reshapes, mirroring the existing keyboard-selection reconciliation
> - The benefit is that mouse hover and keyboard navigation stay in sync
exactly as intended, even while the inbox is polling

## Linked Issues or Issue Description

**Bug report**

- **What happened:** In the Inbox, hovering a row with the mouse and
then pressing `j`/`k` (or another keyboard shortcut) does not continue
from the hovered row. If the list refreshes (which happens on the
inbox's constant polling) in the moment between hovering and pressing a
key, keyboard selection snaps back to the top row instead.
- **Expected behavior:** The keyboard cursor should be in sync with the
hovered row — pressing `j`/`k` after hovering should move relative to
the row the mouse is over.
- **Steps to reproduce:**
  1. Open the Inbox with several rows.
  2. Hover the mouse over a row partway down the list.
  3. Wait for (or trigger) a background poll/refresh of the list.
  4. Press `j` or `k`.
5. Observe selection jumps to the top of the list instead of continuing
from the hovered row.
- **Root cause:** The `[flatNavItems]` effect nulled `hoveredIndexRef`
on every list reshape. Since hover also clears the keyboard selection
band to `-1`, the fallback selection index resolved to `0`.

## What Changed

- Hoisted `navEntryKey` to module scope so it can compute a stable,
index-independent identity for a nav row from both the hover handler and
the reshape effect.
- Added `hoveredNavKeyRef` to track the hovered row's stable key
alongside the existing numeric `hoveredIndexRef`, set whenever the
pointer selects a row.
- On list reshape, re-anchor the hovered index by key (find the row with
the same key) instead of unconditionally dropping it; only drop the
hover when the row is actually gone. This mirrors the existing
`selectedIndex` key-based reconciliation.
- Added a unit test that hovers a row, reshapes the list via a simulated
poll, presses `j`, and asserts selection continues from the hovered row.

## Verification

- `pnpm --filter ./ui exec vitest run src/pages/Inbox.test.tsx` → 15/15
passing, including the new hover→`j`/`k` sync test and the existing
keyboard-nav tests.
- The new test explicitly covers the reshape-during-hover path that the
prior test suite had deferred to live/e2e verification.

## Risks

- Low risk. The change is confined to the Inbox's in-memory
hover/selection bookkeeping (two refs and one effect); it adds no new
renders (hover still paints via CSS `:hover`) and touches no data
fetching, routing, or persistence. Behavior is unchanged when the list
does not reshape; when it does, the hover now follows the same row
instead of being dropped.

## Model Used

- Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, used with
extended thinking and tool use (file edits, running the UI test suite
locally).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 07:50:48 -05:00
Dotta 263316609e
fix(server): avoid hot restart shutdown deadlock (#9670)
## Thinking Path

> - Paperclip is the open source control plane people use to manage AI
agents and their work
> - The server coordinates agent heartbeats and preserves eligible live
runs during a hot restart
> - Shutdown previously waited for all heartbeat scheduler work before
capturing the hot-restart snapshot
> - A deployment heartbeat can itself be in that scheduler set while
waiting for the restart, creating a circular wait
> - The missing snapshot prevents startup from classifying and adopting
the still-running agent process
> - This pull request captures the snapshot first and skips
scheduler/drain waits only for an eligible hot restart
> - The benefit is a single SIGTERM can restart the server without
losing eligible live agent runs

## Linked Issues or Issue Description

- **Preflight:** Searched open and closed PRs for the hot-restart
shutdown deadlock; no duplicate found. Reproduced on `master` and
confirmed this is core Paperclip behavior.
- **What happened:** During a hot restart initiated by a running
deployment heartbeat, the SIGTERM handler waited for
`heartbeatSchedulerInFlight` before calling
`prepareHotRestartShutdown()`. The heartbeat was itself in that set and
waited for restart completion, so shutdown never wrote the adoption
snapshot.
- **Expected behavior:** An eligible hot restart captures its snapshot
before waiting for scheduler work, preserves live child processes, and
exits after one SIGTERM.
- **Steps to reproduce:**
1. Start a heartbeat that remains active while requesting a hot restart.
  2. Send SIGTERM to the server process.
3. Observe shutdown waiting on the active scheduler task and startup
finding an intent without a shutdown snapshot.
- **Paperclip commit:** `992389480a243b97bda214227e0767eb8c3672af`
- **Deployment/install:** Self-hosted server built from source.
- **Adapter:** Not adapter-specific; reproduced with a Codex heartbeat.
- **Database/access:** Embedded PGlite; agent bearer context.
- **Environment:** Node `v22.22.2` on `Linux 6.17.0-1015-aws aarch64
GNU/Linux`.
- **Privacy:** No secrets, private logs, user paths, or internal issue
references are included.

## What Changed

- Add a focused shutdown coordinator that prepares hot-restart state
before waiting for heartbeat scheduler idleness.
- Skip scheduler-idle and graceful-drain waits only when the hot-restart
service returns `skipDrain: true`.
- Preserve normal graceful shutdown behavior when no eligible intent
exists or preparation fails.
- Add regression coverage for pending scheduler work, normal shutdown,
and preparation failure.

## Verification

- `pnpm exec vitest run server/src/shutdown.test.ts` — 3 passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts -t
'hot-restart'` — 3 passed, 88 skipped.
- `git diff --check origin/master...HEAD` — passed.

## Risks

- Low-to-moderate risk: shutdown ordering changes, but only the
explicitly eligible hot-restart path bypasses scheduler-idle and
run-drain waits.
- Normal shutdown and hot-restart preparation failures retain the
existing graceful behavior.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. This is a focused bug fix
and does not duplicate planned roadmap work.

## Model Used

- OpenAI Codex coding agent; exact runtime model ID and context-window
size are not exposed to the agent. Tool use, shell execution, repository
editing, and test execution were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
documentation change required for this internal shutdown-order fix)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 05:09:56 -05:00
Dotta 992389480a
fix(server): restore hot-restart run adoption (#9647)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The local heartbeat/runtime subsystem starts long-running local
agent processes and records their run state.
> - Operators sometimes need to rebuild and restart the Paperclip server
while local agent processes are still alive.
> - A normal restart should remain conservative, but a guarded
production hot restart needs an explicit marker, startup reconciliation,
and an inspectable report.
> - The broader hot-restart PR is currently merge-conflicted, so this
pull request lands the minimal server-side recovery path on current
`master`.
> - The benefit is that deploy operators can restart from a current
branch without reverting production changes and without marking adopted
live runs as `process_lost`.

## Linked Issues or Issue Description

No public GitHub issue exists for this deploy-safety fix.

Bug fix:

- What happened: the current deployable `master` branch did not include
the hot-restart marker CLI, startup adoption report path, or health
version proof needed by guarded service restarts.
- Expected behavior: a deploy operator can write a one-shot marker
before restarting, the old server snapshots eligible running child
processes, the new server reports adopted/finalized/lost runs, and
adopted live runs are not reaped as `process_lost`.
- Steps to reproduce: restart a server with running local child-process
heartbeat runs without the marker/adoption path; startup orphan reaping
has no adoption metadata and treats live detached children as lost.
- Paperclip version/commit: fixed on top of `master` at `b606869a6`.
- Deployment mode: production/local-service style deployments that
rebuild and restart the primary `paperclip.service`.
- Related PR: Refs #9628. This PR intentionally lands a smaller
deploy-safe subset because #9628 is currently merge-conflicted.
- Duplicate search: searched public PRs/issues for `hot restart` and
`process_lost adoption`; #9628 is the directly related prior
implementation.

## What Changed

- Added `scripts/request-hot-restart.ts` to write a one-shot hot-restart
intent marker under `PAPERCLIP_HOME`.
- Added `server/src/services/hot-restart.ts` for intent/report path
resolution, parsing, atomic writes, shutdown snapshots, and marker
cleanup.
- Wired server shutdown/startup so explicit hot restarts snapshot active
runs, skip the normal heartbeat drain, reconcile live child processes on
boot, and write `hot-restart-report.json`.
- Preserved adopted run metadata so normal orphan reaping does not
regress adopted live runs to `process_lost`.
- Added `serverVersion` health proof alongside existing `version`, plus
docs and regression coverage.

## Verification

- `pnpm vitest run server/src/__tests__/health.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts` — 2 files
passed, 100 tests passed.
- `pnpm --filter @paperclipai/server typecheck`
- `env PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/hot-restart-cli-smoke"
pnpm --filter @paperclipai/server exec tsx
../scripts/request-hot-restart.ts --server-pid 12345`
- Branch ancestry checked after `git fetch origin master`:
`origin/master` was `b606869a6`, and `HEAD..origin/master` was empty.

## Risks

- Medium risk: process adoption depends on PID/PGID metadata and the
service manager leaving child processes alive for the guarded restart.
- Normal restarts remain conservative, but an incorrect marker PID
intentionally falls back to graceful drain instead of adoption.
- The PR is server-only and does not include the broader
UI/experimental-setting work from #9628.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5 via Codex coding agent in a Paperclip execution
workspace; tool use and shell/code execution enabled; context window not
surfaced by this runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 02:46:09 -05:00
Dotta cf7711ecc5
feat(ui): explain when a message won't reopen a blocked issue (Rule C, PAP-13554) (#9417)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent's task is `blocked`, humans often comment on the issue
thread expecting it to reopen and resume
> - If the issue stays `blocked` because an unresolved (not-done)
blocker still gates the reopen path, the UI said nothing — the human
"sent a message and nothing happened" (a real silent-failure report)
> - That silence makes the product feel broken even though the gate is
working as designed
> - This pull request adds Rule C copy to the existing
`IssueBlockedNotice` surface: when a comment won't reopen a blocked
issue, it explains why and names the deepest unresolved blocker leaf
with its status
> - The benefit is that the human immediately understands the task is
*paused, not stuck*, and knows exactly which task to act on to unblock
it

## Linked Issues or Issue Description

<!-- No public GitHub issue — describing the underlying problem inline
(path B), following the feature issue template. -->

**Problem or motivation**

A human comments on a `blocked` issue expecting it to move back to
`todo` and resume. When the reopen gate keeps it blocked because an
unresolved (not-`done`) blocker remains, nothing in the UI communicates
this, so the user perceives a dropped message ("I sent a message and
nothing happened").

**Proposed solution**

Reuse the existing amber `IssueBlockedNotice` surface (the designated
blocked/recovery surface — no new component). When a message won't
reopen a blocked issue, the notice states that it won't reopen yet,
names the unresolved blocker leaf with its status (e.g. "Still blocked
by PAP-XXXXX (in progress)"), and reassures that it reopens
automatically once the blocker is done. UI-only; no server change.

**Alternatives considered**

Adding a server signal (a reopen-suppressed reason on the comment/notice
payload) was considered but rejected as unnecessary — the
unresolved-blocker set is already available client-side, so the copy is
derived in the component. Done-but-pending-finalize blockers are `done`,
so they fall out of the unresolved set into the standard reopen (Rule B)
path and are correctly not shown as reopen-suppressed.

## What Changed

- `ui/src/components/IssueBlockedNotice.tsx`: added reopen-suppressed
messaging for `blocked` issues that still have unresolved blockers — a
lead sentence ("a message won't reopen it yet, then it reopens
automatically"), the named unresolved blocker leaf with its status, and
an "and N other task(s)" summarization when multiple blockers remain.
- `ui/src/components/IssueBlockedNotice.test.tsx`: added/updated tests
covering the single, nested-chain (deepest-leaf), multiple-blocker, and
empty-blocker states, plus the not-a-reopen-case (`in_progress`) path.
- `ui/storybook/stories/issue-blocked-notice.stories.tsx`: new Storybook
stories rendering each notice state for copy review.

## Verification

- `cd ui && tsc -b` — typecheck clean (previously failed TS2322 on the
story meta; fixed by a default `args`).
- Vitest: `IssueBlockedNotice.test.tsx` green (single / nested /
multiple / empty / in-progress states).
- Storybook stories rendered at 1440×900 for all four states;
screenshots posted on the tracking issue.
- UXDesigner reviewed the notice copy and signed off (no copy edits
required).

## Risks

Low risk. UI-only, additive copy on an already-shipped amber notice
surface; no server or schema changes. The reopen behavior itself is
unchanged — this only explains the existing gate. Worst case is copy
wording, which had a design review.

## Model Used

Claude — claude-opus-4-8 (Opus 4.8), extended thinking, with tool use /
code execution in the Paperclip harness.

---

- [x] I searched the GitHub PR list for similar/duplicate PRs before
opening this one.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 02:34:53 -05:00
Dotta 4f9894df44
fix(server): bound accepted-interaction continuation recovery (#9656)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Heartbeat recovery keeps assigned issues moving when a run or
continuation path disappears
> - Accepted issue-thread interactions can create a continuation wake
after an agent previously parked for review
> - The recovery sweep could requeue that accepted-interaction wake
while the queued-run gate cancelled it using the older pre-acceptance
park summary
> - That cancellation path had no bound, so recovery could repeat the
same wake and cancellation indefinitely
> - This pull request makes accepted-interaction evidence supersede the
older park and caps repeated recovery cancellations at three attempts
> - The benefit is that accepted work resumes normally, while genuine
repeated failures become a visible dependency wait or escalation instead
of a cancel loop

## Linked Issues or Issue Description

Refs #9331

The accepted-interaction continuation recovery added by #9331 can
encounter a stale continuation summary written before approval. The
sweep requeues a continuation carrying the accepted interaction
timestamp, but queued-run invalidation cancels it because the older
summary says to wait for review. Recovery then sees the accepted
interaction without a successful run and requeues again. This PR
prevents that stale-summary cancellation and adds a bounded fallback if
three equivalent cancellations have already occurred.

## What Changed

- Let queued continuation wakes with a parseable `interactionResolvedAt`
bypass a pre-acceptance waiting-for-review park summary.
- Count consecutive unsuccessful continuation runs for the same issue
and agent since interaction acceptance; after three review-park
cancellations, convert a real dependency wait or use the existing
visible escalation path.
- Add focused regression coverage for the park bypass, unchanged
non-interaction park behavior, below-cap requeue, cap escalation, and
successful-run skip.
- Document the accepted-interaction precedence and bounded requeue
contract in execution semantics §9.2.

## Verification

- `pnpm vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts -t "accepted
interaction continuation recovery|accepted interaction recovery after
its continuation succeeds|requeues accepted interaction continuations
stranded"`
- `pnpm vitest run
server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts -t
"pre-acceptance review park|continuation summary parks executor work"`
- `pnpm --filter @paperclipai/server typecheck`

## Risks

- Low risk: the park bypass only applies when the queued context
contains a parseable interaction resolution timestamp.
- The retry bound is scoped to unsuccessful `issue_continuation_needed`
runs for the same company, issue, agent, error code, and post-acceptance
time window.
- No schema, migration, API, or UI changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model ID `gpt-5.5`, high-reasoning coding mode
with repository tool use and command execution; context-window size was
not exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-16 02:34:20 -05:00
Dotta 3124dd0f1e
feat(server): recovery observability report and rate alert (#9644)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - When a run is stranded (process lost, adapter failure, a finished
run with no disposition, an over-eager inactivity kill), the harness
opens a *recovery action* and wakes an owner to recover it
> - Recovery volume regressed sharply in one week — 3.26% of all runs vs
a ~1.2% monthly norm, 5–8x the prior volume — and nobody noticed until
it was ~194 actions deep, because there was no way to *see* the recovery
rate
> - We also could not see which causes drive recovery, nor how often a
manager ends up doing the deliverable work themselves instead of handing
it back to the original owner (the product goal is that managers doing
the work stays rare)
> - This pull request adds a recovery-observability report + API
endpoint: weekly rate normalized per run, a threshold alert, the cause
taxonomy live from the ledger, and the handed-back vs owner-completed
ratio and per-cause routing outcomes
> - The benefit is that a recovery regression like that week is caught
by a threshold instead of by a human noticing it by feel, and each
recovery playbook row can be verified in production

## Linked Issues or Issue Description

**Feature.**

**Problem or motivation**

Recovery takeovers are a first-class exception path
(`issue_recovery_actions`), but there is no aggregate view of them. A
week where the recovery rate tripled went unnoticed until it was deep.
There is no signal for (a) the per-run recovery rate over time, (b)
which cause + run error code drives it, or (c) whether the recovery
owner hands the task back to the original assignee or ends up doing the
deliverable work themselves.

**Proposed solution**

A read-only report service and `GET
/companies/:companyId/recovery-observability` endpoint that surfaces the
weekly rate, a threshold alert, the cause taxonomy, the hand-back ratio,
and per-cause routing outcomes.

**Alternatives considered**

Adding `handed_back` / `owner_completed` to the recovery-action outcome
vocabulary and writing them at resolution time. Rejected for this
change: the distinction is derivable from the recovery owner, the
recorded return owner, and where the source issue actually landed, so
the report works against all historical data without a backfill.

**Roadmap alignment**

Implements the recovery-observability line of the approved
recovery-takeover plan (make regressions visible via a threshold rather
than by human feel); no schema or write-path change.

## What Changed

- Add `server/src/services/recovery-observability.ts`:
- `recoveryObservabilityService(db).report(companyId, { weeks,
thresholdPercent, now })` returns weekly rates (recovery actions / runs,
Monday-anchored to match the retrospective), a `cause` +
`latestRunErrorCode` breakdown, a handed-back vs owner-completed
summary, and per-cause routing outcomes.
- `evaluateRecoveryRateAlert(weekly, thresholdPercent)` — a pure
function (default threshold 2% of runs) returning the breached weeks and
whether the latest week regressed.
- `classifyRecoveryHandoff(...)` — a pure classifier deriving
`self_recovery` / `handed_back` / `owner_completed` from the recovery
owner, return owner, and final issue landing.
- Add `GET /companies/:companyId/recovery-observability` (optional
`weeks` and `threshold` query params) to the existing dashboard router.
- The `weeks` window is bounded (`MAX_WINDOW_WEEKS = 104`,
service-authoritative and re-clamped at the route) so a large query
value can't over-allocate the per-week array.
- Add tests: unit coverage for the alert and the classifier, plus an
embedded-Postgres integration test that seeds synthetic runs and
recovery actions crossing 2% and asserts the alert fires and the
hand-back ratio is computed.

## Verification

- `CI=1 NODE_ENV=development npx vitest run
server/src/__tests__/recovery-observability.test.ts` — 10/10 pass
(includes the synthetic 2%-crossing alert case and the hand-back ratio
case).
- Rendered against a live database of 300+ recovery actions: the weekly
rates reproduce the retrospective (e.g. 1.37% / 1.53% / 0.65% / 0.86% /
1.37% for early-June weeks), the alert fires on the two most recent
weeks (3.15% and 3.05%, both over 2%), and the hand-back summary shows
owner-completed ≈ 73% vs handed-back ≈ 27% — matching the observed
"managers keep ~80% of takeovers".

## Risks

- Low risk. Read-only: adds one GET endpoint and a service; no schema,
migration, or write-path changes. The hand-back classification reads the
source issue's current assignee/status, so a much-later reassignment
could reclassify a historical action — acceptable for an aggregate trend
view.

## Model Used

- Claude, `claude-opus-4-8` (Opus 4.8), extended thinking, tool use /
code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 02:34:00 -05:00
Jakub Mikiciuk f44a002b8d
fix(ui): restore prefix-aware company export/import links (#6648)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Each company workspace in the web UI is mounted under a URL prefix
(e.g. `/NEU/company/...`), and `Link` from `@/lib/router` applies that
prefix automatically via `applyCompanyPrefix`
> - Company Settings rendered its Org Chart, Export, Import, and Cloud
Upstream buttons as raw `<a href>` anchors, which drop the prefix, so
those pages 404 on prefixed instances (#2910);
`CompanyExport.filePathFromLocation` also failed to locate
`/company/export/files/` inside prefixed URLs, breaking export file
previews
> - #2951 fixed the Settings links with `<Link to>` plus tests, but the
sandbox-settings work in #4415 reverted the links back to `<a href>`,
silently reintroducing the bug (#6647)
> - This pull request restores the prefix-aware `<Link>` for all four
Settings links, normalizes the pathname with `toCompanyRelativePath()`
before matching the export-files marker, and adds regression tests
covering every route so the fix cannot be lost again
> - The benefit is that export/import/org-chart/cloud-upstream
navigation and export file previews work on every prefixed deployment

## Linked Issues or Issue Description

Fixes: #6647
Refs #2910 (original report: `/company/export` → prefix `COMPANY` → not
found)
Refs #2951 (original fix with `<Link to>` + tests — merged, then lost)
Refs #4415 (sandbox settings PR that reverted Settings back to `<a
href>`)

## What Changed

- **`ui/src/pages/CompanySettings.tsx`**: use `Link` from `@/lib/router`
for the Org Chart, Export, Import, and Cloud Upstream buttons (replacing
raw `<a href>`)
- **`ui/src/pages/CompanyExport.tsx`**: resolve file paths from prefixed
URLs by normalizing with `toCompanyRelativePath()` before matching
`/company/export/files/`
- **`ui/src/lib/company-routes.test.ts`**: regression tests for
export/import/cloud-upstream/org prefix rewriting, double-prefix
prevention, and export file URL normalization

## Verification

```bash
pnpm vitest run ui/src/lib/company-routes.test.ts
```

Manual:

1. Open `http://localhost:3100/NEU/company/settings` (or your company
prefix).
2. Click **Export** / **Import** — URL should stay under
`/:prefix/company/...`.
3. On export, select a file — URL should be
`/:prefix/company/export/files/...` and preview should load.

The change is navigation-target-only (no visual/layout changes), so
before/after is shown as the resolved URLs:

| Link | Before (prefix dropped → not found) | After |
|------|-------------------------------------|-------|
| Export | `/company/export` | `/NEU/company/export` |
| Import | `/company/import` | `/NEU/company/import` |
| Org Chart | `/org` | `/NEU/org` |
| Cloud Upstream | `/company/settings/cloud-upstream` |
`/NEU/company/settings/cloud-upstream` |

## Risks

Low — same approach as #2951; only navigation/parsing, no API changes.

## Model Used

- Original implementation: authored by @qbamca in Cursor (agentic
editor; the session's exact model ID was not recorded)
- Follow-up commit (merge-conflict resolution) and this description
update: Claude Fable 5 (Anthropic, `claude-fable-5`, extended thinking,
agentic tool use), operated by the Commit Capital triage team

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
doc changes required — behavior matches documented routing)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending re-review of the conflict-resolution commit)
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Jakub Mikiciuk <jmikiciuk@igus.net>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:52:40 -05:00
Dotta df2404d443
fix(codex): count raw child output as activity (#9632)
## Thinking Path

> - Paperclip is the open source control plane teams use to manage AI
agents and their work
> - The Codex local adapter supervises CLI child processes and
terminates genuinely silent runs
> - The existing inactivity timer only recognized parsed JSONL stdout
events as activity
> - Long verification commands can emit ordinary stdout or stderr while
producing no JSONL, so healthy children could be killed
> - This pull request makes the watchdog observe raw child-process
output before filtering or parsing
> - The benefit is that long typecheck, build, and test phases survive
while truly silent children remain bounded

## Linked Issues or Issue Description

- **Bug:** The Codex output-inactivity watchdog could terminate healthy
runs during long verification phases because non-JSON stdout and stderr
did not reset its timer.
- **Expected:** Any bytes emitted by the child process count as output
activity; only a child with no stdout or stderr for the configured
interval is terminated.
- **Reproduction:** Configure a short `outputInactivityTimeoutMs`, run a
Codex child that periodically emits plain-text verification progress
without JSONL events, and observe the old monitor firing despite
continued output.
- **Deployment mode:** Local `codex_local` adapter execution.
- Related implementation: #5017
- Related recovery behavior: #8680

## What Changed

- Reset the Codex inactivity monitor on every non-empty stdout or stderr
chunk before stderr noise filtering.
- Track raw output chunk and byte counts in monitor diagnostics while
retaining parsed JSONL event counts.
- Add regression coverage for more than 21 simulated minutes of non-JSON
verification output and retain silent-child termination coverage.
- Document that `outputInactivityTimeoutMs` observes raw child output
and that `null` still disables the monitor.

## Verification

- `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/vitest
run
packages/adapters/codex-local/src/server/output-inactivity-monitor.test.ts
packages/adapters/codex-local/src/server/output-inactivity-monitor.integration.test.ts`
- `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/tsc -p
packages/adapters/codex-local/tsconfig.json --noEmit`

## Risks

- Low risk: the monitor becomes more conservative and may allow a
noisy-but-stuck child to run longer, but the configured hard timeout and
platform silent-run safety net remain unchanged.
- No schema, API, migration, or UI changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex CLI coding agent; runtime model ID and context-window
size were not exposed to the agent. Used repository/tool access, code
execution, and focused test verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 21:49:30 -05:00
Dotta 8eff54bc47
[codex] Explain AWS secret creation failures in the UI (#9645)
## Thinking Path

> - Paperclip is the open source control plane operators use to manage
AI-agent companies.
> - Operators can store runtime secrets in provider vaults such as AWS
Secrets Manager.
> - The server already preserves sanitized AWS failure details,
including the failed operation, required IAM capability, region, and
safe recovery options.
> - The create-secret dialog reduced that structured response to a
generic message, leaving operators unable to understand or fix failed
AWS writes.
> - This pull request keeps the safe structured error through the UI and
presents concise, actionable diagnostics without exposing raw AWS
principals or account details.
> - The benefit is that operators can correct IAM access or link an
existing AWS secret immediately instead of debugging an opaque failure.

## Linked Issues or Issue Description

No public issue was found for this exact UI gap.

Bug report:
- What happened: creating a Paperclip-managed value in AWS Secrets
Manager could fail with a generic dialog error even though the API
returned safe, actionable provider details.
- Expected behavior: the dialog should identify the AWS operation,
required IAM capability, region, provider vault, and safe alternative
while keeping raw cloud-provider details redacted.
- Steps to reproduce: configure an AWS Secrets Manager provider vault
without `secretsmanager:CreateSecret`, then create a Paperclip-managed
secret using that vault.
- Paperclip version/commit: current `master` before this PR.
- Deployment mode: Paperclip server with an AWS Secrets Manager provider
vault.

Related prior server-side propagation work:
- Refs #9161

## What Changed

- Preserve the structured `ApiError` returned by failed create-secret
mutations instead of reducing it to a string.
- Render AWS-specific, sanitized diagnostics with the required IAM
capability, region, provider vault, operation, and external-reference
recovery option.
- Add a full dialog render regression test that verifies actionable
details appear and raw AWS ARN/account information does not.

## Verification

- `pnpm exec vitest run ui/src/pages/Secrets.render.test.tsx` — 1 file
passed, 15 tests passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm check:token-gates` — all gates clean.
- `git diff --check public-gh/master...HEAD` — passed.

### Visual Verification

QA verified both states in Chromium at head `357cd2271` using the real
`Secrets` component and confirmed that raw AWS account, ARN,
assumed-role, and provider exception details are absent from the
rendered DOM.

**AWS access-denied diagnostics**

![AWS access-denied
diagnostics](https://pages.paperclip.ing/pap-14130-secret-error/aws-access-denied.png)

**Generic non-AWS fallback**

![Generic non-AWS
fallback](https://pages.paperclip.ing/pap-14130-secret-error/generic-error.png)

## Risks

Low risk. The change only affects failed create-secret presentation in
the UI; successful secret creation, API contracts, schema, and
migrations are unchanged. The structured details are server-sanitized,
and the regression test confirms raw AWS principal/account data is not
rendered.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI `gpt-5.4` via Codex CLI, with tool-enabled repository editing,
shell execution, Git, GitHub, and Paperclip API access. Reasoning mode
and exact context-window size were not exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 21:47:07 -05:00
Dotta 8368fb30b0
fix(routines): coalesce sub-hourly catch-up runs (#9649)
## Thinking Path

> - Paperclip is the open source control plane people use to run
AI-agent companies.
> - Scheduled routines support catch-up policies when the server resumes
after missed cron ticks.
> - The existing capped replay policy dispatched once per missed tick,
which can flood the board after downtime for frequent schedules.
> - Sub-hourly routines usually need one prompt catch-up execution
rather than historical per-tick replay, while hourly-or-slower schedules
may rely on the existing behavior.
> - This pull request coalesces missed sub-hourly ticks into one
execution and keeps the slower-schedule behavior unchanged.
> - The benefit is bounded recovery work without changing the semantics
of lower-frequency scheduled routines.

## Linked Issues or Issue Description

### What happened?

When a scheduled routine using `enqueue_missed_with_cap` resumes after
several missed sub-hourly cron ticks, Paperclip dispatches one catch-up
execution for every missed tick. Those executions arrive in a
same-second burst and can flood the board with duplicate-looking work.

### Expected behavior

Sub-hourly schedules should advance past all missed ticks but dispatch
exactly one catch-up execution. Hourly-or-slower schedules should retain
capped per-tick replay.

### Steps to reproduce

1. Build Paperclip from `master` and create a routine with a sub-hourly
cron schedule and `catchUpPolicy: enqueue_missed_with_cap`.
2. Set its persisted `nextRunAt` far enough in the past to cover several
scheduled occurrences.
3. Run routine catch-up processing.
4. Observe multiple catch-up dispatches instead of one coalesced
execution.

### Paperclip version or commit

Reproduced on `master` before this PR.

### Deployment mode

Built from source in local development with embedded PGlite.

## What Changed

- Classify sub-hourly cadence from timezone-aware scheduled occurrences,
avoiding daily multi-minute false positives while supporting schedules
restricted to active days.
- Coalesce all missed sub-hourly ticks into one catch-up dispatch while
advancing `nextRunAt` to the next future occurrence.
- Preserve capped per-tick replay for hourly-or-slower schedules.
- Clarify the catch-up policy labels in both routine editing surfaces.
- Add regression coverage for both the coalesced and preserved
behaviors.

## Verification

- `pnpm exec vitest run server/src/__tests__/routines-service.test.ts
--testNamePattern='coalesces multiple missed sub-hourly ticks|continues
replaying each missed hourly tick|continues replaying missed ticks for
daily schedules with multiple minute values|coalesces sub-hourly
schedules restricted to weekdays'` — 4 passed.
- `pnpm check:token-gates` — all gates clean.
- `git diff --check origin/master...HEAD` — clean.

## Risks

- Low-to-moderate behavioral risk: sub-hourly routines using
`enqueue_missed_with_cap` now intentionally receive one recovery
execution instead of one per missed tick.
- Hourly-or-slower schedules retain their previous capped replay
behavior, limiting the compatibility surface.
- No schema, migration, workflow, or lockfile changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex CLI with GPT-5.5, medium reasoning, code execution and
repository tool use; the runtime did not expose a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 21:46:10 -05:00
Dotta bd7c0d5f83
fix(issues): deduplicate repeated creates (#9650)
## Thinking Path

> Paperclip already treats issue creation as a company-scoped mutation,
but retries and parallel agent heartbeats can submit the same create
more than once. Client instructions cannot provide at-most-once behavior
under concurrency, so the guard belongs in the server transaction. This
change adds an explicit company-scoped idempotency contract, a
conservative fallback for recent open same-parent titles, and run
attribution for auditability. Advisory transaction locks serialize
competing requests before lookup/insert, avoiding the race that affected
the prior attempt.

## Linked Issues or Issue Description

Fixes #6529.

This is a clean replacement for #6936, which was closed because it mixed
unrelated changes and its check-then-insert implementation was not
concurrency-safe. Unlike that attempt, this PR is scoped to eight files,
uses a dedicated idempotency-key table, and serializes duplicate
candidates inside the create transaction.

## What Changed

- Accept optional `idempotencyKey` and `allowDuplicate` fields on issue
creation.
- Replay the existing issue with HTTP 200 and deduplication metadata for
a repeated company/key pair.
- Deduplicate recent open issues with the same company, parent, and
normalized title for 48 hours unless `allowDuplicate: true` is supplied.
- Persist idempotency mappings in a company-scoped table and serialize
competing creates with transaction advisory locks.
- Populate `originRunId` from `X-Paperclip-Run-Id` for agent/manual
creates when the body does not provide an origin run.
- Add route integration coverage for key replay, title fallback, bypass,
closed/old recreation, company scoping, and run attribution.

## Verification

- `pnpm exec vitest run
server/src/__tests__/issue-create-deduplication-routes.test.ts` — 7
tests passed.
- `pnpm --filter @paperclipai/db typecheck` — passed, including
migration numbering and safety checks.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check origin/master...HEAD` — passed.
- `pnpm exec vitest run
server/src/__tests__/issue-assigned-backlog-contract-routes.test.ts
server/src/__tests__/issue-create-deduplication-routes.test.ts` — 10
tests passed after the service-contract compatibility fix.

## Risks

- The title fallback intentionally treats normalized same-parent titles
as duplicates for 48 hours; callers creating intentionally repeated
titles must send `allowDuplicate: true`.
- Advisory locks use hashed duplicate keys, so an extremely unlikely
hash collision can serialize unrelated creates but cannot merge their
lookup results.
- Deleting an issue cascades its idempotency mapping, allowing the same
key to create a replacement later.

## Model Used

- OpenAI `gpt-5.6-sol`, high reasoning effort, Codex CLI with
repository, shell, GitHub CLI, and Paperclip API tool access.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 21:45:11 -05:00
Dotta ea0e899905
fix(search): honor extract match limits + harden pr-gardening candidate discovery (#9652)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The `/pr-gardening` skill drives a bundled agent that scans a
company's issues for those linked to open GitHub PRs, then reports on
their state; it relies on the server's company-search **extract**
endpoint to pull PR references out of issue bodies
> - Two gaps surfaced during end-to-end QA of the gardening workflow:
the extract service silently ignored a per-issue match cap, so callers
could not bound how many matches came back per issue, and the skill's
candidate-discovery scripts fell over on large repos and on issues that
referenced deleted PRs
> - Left unaddressed, the gardener either truncated its scan
unpredictably or aborted outright, so it could not reliably enumerate PR
candidates
> - This pull request honors an explicit `matchesPerIssue` limit in the
extract search API and hardens the skill's candidate discovery against
missing/unavailable PRs and oversized `gh` output
> - The benefit is a PR-gardening workflow that scans deterministically
and finishes cleanly on real-world companies

## Linked Issues or Issue Description

No pre-existing public GitHub issue — describing the bug in-PR following
the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`).

### What happened?

The company-search extract endpoint accepted a per-issue match limit but
did not apply it, returning matches capped only by the old hardcoded
constant regardless of the caller's request. Separately, the
`/pr-gardening` skill's candidate-discovery scripts crashed when a
scanned issue referenced a deleted PR (GitHub `Not Found (HTTP 404)` /
GraphQL `Could not resolve to a PullRequest`) and could exceed the
default `gh` output buffer on large result sets, aborting the whole
scan.

### Expected behavior

The extract API bounds matches per issue when a caller passes
`matchesPerIssue` (default 20, max 200), and omitting it preserves the
previous default. The gardening scripts skip PRs that are
deleted/unavailable and tolerate large `gh` responses without aborting
the scan.

### Steps to reproduce

1. Call the company-search extract endpoint with a `matchesPerIssue`
value against an issue containing many PR references — previously the
value was ignored.
2. Run the pr-gardening candidate scan against a company whose issues
reference a since-deleted PR — previously the scan threw instead of
skipping that PR.

### Paperclip version or commit

`master` at the base of this PR (branch cut from current
`origin/master`).

### Deployment mode

Local Paperclip instance / self-hosted.

## What Changed

- **Extract search honors `matchesPerIssue`**: added the
`matchesPerIssue` field to the shared search validator/types and applied
the cap in `company-search-extract` so results are bounded per issue
(`packages/shared`, `server/src/services/company-search-extract.ts`,
`doc/SPEC-implementation.md`).
- **Hardened pr-gardening candidate discovery**: `find-candidates.mjs` /
`lib.mjs` now request `matchesPerIssue=200`, treat missing/unavailable
PRs (deleted PR → `isMissingPullRequestError` / `unavailable`) as skips
instead of fatal errors, and raise the `gh` `maxBuffer` to 50 MB for
large repos.
- **Tests**: expanded `company-search-extract-{routes,service}.test.ts`
for the new limit and added coverage in `pr-gardening.test.mjs`.

## Verification

Re-run on a fresh worktree cherry-picked onto current `master`:

- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` → 8/8 pass
- `pnpm vitest run
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-extract-service.test.ts` → 10/10
pass

## Risks

Low risk. `matchesPerIssue` is optional and backward-compatible
(omitting it preserves prior behavior). The skill changes only add
skip/tolerance paths and a larger buffer; no schema or migration
changes.

## Model Used

Claude — Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use /
code execution via the Claude Agent SDK.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 21:44:21 -05:00
Dotta 5588ddf681
fix(server): prevent recurring worktree port conflicts (#9642)
## Thinking Path

> - Paperclip is the control plane operators use to run AI-agent
companies and their isolated development workspaces.
> - Worktree startup assigns each workspace a server port and an
embedded PostgreSQL port.
> - Existing collision detection depended on discovering sibling configs
from the current repository layout, so worktrees in different repository
roots could select the same ports.
> - Concurrent startup also had no shared critical section, allowing two
worktrees to observe the same available ports before either persisted
its selection.
> - Repeated collisions prevented otherwise isolated workspaces from
starting reliably and could recur after a port was repaired once.
> - This pull request adds a shared, locked registry of active worktree
config paths and uses it during port selection and repair.
> - The benefit is stable, persisted, cross-repository port isolation
for both the Paperclip server and embedded PostgreSQL.

## Linked Issues or Issue Description

### What happened?

When multiple Paperclip worktrees shared the same worktree home but
lived under different repository roots, startup could assign duplicate
server and embedded PostgreSQL ports. The prior sibling scan did not
reliably discover configs outside the current repository, and
simultaneous repairs were not serialized.

### Expected behavior

Each active worktree should reserve unique server and database ports
across repository roots, persist any repaired selection, and reuse the
persisted ports on subsequent starts.

### Steps to reproduce

1. Create two Paperclip worktrees in different repository roots that
share `PAPERCLIP_WORKTREES_DIR`.
2. Give both worktree configs the same server and embedded PostgreSQL
ports.
3. Start or repair both worktrees.
4. Observe that both can retain the same ports because neither reliably
discovers the other configuration.

### Environment

- Version: reproducible on `master` before this change
- Deployment: local development worktrees built from source
- Adapter: not adapter-specific
- Database: embedded PostgreSQL

Related prior reliability work: #1829. Related documentation for
recovering port conflicts: #9407.

## What Changed

- Add a shared `worktree-port-reservations.json` registry under the
worktree home, containing live worktree config paths.
- Serialize registry reads, collision detection, config repair, and
registry updates with a stale-safe filesystem lock.
- Include registered configs and isolated instance configs when
collecting reserved server and embedded PostgreSQL ports.
- Atomically prune stale registry entries and persist repaired ports
plus the matching public base URL.
- Add regression coverage for cross-repository collisions, persisted
repairs, and repeat startup behavior.

## Verification

- `pnpm exec vitest run server/src/__tests__/worktree-config.test.ts` —
14 tests passed, including stale-lock recovery.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- Rebased onto current `public-gh/master` before verification.

## Risks

- Low-to-moderate risk: worktree startup now briefly acquires a
filesystem lock in the shared worktree home.
- The lock has a 10-second acquisition timeout and removes lock
directories older than 5 seconds so interrupted owners are recoverable
within the wait window.
- Registry writes are atomic and stale config paths are pruned, limiting
persistent state to existing worktree configs.
- The change is scoped to worktree runtime configuration and does not
affect normal main-instance configuration.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.3 Codex and GPT-5.4 with repository access,
terminal execution, and code-review tooling.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 20:09:44 -05:00
Dotta f1508a7929
feat(decisions): expand image rows into gallery (#9532)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies and review work that needs human attention.
> - The Decisions page condenses approvals, failed runs, reviews, and
issue interactions into a scannable attention queue.
> - Some decision rows already include screenshot evidence, but the
collapsed thumbnail stack is too small for meaningful inspection.
> - Image-only rows were not expandable because expansion was previously
reserved for inline decision resolvers.
> - Reviewers therefore had to leave the Decisions page before they
could understand the visual evidence attached to a row.
> - This pull request makes rows with images expandable and presents a
readable, linked gallery while preserving the compact collapsed view.
> - The benefit is faster evidence review without sacrificing queue
density or existing inline-resolution behavior.

## Linked Issues or Issue Description

No public GitHub issue currently tracks this feature.

### Subsystem affected

`ui/` — React + Vite board UI

### Problem or motivation

Screenshot evidence in Decisions rows is only visible as small
overlapping thumbnails, and non-inline rows cannot expand to show it.
This makes visual review unnecessarily slow and forces reviewers to
navigate away from the queue.

### Proposed solution

Treat active rows with images as expandable, render the first three
images at a readable size in an expanded gallery, and link images plus
any remaining-image affordance to the related issue.

### Alternatives considered

Always rendering large images would make the queue difficult to scan;
opening the issue immediately preserves density but prevents in-context
review. An explicit expandable gallery keeps both behaviors available.

### Roadmap alignment

`ROADMAP.md` does not list overlapping Decisions image-gallery work.
This is a focused improvement to the existing review surface rather than
a new product area.

### Additional context

The Storybook variants document both the collapsed thumbnail treatment
and deterministic expanded gallery state for reviewer inspection.

## What Changed

- Allow active Decisions rows with screenshot evidence to expand even
when they have no inline resolver.
- Keep compact thumbnails in collapsed rows and render up to three
larger, linked images when expanded.
- Add an accessible remaining-image tile that links to the related issue
when more screenshots exist.
- Add component coverage for image-only expansion and the
remaining-image issue link.
- Add collapsed and expanded image-gallery Storybook variants, including
deterministic initial expansion.

## Verification

- `pnpm exec vitest run ui/src/components/AttentionQueueRow.test.tsx`
- `pnpm check:token-gates`
- `pnpm --dir ui typecheck`
- `pnpm --dir ui build-storybook`
- QA visual verification (light + dark, no defects):
https://github.com/paperclipai/paperclip/pull/9532#issuecomment-4964769485

## Risks

- Low risk: the change is isolated to Decisions row rendering and
Storybook fixtures.
- Rows with images gain a new expansion interaction, but existing inline
resolver behavior and deep links remain intact.
- The gallery intentionally limits the in-row preview to three images to
avoid unbounded row height.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Anthropic Claude Opus 4.8 assisted with the original implementation
and tests.
- OpenAI `gpt-5.4` via Codex CLI assisted with current-master rebase
integration, verification, and PR preparation. The runtime
context-window size is not exposed; capabilities used include reasoning,
repository tool use, code execution, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 20:08:21 -05:00
Dotta d32ed88443
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI-agent companies and their work.
> - Its recovery subsystem detects stranded issue execution and decides
whether to retry, escalate, or request operator intervention.
> - The existing recovery path used a mostly generic owner ladder and
generic execution contract, so transient failures could wake a manager
who then performed the deliverable instead of repairing and returning
the task.
> - Provider quota failures also entered the same takeover path even
when the correct action was to wait for capacity and retry the original
assignee.
> - Recovery actions already retain the source owner and evidence needed
to choose a cause-specific route, render a scoped contract, and measure
whether work was handed back.
> - This pull request adds a cause-keyed recovery playbook, propagates
its contract through every built-in adapter, and makes resolved recovery
actions return work to the original owner by default.
> - The benefit is bounded self-recovery that preserves task ownership,
avoids needless management takeover, and makes recovery outcomes
observable.

## Linked Issues or Issue Description

No matching public GitHub issue was found.

Related recovery work was reviewed but is not duplicated here: #9630
restores bounded recovery continuations, #8807 changes one
assignee-ranking case, and #9404 records runtime-failure transition
evidence. This change instead introduces cause-specific routing and
recovery contracts across the recovery lifecycle.

### What happened?

When an issue became stranded, recovery generally selected an owner
through the same fallback ladder and rendered the normal execution
contract. That made the recovery wake look like ordinary deliverable
work, even when the correct action was to retry the original agent,
repair its runtime, or wait for a provider quota reset.

### Expected behavior

Recovery should select a response by failure cause, tell the recipient
to recover rather than complete the deliverable, suppress takeover wakes
for provider quota waits, and return repaired work to its original
assignee unless the recovery owner explicitly completes it.

### Actual behavior

Recovery could escalate transient failures to management, omit the
cause-specific next action from the wake, and leave the recovery owner
assigned after the runtime problem was resolved.

### Impact

The generic path creates avoidable management work, ownership churn, and
budget consumption while obscuring whether recovery successfully
returned work to the responsible agent.

## What Changed

- Added cause-keyed routing for process loss, missing disposition,
provider quota limits, Codex output inactivity, workspace validation
failures, and fallback recovery causes.
- Added recovery-scoped wake rendering that replaces the generic
execution contract with the failure summary, original assignee, attempt
count, next action, and cause-specific playbook instruction.
- Propagated the structured recovery contract through all built-in
adapter execution paths, including Hermes local and gateway adapters.
- Added provider-quota wait monitoring so capacity failures schedule the
original assignee instead of enqueueing a takeover wake.
- Added hand-back behavior and `handed_back` / `owner_completed` outcome
accounting when recovery actions are resolved.
- Added focused routing, renderer, quota-monitor, and hand-back
regression coverage plus implementation-spec documentation.

## Verification

- `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts
server/src/__tests__/issue-recovery-actions.test.ts`
  - 4 test files passed; 194 tests passed.
- Targeted `pnpm --filter ... typecheck` across
`@paperclipai/adapter-utils`, `@paperclipai/shared`,
`@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter
packages.
  - 13 affected workspace packages passed typecheck.
- `pnpm check:token-gates`
  - All UI token gates passed.

## Risks

- Recovery routing behavior changes for stranded work, so an incorrectly
classified cause could select a different recipient than before;
fallback causes retain the existing management ladder.
- Provider quota detection depends on structured failure evidence and
conservative text matching; unmatched failures continue through fallback
recovery.
- Adapter prompt plumbing changes across built-ins, covered by shared
renderer tests and compile-time call signatures.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool
use, and code execution. The runtime does not expose its configured
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 20:04:42 -05:00
Dotta 16b95eece5
fix(server): preserve source SHA without Git metadata (#9638)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies
> - Operators need to identify the exact source build running from the
persistent account menu
> - PR #9508 added linked source SHA metadata when the server can
inspect its Git checkout
> - Production images and packaged deployments may not include a `.git`
directory even though their build commit is known
> - Falling back to the package version in those environments makes the
UI look like a formal release and hides the source SHA
> - This pull request reads a validated deployment commit marker when
Git metadata is unavailable and uses it consistently for server version
and server-info responses
> - The benefit is that unreleased deployments keep showing an
inspectable SHA without changing exact-tag release versions

## Linked Issues or Issue Description

Follow-up to #9508.

### Pre-submission checklist

- [x] I searched existing open and closed issues and found no duplicate
for the no-`.git` deployment fallback.
- [x] The behavior reproduces when the server runs without Git metadata
but has a known build commit.
- [x] The behavior originates in Paperclip's core server build metadata
handling, not an adapter, provider, or local configuration.

### What happened?

PR #9508 displays source branch and SHA metadata for unreleased builds,
but server version and server-info resolution still fall back to the
package version when the runtime has no `.git` directory. This is common
in production images and packaged deployments.

### Expected behavior

When a validated deployment commit is available through
`PAPERCLIP_BUILD_COMMIT` or `/app/.paperclip-build-commit`, the server
should retain a derived source version and expose SHA metadata even if
Git commands are unavailable. Exact release tags should continue using
the formal package version.

### Steps to reproduce

1. Build or run Paperclip without a `.git` directory.
2. Provide a full commit SHA through `PAPERCLIP_BUILD_COMMIT` or
`/app/.paperclip-build-commit`.
3. Start the server and inspect the version and server-info output.
4. Observe that current `master` returns only the package version and
reports Git metadata unavailable.

### Paperclip version or commit

Current `master` after #9508.

### Deployment mode

Packaged or containerized deployments without runtime Git metadata.

### Installation method

Built from source or deployment image.

## What Changed

- Add validated build-commit parsing from `PAPERCLIP_BUILD_COMMIT` and
`/app/.paperclip-build-commit`.
- Preserve source-derived server versions when Git commands are
unavailable.
- Expose fallback SHA metadata through server-info with an explicit
unavailable local-status state.
- Keep exact release-tag builds on the formal package version.
- Add focused regression tests for parsing, version resolution, and
server-info fallback behavior.

## Verification

- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/build-commit.test.ts src/__tests__/server-info.test.ts
src/__tests__/version.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- `git diff --check public/master...HEAD`

## Risks

- Low risk: only full 40-character hexadecimal commit values are
accepted; malformed or truncated markers preserve the existing fallback
behavior.
- Deployment tooling must set `PAPERCLIP_BUILD_COMMIT` or write
`/app/.paperclip-build-commit` for the fallback to activate.
- Fallback server-info cannot provide branch, subject, commit time, or
working-tree status without Git metadata, so those fields remain
explicitly unavailable.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.4 with medium reasoning, repository/tool
access, shell execution, and code editing; context-window size was not
exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates pass
- [x] Greptile review is 5/5 with no open P2-or-higher comments,
recommendations, or follow-ups

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 20:03:52 -05:00
Dotta b606869a6a
feat(skills): add active PR gardening workflow (#9510)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its skills layer gives agents repeatable operational workflows
without embedding every procedure in core orchestration
> - Pull requests referenced across active issue threads currently
require expensive manual discovery and inconsistent readiness checks
> - Candidate extraction and current-head verification can be
deterministic, read-only code paths instead of LLM scanning
> - This pull request adds an active PR gardening skill with discovery,
readiness, reporting, and originating-issue follow-up instructions
> - The benefit is a token-efficient, auditable way to identify PRs that
are ready or need attention while preserving a strict no-merge guardrail

## Linked Issues or Issue Description

- **Problem:** Paperclip operators lack a fleet-level workflow to
discover PRs mentioned by recently active issues, verify their exact
current-head readiness, and route actionable failures back to the issue
that owns the work.
- **Desired behavior:** Use the company extract-search API plus
read-only GitHub inspection to deduplicate open PRs, classify readiness
and confidence, avoid nagging drafts, and render an inspectable report.
- **Safety requirements:** The workflow must never merge, approve, or
close PRs; must never issue mutating GitHub calls; and must suppress
Paperclip comments in dry-run mode.
- No duplicate PR was found. Related historical PR #3725 concerns skill
endpoint permissions rather than PR gardening.

## What Changed

- Added `.agents/skills/pr-gardening/SKILL.md` with discover, verify,
comment, monitor, and report stages plus cooldown and max-round
guidance.
- Added `find-candidates.mjs` to page through extract-search results,
normalize/deduplicate PRs, map source issues and work products, and drop
closed GitHub PRs.
- Added `check-readiness.mjs` to inspect current-head checks, Greptile
freshness, conflicts, reviews, and base distance with machine-readable
reasons.
- Added `render-report.mjs` to group open PRs into High, Medium, and Low
merge-confidence sections.
- Added focused Node tests, including an explicit assertion that scripts
contain no mutating GitHub commands.

## Verification

- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 8 tests
passed.
- Live read-only GitHub dry run at current heads: PR #9493 classified
High/ready; PR #9473 classified Medium because it is two commits behind
`master`; merged PR #9470 is excluded from readiness candidates.
- The deployed control plane currently returns 404 for
`/api/companies/:companyId/search/extract`; direct Stage A live
execution will be repeated after the separately prepared extract-search
endpoint is deployed.

## Risks

- Low runtime risk: this adds a repo skill and read-only scripts, not a
server execution path.
- The discovery script intentionally fails closed when extract-search
reports truncated matches, preventing silently incomplete reports.
- Readiness reflects GitHub state at execution time and must be rerun
after any head update.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex using GPT-5.4 with medium reasoning, repository tool use,
shell execution, and live GitHub/Paperclip API verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 19:05:44 -05:00
Dotta ae77908618
feat(search): add bulk extract endpoint (#9507)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI-agent companies
> - Agents and operators need company-scoped search to discover relevant
issue history safely
> - The interactive search endpoint intentionally returns compact
excerpts and low pagination caps for UI use
> - Automation that inventories repeated references, such as
pull-request URLs, needs exhaustive distinct matches without loading
full issue objects into an LLM context
> - Client-provided regular expressions would create an unsafe and
expensive query surface, so extraction must remain literal with
server-owned expansion modes
> - This pull request adds a bounded agent-oriented extraction endpoint
with explicit truncation
> - The benefit is deterministic, compact bulk discovery across issues,
comments, and documents while preserving company authorization and rate
limits

## Linked Issues or Issue Description

### Subsystem affected

`server/` REST API and `packages/shared/` contracts.

### Problem or motivation

The existing interactive company search caps issue pagination and
snippets, so automation cannot reliably enumerate every distinct literal
or pull-request URL across issue descriptions, comments, and linked
documents without fetching large full issue payloads.

### Proposed solution

Add `GET /api/companies/:companyId/search/extract` with escaped literal
matching, optional server-owned URL token expansion,
issue/comment/document scopes, status/date filters, higher issue-level
pagination caps, compact source references, and explicit
pagination/match truncation flags.

### Alternatives considered

Reusing `GET /issues?q=` would return unnecessarily large issue objects;
increasing interactive-search snippet limits would make the UI API
heavier; accepting arbitrary client regex would expose avoidable
database cost and ReDoS risk.

### Roadmap alignment

`ROADMAP.md` does not currently list a conflicting company-search or
bulk-extraction initiative. GitHub searches found no directly
duplicative open issue or pull request.

## What Changed

- Added shared query validation and response contracts for literal and
URL extraction.
- Added a company-scoped extraction service that pages issues, gathers
matching issue/comment/document sources, expands URL tokens,
deduplicates values, and reports truncation explicitly.
- Added the authenticated route using the existing company-search
authorization decision and rate limiter.
- Added targeted Vitest coverage for URL extraction, multi-source
dedupe, date/status filters, match caps, cross-company denial, and rate
limiting.
- Documented the extraction surface in the implementation specification.

## Verification

- `pnpm exec vitest run
server/src/__tests__/company-search-extract-service.test.ts
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-rate-limit-routes.test.ts
server/src/__tests__/company-search-service.test.ts` — 30 tests passed.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check` — passed.

## Risks

- Bulk substring search can scan large text columns. The endpoint
mitigates this with a minimum literal length, bounded issue pagination,
a 20-distinct-match cap per issue, explicit truncation, existing
company-search rate limiting, and no client-provided regex.
- URL expansion uses a fixed server-owned pattern plus an escaped
literal. A security review is requested as part of PR review to confirm
the pattern and abuse controls.
- No database migration or existing API response shape changes are
included.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex CLI coding agent; exact runtime model ID and
context-window size were not exposed to the session. Tool-enabled code
execution and repository editing were used with medium reasoning effort.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 19:05:06 -05:00
Dotta 3ae2c30f2f
feat(skills): import skills from projects (#9620)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies.
> - Company skills make reusable agent behavior discoverable and
editable from one place.
> - Projects already contain skill directories, but operators had to
import each skill path manually.
> - Copying those skills would break the desired write-through workflow
between Skill Studio and the source project.
> - The server therefore needs a safe preview/select/import contract
that only accepts rediscovered, workspace-contained candidates.
> - The UI needs a guided project picker that explains reference
semantics, handles conflicts, and remains usable on mobile.
> - This pull request adds that end-to-end project skill import flow
with authorization, tenant-scope, traversal, and symlink regression
coverage.
> - The benefit is faster bulk onboarding while keeping project files as
the single source of truth.

## Linked Issues or Issue Description

**Feature request**

**Problem:** Importing several skills already stored in a Paperclip
project requires operators to discover and submit each local path
individually. This is slow, hides which well-known directories were
searched, and makes conflict/already-imported states difficult to
evaluate before mutation.

**Proposed solution:** Add an “Import skills from project” flow that
previews skills from well-known directories, lets operators selectively
import eligible candidates, and stores local-path references so Skill
Studio edits write through to the project files.

**Alternatives considered:** Copying files into company-managed skill
storage was rejected because it creates divergent copies. Trusting
client-supplied paths was rejected because imports must be constrained
to server-rediscovered, workspace-contained candidates.

**Additional context:** GitHub duplicate search found no existing issue
or PR for this exact workflow. Refs #3799 for related skill-import
inventory behavior; this PR does not claim to close that issue.

## What Changed

- Extend `scan-projects` with backward-compatible preview and
selective-import modes, typed validation, candidate statuses, and
OpenAPI coverage.
- Discover project skills under `skills`, `.agents/skills`,
`.claude/skills`, `.codex/skills`, `.cursor/skills`, `.opencode/skills`,
and `.gemini/skills`.
- Re-discover selections server-side, enforce company/project/workspace
scope, and reject traversal or symlink escapes before creating
`local_path` references.
- Add the Skills-page menu entry and responsive project import dialog
with project selection, grouped candidates, select all/deselect all,
conflicts, empty/error/403 states, and import results.
- Add route, service, and component regressions for preview
authorization, cross-tenant selections, traversal/symlink safety,
selection counts, grouping, and result semantics.

### Screenshots

**Choose a project**

![Choose a
project](https://raw.githubusercontent.com/cryppadotta/paperclip-prs/refs/heads/pr-assets/import-skills-from-project/01-pick-project.png)

**Review discovered skills**

![Review discovered
skills](https://raw.githubusercontent.com/cryppadotta/paperclip-prs/refs/heads/pr-assets/import-skills-from-project/03-select.png)

**Mobile selection footer**

![Mobile selection
footer](https://raw.githubusercontent.com/cryppadotta/paperclip-prs/refs/heads/pr-assets/import-skills-from-project/select-390.png)

**Import result**

![Import
result](https://raw.githubusercontent.com/cryppadotta/paperclip-prs/refs/heads/pr-assets/import-skills-from-project/06-result.png)

## Verification

- `pnpm exec vitest run
server/src/__tests__/company-skills-service.test.ts
server/src/__tests__/company-skills-routes.test.ts
ui/src/pages/skills/ImportSkillsFromProjectDialog.test.tsx` — 3 files,
81 tests passed.
- `pnpm check:token-gates` — all token gates clean.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- Security review passed after adding tenant-scope and
unauthorized-preview regressions; UX re-review approved desktop/mobile
surfaces; QA passed all seven acceptance areas including write-through
editing, deduplication, conflicts, empty state, and permission denial.

## Risks

- Files remain referenced in project workspaces, so moving or deleting a
source directory can make an imported skill unavailable; the UI
explicitly communicates the reference behavior.
- New well-known directory scans may discover more candidates than older
versions, but preview mode prevents mutation until the operator confirms
a selection.
- The endpoint remains backward compatible: omitting `mode` preserves
the prior full-import behavior.
- No schema migration or telemetry event changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Anthropic Claude Opus 4.8 with tool use/code execution assisted with
the UI implementation and UX polish. OpenAI Codex CLI with tool use/code
execution assisted with server implementation, security fixes,
regression coverage, integration, and PR preparation; the runtime did
not expose Codex's exact backing model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 18:01:44 -05:00
Dotta 9af96461d5
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI agents and their work.
> - Its server recovery layer classifies blocked issue graphs and
restores interrupted heartbeat execution.
> - A dependent issue could remain dispatch-suppressed by a cancelled
blocker without producing operator-visible attention when the dependent
still displayed as todo or backlog.
> - Separately, a monitor-triggered run that lost its process before
disposition could consume the monitor's one-shot wake without scheduling
the existing bounded continuation.
> - Both gaps strand useful work even though Paperclip already has the
relevant blocker-attention and process-loss recovery mechanisms.
> - This pull request widens the existing classification path and reuses
the single process-loss retry for monitor dispatches with no future
wake.
> - The benefit is visible, routable recovery without weakening
dependency checkout rules or introducing an unbounded retry loop.

## Linked Issues or Issue Description

No matching public GitHub issue or pull request was found.

### What happened?

Two server recovery cases could leave work stranded:

1. A non-terminal, agent-assigned issue with an unresolved cancelled
blocker remained ineligible for checkout, but blocked-chain liveness
classification only inspected issues already displaying `blocked` or
`in_review`, so the existing `blocked_by_cancelled_issue` attention was
not surfaced.
2. A one-shot issue monitor cleared its next check when dispatched. If
that monitor-triggered run ended as `process_lost` without a tracked
local child, the existing bounded retry gate rejected it and no future
monitor wake remained.

### Expected behavior

- Cancelled blockers continue to be unresolved dependencies, and their
dependents receive blocker attention regardless of whether the dependent
currently displays as backlog, todo, blocked, or in review.
- A monitor-triggered run lost before disposition receives exactly one
bounded continuation when no future monitor check exists; a second loss
follows the normal recovery-action escalation path.

### Steps to reproduce

1. Create an agent-assigned todo issue blocked by a cancelled issue and
run issue-graph liveness classification.
2. Observe that no cancelled-blocker finding appears before this change.
3. Dispatch a due issue monitor, clear its one-shot
`monitorNextCheckAt`, and mark the resulting untracked run
`process_lost`.
4. Observe that no retry is queued before this change.

### Environment

- Paperclip commit: `3e348b96b`
- Deployment: built from source / local test environment
- Adapter: not adapter-specific; core server recovery
- Database: embedded test database

## What Changed

- Inspect non-terminal, agent-assigned issues with unresolved blocker
edges during blocked-chain liveness classification.
- Include cancelled dependents in the existing blocked-inbox attention
query while preserving company-scoped relation checks.
- Allow monitor-triggered `process_lost` runs with no future monitor
wake to use the existing single bounded retry.
- Mark monitor recovery retries as continuation-needed context and
retain the existing second-loss escalation behavior.
- Document cancelled-blocker and monitor-dispatch recovery semantics.
- Add focused regressions for liveness findings, attention propagation,
one retry, and second-loss escalation.

## Verification

- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/issue-blocker-attention.test.ts
server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests
passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.

## Risks

- Low risk and server-only. The liveness scan inspects more unresolved
dependency shapes, which can produce additional existing attention
entries for previously invisible cancelled blockers.
- Monitor recovery remains bounded by `processLossRetryCount < 1`, and
the extra path only applies when the dispatch was monitor-triggered and
no future monitor check exists.
- No schema, migration, authorization, API-contract, or UI changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool
use, command execution, and test execution capabilities.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 16:41:39 -05:00
Granis87 bbb3e19c6f
docs: add safe local worktree bootstrap guidance (#9195)
## Thinking Path

> - Parallel local agent experiments should not reuse the primary
Paperclip instance.
> - Paperclip already creates isolated worktree instances with generated
names and environment files.
> - The local development guide lacked a complete bootstrap and recovery
path.
> - This pull request adds that path using the registered CLI forms and
generated selector.
> - The benefit is safer setup, recovery, and cleanup for local
experiments.

## Linked Issues or Issue Description

No public issue was found. The local development guide did not connect
worktree creation, environment loading, startup, repair/reseed, and
cleanup into one safe sequence.

## What Changed

- Added an isolated worktree bootstrap example.
- Documented bash/zsh environment loading without presenting invalid
PowerShell syntax.
- Used registered repair/reseed commands and the generated
paperclip-local-lab selector.
- Added explicit cleanup guidance.

## Verification

- git diff --check upstream/master...HEAD
- Verified command registration and generated naming in
cli/src/commands/worktree.ts.

## Risks

Low. Documentation-only; reviewed command and selector mismatches are
corrected.

## Model Used

OpenAI GPT-5.3 Codex Spark for initial branch work; OpenAI GPT-5 Codex
for review and follow-up fixes, with repository and GitHub tool use.
Context-window sizes were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this docs PR does not
duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked or
described the problem above
- [x] I have either linked an existing issue or described the issue
in-PR
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run focused validation locally
- [x] I have added or updated tests where applicable
(documentation-only; no runtime tests needed)
- [x] I have updated the relevant documentation
- [x] I have considered and documented risks above
- [x] All Paperclip CI gates are green on the latest commit
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the latest commit
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: RobinALG87 <RobinALG87@users.noreply.github.com>
2026-07-15 15:51:53 -05:00
Dotta 02a4f52277
perf(ui): event-source the company live-runs list (Phase 1) (#9627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its web UI keeps live views fresh with React Query, coordinated by a
live-events websocket (`/api/companies/:id/events/ws`) and cross-tab
polling
> - We already cut the worst live-update churn (#9569, #9624), but the
deeper issue is that even the "push" path is *push-the-signal,
pull-the-data*: a websocket event triggers `invalidateQueries` → an HTTP
refetch
> - The company live-runs list (`queryKeys.liveRuns`) is the
most-observed resource — the sidebar renders it on nearly every page —
so its refetch is the most ambient source of churn, fired on every
`heartbeat.run.queued` / `heartbeat.run.status` event
> - Those events already carry enough (`runId`, `status`) to update the
cached list directly, so this pull request event-sources that list
instead of refetching it
> - The benefit is that the always-observed live-runs list stops
refetching on run lifecycle events — the first concrete step of the
push-over-poll redesign

## Linked Issues or Issue Description

No public GitHub issue exists; describing inline per CONTRIBUTING.md →
"Link Issues or Describe Them In-PR", following the bug report template.
This continues the memory/CPU-churn work from #9569 and #9624.

**What happened?**

Live agent-run tabs accrue high idle CPU and off-heap memory because
live-update events cause HTTP refetches. Profiling showed the company
live-runs list — observed on almost every page via the sidebar — being
refetched on every run status change, one of the most frequent ambient
refetches.

**Expected behavior**

A websocket event that already carries the changed data should update
the cached list directly, without an HTTP round-trip, so the
always-observed live-runs list does no refetch on routine run lifecycle
events.

**Steps to reproduce**

Open the app with agents running and watch the network panel: `GET
/api/companies/:id/live-runs` fires on each `heartbeat.run.status` /
`heartbeat.run.queued` event even though the event payload already
describes the change.

**Paperclip version or commit**

Branch `perf/live-runs-event-sourced`, off `master` (after #9624).

**Deployment mode**

Local dev (`pnpm dev`), web UI. Core UI live-updates plumbing; not
adapter-specific.

## What Changed

- **`ui/src/lib/live-runs-cache.ts` (new)** — pure `removeRunFromList` /
`patchRunStatusInList` helpers for the cached `LiveRunForIssue[]`.
- **`ui/src/context/LiveUpdatesProvider.tsx`** — on
`heartbeat.run.queued` / `heartbeat.run.status`, patch
`liveRuns(companyId)` in place instead of invalidating it:
  - terminal status → remove the run from the list,
  - status change on a run already in the list → update it in place,
- a genuinely new run (can't be reconstructed from the event) → fall
back to a single `invalidateQueries` refetch.
- Removed the blanket `liveRuns` invalidation from
`invalidateHeartbeatQueries`.
- On websocket **reconnect**, refetch `liveRuns` once to reconcile
events missed while disconnected (durable replay is a later phase).
- Other resources these events invalidate (`dashboard`, `costs`,
`sidebarBadges`, `agents.list`, agent detail) are unchanged — they're
lower-frequency / less-often-observed and are follow-up phases. This
keeps the change scoped and **client-only** (no server changes).

## Verification

- `vitest`: new `live-runs-cache.test.ts` (remove/patch/no-op/undefined)
and new lifecycle-handler cases in `LiveUpdatesProvider.test.ts`
(terminal→remove, present→patch, new→needs-refetch) via
`__liveUpdatesTestUtils`. All existing `LiveUpdatesProvider` tests still
pass (33 total across the two files).
- `tsc -b` clean.
- Runtime: the event-sourced path is covered by unit tests; end-to-end
refetch reduction should be re-measured against a rebuilt bundle with
the network panel / MCP instrumentation.

## Risks

Low, and client-only.
- **Staleness across a dropped connection:** an event missed while the
socket is down isn't replayed yet, so the reconnect handler refetches
`liveRuns` once to reconcile. Durable event replay (Last-Event-ID) is a
planned later phase; until then reconnect-reconcile covers the gap.
- **New-run fallback:** a genuinely new run still triggers one refetch
(it can't be reconstructed from the event alone), so no new runs are
missed.
- Aggregate resources (dashboard/costs/badges) are untouched and still
invalidate (already coalesced), so their behavior is unchanged.

Follow-up phases (from the design discussion): event-source
`activity`/comments and the remaining class-B resources; give pure-poll
resources events and drop their intervals; add durable event sequence +
reconnect replay; and a shared bus (Postgres `LISTEN/NOTIFY`) only when
the API tier scales to >1 replica.

## Model Used

- **Provider:** Anthropic, via the Claude Code CLI.
- **Model:** Claude Opus 4.8 (`claude-opus-4-8`).
- **Reasoning mode:** Extended thinking enabled.
- **Capabilities used:** tool use (shell, file editing), sub-agent
fan-out to inventory the client polling and server push infrastructure,
and the Chrome DevTools MCP to reproduce/profile the churn that
motivated this redesign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work (a perf/plumbing change, not planned core feature
work)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (continues #9569 / #9624; no duplicates)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered and documented any risks above
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no user-facing docs; rationale documented inline)
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 15:40:35 -05:00
Dotta 3e348b96b9
perf(ui): cut live-updates churn that inflates tab memory (#9624)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its web UI keeps live views (issue threads, run transcripts,
dashboard) fresh via React Query polling plus a live-events websocket,
coordinated across tabs with a `BroadcastChannel` layer
> - Long-lived tabs viewing live agent runs grew to multi-GB memory
footprints while their JS heap stayed ~60–200 MB — so the memory is
off-heap (Blink/native + committed allocator arenas), not a classic JS
leak
> - Live profiling of a reproduced 2.7 GB / 66 MB tab showed 15–30% idle
CPU, ~3 fetches/sec across overlapping poll loops, and ~8 `setInterval`
create/clear cycles per second whose rate grew ~7× as the tab aged —
relentless allocation churn that inflates committed memory the OS never
reclaims, amplified across tabs by the cross-tab fan-out
> - This pull request cuts that churn at its four largest sources
(invalidation storm, per-instance 1 s timers, redundant polling,
unbounded streamed-run set)
> - The benefit is that idle tabs do far less periodic work, so their
off-heap footprint stops ballooning over a long session

## Linked Issues or Issue Description

No public GitHub issue exists; describing inline per CONTRIBUTING.md →
"Link Issues or Describe Them In-PR", following the bug report template.

**What happened?**

Browser tabs viewing live agent runs grew to 8–16 GB memory footprint
over a long session (multiple tabs open), while each tab's "live" JS
heap stayed only ~150–250 MB. Every idle tab also burned 15–30% CPU.
Tabs eventually approached the ~4 GB V8 heap ceiling / OS pressure and
could crash.

**Expected behavior**

Tabs viewing live runs should hold a bounded footprint and do minimal
work while idle, regardless of how long they stay open or how many tabs
are open.

**Steps to reproduce**

Open several issue/run tabs that have agents actively streaming and
leave them open for a while. Watch Chrome's Task Manager: Memory
Footprint climbs into the GBs while "JavaScript Memory" stays small, and
CPU stays high on idle tabs. Reproduced in ~90 minutes: a tab reached
2.7 GB footprint on a 66 MB JS heap, and the per-second timer-churn rate
was ~7× higher on a 90-minute-old tab than a fresh one.

**Paperclip version or commit**

Branch `fix/live-updates-churn`, off `master`.

**Deployment mode**

Local dev (`pnpm dev`), web UI. Not adapter-specific — core UI
live-updates plumbing (observed with `claude_local` / `codex_local`
runs).

## What Changed

- **`ui/src/lib/query-invalidation-batcher.ts` (new)** —
`createInvalidationBatcher` throttles + de-dupes React Query
invalidations into one flush per ~300 ms, and
`createCoalescingQueryClient` wraps the client via a `Proxy` so only
`invalidateQueries` is batched (optimistic `setQueryData` writes stay
immediate). Wired into `LiveUpdatesProvider`, which previously
invalidated synchronously on every websocket event.
- **`ui/src/hooks/useSecondTick.ts` (new)** — one shared, ref-counted,
page-wide 1 s ticker. `useLiveElapsed` in `IssueChatThread` now uses it
instead of a per-instance `setInterval` that forced a full-thread
re-render every second per live element.
- **`ui/src/components/transcript/useLiveRunTranscripts.ts`** — when the
realtime websocket is enabled, the recurring log poll backs off to a 30
s safety-net cadence instead of polling every 2 s on top of the live
stream. Added a marker for the durable poll→push rearchitecture.
- **`ui/src/lib/issueChatTranscriptRuns.ts`** —
`resolveIssueChatTranscriptRuns` now caps the streamed run set
(live/active runs always kept; most-recent linked runs fill up to 20) so
a large run history can't open a live-transcript poll per historical
run.
- **`ui/src/main.tsx`** — explicit `gcTime` so cross-tab-published cache
entries for unobserved resources are collected promptly.
- Tests for the batcher, shared ticker, and run cap.

## Verification

- `vitest`: new suites `query-invalidation-batcher.test.ts` (batcher
collapses 20 invalidations → 1 flush; keeps distinct keys/variants;
dispose cancels; proxy passes non-invalidate methods through),
`useSecondTick.test.tsx` (single ref-counted timer, stops when idle),
`issueChatTranscriptRuns.test.ts` (cap keeps newest + live). All pass.
- Existing affected suites pass: `LiveUpdatesProvider` (23),
`IssueChatThread` (), `useLiveRunTranscripts`,
`AgentDetail.instructions` — 109 tests across affected files.
- `tsc -b` clean.
- Behavior confirmed by live profiling before the change (2.7 GB / 66 MB
tab, ~8 interval churns/sec growing 7× with age). Runtime churn
reduction should be re-measured against a rebuilt bundle with the same
instrumentation.

## Risks

Low-to-moderate; all changes reduce work rather than add features.
- **Invalidation batching** delays live-driven refetches by up to ~300
ms. Optimistic `setQueryData` writes (e.g. the visible issue's new
comment) remain immediate, so foreground updates still feel instant;
only the safety-net refetch is throttled. Non-live invalidations (user
actions, mutations) are unaffected — they use the real client.
- **Poll back-off** relies on the websocket as the live source when
realtime is enabled; a 30 s fallback poll still covers gaps/reconnects
(both the transcript hook and `LiveUpdatesProvider` also
auto-reconnect).
- **Run cap (20)** means an issue with a very large run history streams
live transcripts only for its live/active + 20 most-recent runs; older
runs still open normally via their run pages.
- Downstream test fallout (timing-sensitive tests around
invalidation/polling) may need adjustment — flagged intentionally for
follow-up.

Durable follow-up (out of scope, marked in code): replace transcript/run
polling with server push (SSE/websocket deltas) so idle tabs do no
periodic work at all.

## Model Used

- **Provider:** Anthropic, via the Claude Code CLI.
- **Model:** Claude Opus 4.8 (`claude-opus-4-8`).
- **Reasoning mode:** Extended thinking enabled.
- **Capabilities used:** tool use (shell, file editing), sub-agent
fan-out for codebase analysis, and the Chrome DevTools MCP to reproduce
and profile the memory/CPU churn on a live instance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work (a bug/perf fix, not planned core feature work)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered and documented any risks above
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no user-facing docs; rationale documented inline)
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 14:17:37 -05:00
Nicky Leach 3a727bf780
fix(codex): warn when sandbox auth is shadowed (#9259) 2026-07-15 10:02:28 -07:00
Dotta 89ce36d7af
feat(skills): open-by-default company skill policy and core UX (#9564)
## Thinking Path

> - Paperclip uses company skills to make agent capabilities reusable
across an organization.
> - Skill operations currently mix capability availability with
permission checks, which creates avoidable setup friction and
inconsistent denial handling.
> - The policy contract needs to remain open by default while allowing
company-scoped restrictions for governed deployments.
> - Core owns the canonical policy actions, persistence, evaluation, API
behavior, safe import boundaries, and generic denial/read-only UI.
> - Enterprise policy-editor implementation belongs in the separate
`paperclip-ee` repository and is intentionally excluded from this PR.

### Problem or motivation

Company skill operations can encounter permission dead ends even when no
explicit restriction has been configured, and import-source
classification can drift between policy evaluation and execution.

### Proposed solution

Define eight canonical skill policy actions, default all actions to
allowed, persist company-scoped restrictions, expose policy evaluation
APIs, normalize import sources at the boundary, and update Skill Studio
to present actionable restriction states without embedding Enterprise
Edition implementation in the core repository.

### Alternatives considered

Keeping capability checks distributed across routes and UI surfaces was
rejected because it duplicates policy logic and makes denial behavior
inconsistent. Shipping the Enterprise policy editor in this repository
was rejected because `paperclip-ee` is a separate repository and must
receive its own PR.

### Roadmap alignment

Extends the completed **Skills Manager** roadmap area by adding coherent
governance and removing workflow dead ends.

### Additional context

The core API contract remains suitable for a separate Enterprise Edition
editor, but this PR contains no `paperclip-ee` package or EE-specific UI
integration code.

## What Changed

- Added the company skill policy contract to product and implementation
documentation, including the open-by-default rule, eight canonical
actions, decision shape, and core/EE ownership boundary.
- Added the company-scoped policy schema, migration `0170`, shared
validators, policy service, REST routes, OpenAPI coverage, and focused
server tests.
- Hardened import policy enforcement by normalizing import sources and
keeping source classification consistent between policy evaluation and
execution.
- Updated core Skill Studio behavior to remove generic permission dead
ends and show actionable policy/platform denial states only when an
operation is actually denied.
- Removed the `plugin-paperclip-ee` package, Docker wiring, EE
discovery/deep-link helpers, and EE-specific UI tests/stories from this
PR so that implementation can be submitted separately to the EE
repository.
- Preserved open-by-default behavior when no explicit company
restriction exists.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/skill-studio/SkillPolicySurfaces.test.tsx
src/lib/skill-policy-denial.test.ts` — 20/20 passed.
- `pnpm --filter @paperclipai/ui exec tsc --noEmit` — passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/worktree-config.test.ts` — 12/12 passed.
- `pnpm check:token-gates` — passed with all gates clean.
- `git diff --check` — passed.
- `git diff --name-only origin/master | rg
'paperclip-ee|ee-skill-policy'` — no matches.

## Risks

- Migration `0170` introduces company policy persistence; rollout
depends on the migration applying before policy routes are exercised.
- Open-by-default is an intentional behavioral policy: deployments
expecting implicit denials must configure explicit restrictions.
- Import normalization is security-sensitive and should retain focused
review.
- The separate EE editor must stay contract-compatible with the core
policy API as policy actions evolve.

## Model Used

- OpenAI Codex CLI, runtime model identifier and context-window size not
exposed by this execution environment; reasoning, repository tool use,
shell execution, and code review capabilities enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available to this runtime)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked or
described the result above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green on the latest head
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Evyatar Bluzer <bluzername@users.noreply.github.com>
2026-07-15 11:42:40 -05:00
yismail 24bd860280
Stop cancelled productivity review loops (#5210)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies.
> - Productivity review reconciliation creates manager-owned review
issues when assigned work shows no-comment, long-active, or high-churn
patterns.
> - SplatImmo hit a loop because productivity-review issues were
auto-cancelled while the source issue still matched the same trigger.
> - The service already snoozed recently completed reviews, but
cancelled reviews were ignored for that snooze check.
> - This pull request treats recently cancelled productivity reviews as
terminal snooze evidence.
> - The benefit is that cancelling a review now suppresses immediate
recreation without disabling useful future productivity reviews.

## What Changed

- Renamed the recent-review lookup to terminal-review semantics and
included `cancelled` alongside `done`.
- Added a regression test proving a recently cancelled productivity
review produces `snoozed` instead of creating another review.

## Verification

- `pnpm exec vitest run
server/src/__tests__/productivity-review-service.test.ts` passes: 1
file, 12 tests.
- Queried the SplatImmo Paperclip instance for existing `Review
productivity` issues: 500 `issue_productivity_review` issues found, all
already `cancelled`, 0 active.

## Risks

- Low risk: this only affects the reconciliation branch after a terminal
productivity-review issue exists.
- Operators who cancel a productivity review now get the same default
6-hour quiet window as completed reviews; after that window, persistent
evidence can still create a fresh review.

## Model Used

- OpenAI Codex coding agent, GPT-5 class model, tool-enabled code
editing and local command execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Yanis Ismail <yanis.ismail@emissive.fr>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 11:38:37 -05:00
Dotta 0ecae2cd7e
fix(adapters): forward resolved adapter env to local agents, keep runtime env authoritative (#9617)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Local coding adapters (Claude, Codex) run each agent heartbeat in a
spawned child process; Paperclip resolves adapter-configured env —
including `secret_ref` bindings — into the env for that process
> - The resolved adapter env was not being forwarded reliably: config
env could overwrite Paperclip's own runtime vars, and a warm/resumable
ACP session could keep serving stale env because its fingerprint ignored
the resolved env
> - This mattered because a configured API key or other secret could be
silently absent from the agent shell, and a resumed session would never
pick up an updated value — while a config binding could also override
runtime identity/wake vars
> - This pull request keeps Paperclip-managed `PAPERCLIP_*` runtime env
authoritative over config, and folds a stable hash of the applied
adapter env into the session fingerprint so an env change forces a fresh
launch
> - The benefit is that adapter-configured env (plain values and
resolved secrets) reliably reaches the agent process, updates are picked
up on the next launch, and runtime identity can never be clobbered by
config

## Linked Issues or Issue Description

No public GitHub issue exists, so the underlying bug is described inline
following the bug-report template.

**What happened**

Env keys configured on a local adapter (plain values and `secret_ref`
bindings, resolved server-side into plain strings) did not reliably
reach the spawned agent process. Two distinct gaps: (1) when merging
config env into the process env, a config key in the reserved
`PAPERCLIP_*` namespace could overwrite a Paperclip-managed runtime
variable (identity, wake, workspace, API access); (2) a warm-handle /
resumable ACP session computed its reuse fingerprint from
`secretManifestHash` only, which misses plain-value edits and
same-version secret rotations — so a resumed session kept serving stale
env and never re-launched with updated values.

**Expected behavior**

Non-`PAPERCLIP_*` adapter env (plain + resolved secret values) is
forwarded to the agent process; a change to any applied forwarded value
invalidates a warm/resumable session so the next launch sources the
latest env; configured `PAPERCLIP_*` entries can never override
Paperclip runtime env, while an explicitly configured
`PAPERCLIP_API_KEY` (stable per-run config) is still honored and its
rotation also busts the session.

**Steps to reproduce**

Configure an adapter with an env key (e.g. a `secret_ref` API key) and a
resumable ACP session. On resume, the updated env value is not sourced;
separately, a `PAPERCLIP_*` config key overrides the runtime value.

**Deployment mode**

Local adapters (Claude / Codex) via the shared adapter-utils execution
path.

## What Changed

- `packages/adapter-utils/src/server-utils.ts`: add
`isPaperclipRuntimeEnvKey` and, in
`refreshPaperclipWorkspaceEnvForExecution` (used by all local adapters),
skip a `PAPERCLIP_*` config key when Paperclip has already assigned it
this run; all other keys still forward.
- `packages/adapter-utils/src/acpx-engine/execute.ts`: apply the same
`PAPERCLIP_*` non-override rule (via the shared helper) when merging
config env, capture the applied config env in `resolvedAdapterEnv`, and
fold a stable `adapterEnvHash` of it into the session fingerprint so an
env change forces a fresh launch. Per-wake `PAPERCLIP_*` runtime vars
are assigned earlier and never enter that map, so they stay out of the
hash; stable configured `PAPERCLIP_*` values (e.g. an explicit
`PAPERCLIP_API_KEY`) are included so rotating one busts the session.
- Added unit/integration tests for plain + secret forwarding,
`PAPERCLIP_*` non-override, the explicit-API-key path, fingerprint
refresh-on-env-change vs. stable-across-wakes, and rotation of a
configured `PAPERCLIP_API_KEY`.

## Verification

- `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts
packages/adapter-utils/src/server-utils.test.ts` → 2 files, 111 tests
passing (includes the new cases).
- Tests assert: forwarded plain/secret values appear in the spawned
wrapper `.env`; a `PAPERCLIP_*` config key does not override the runtime
value; changing an applied forwarded env value (including a rotated
`PAPERCLIP_API_KEY`) changes `configFingerprint`, while a new wake with
the same config env keeps it stable.

## Risks

Low risk. Behavior change is limited to (a) config env no longer
overriding `PAPERCLIP_*` runtime vars — a security-positive tightening —
and (b) a resumable session re-launching when its applied config env
changes, which is the intended fix. Per-wake `PAPERCLIP_*` churn is
deliberately excluded from the fingerprint so normal sessions still
resume across heartbeats. Existing sessions get a new fingerprint once
on first deploy (the added `adapterEnvHash` field), which is expected.
No secret values are logged (key-name redaction in the adapter plus
manifest-driven redaction of `meta.env`).

## Model Used

Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended
reasoning with tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 11:26:47 -05:00
Evyatar Bluzer b287281940
docs: add inline comments to docker quickstart compose file (#2431)
## Problem

The quickstart docker-compose file was recently moved to
\`docker/docker-compose.quickstart.yml\` during the Docker
reorganization but still have zero inline comments. When new user copy
this file for self-hosting, they see variables like:
- \`BETTER_AUTH_SECRET\` - what is this? How to generate?
- \`PAPERCLIP_DEPLOYMENT_MODE: "authenticated"\` - what other modes
available?
- \`PAPERCLIP_DEPLOYMENT_EXPOSURE: "private"\` - what does private vs
public mean?
- \`OPENAI_API_KEY\` and \`ANTHROPIC_API_KEY\` - both required? Or just
one?

They have to go read DOCKER.md or other docs to understand each
variable.

## What I changed

Added inline YAML comments directly in the file:
- Header block with step-by-step quickstart commands (cd docker, export,
docker compose up)
- Section headers grouping LLM keys, deployment settings, and secrets
- Comment explaining each non-obvious variable with valid values
- Note about \`BETTER_AUTH_SECRET\` with openssl generation command
- Comment on the volume explaining what data it persist

No functional change - only YAML comments added.
2026-07-15 09:39:20 -05:00
Dotta ea66ea81e6
fix(auth): honor responsible-user grants for company skills (#9571)
## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies
> - Company skills are governed resources, so board users and agents
acting for responsible users must be authorized consistently before
mutating skill configuration
> - The responsible-user authorization intersection handled several task
permissions but did not map company-skill mutation actions to the
corresponding `skills:create`, `skills:update`, and `skills:delete`
grants
> - That gap caused valid skill import and mutation requests to be
rejected even when the responsible user held the exact direct permission
required by the route
> - The branch also introduces the repo-sourced `prepare-paperclip-pr`
skill so the standard PR preparation process is versioned and reviewable
alongside the code
> - This pull request adds the missing authorization mappings, covers
board, agent, JWT-route, and denial behavior with regression tests, and
adds the renamed PR-preparation skill
> - The benefit is that governed company-skill workflows honor explicit
grants without weakening the responsible-user permission intersection

## Linked Issues or Issue Description

No public issue exists. Bug-report shape:

- **Affected area**: company skill authorization and skill import routes
- **Observed behavior**: agents acting under a responsible user could
receive `403` responses for company-skill mutations even when that user
had the matching direct `skills:create`, `skills:update`, or
`skills:delete` grant
- **Expected behavior**: the responsible-user authorization intersection
should accept exact company-skill grants while preserving denials for
missing or unrelated grants
- **Reproduction**: authenticate as an agent with a responsible user,
grant that user the relevant company-skill permission, then import or
mutate a company skill
- **Additional repository change**: adds the renamed
`prepare-paperclip-pr` skill as the versioned source of truth for PR
preparation

Supersedes #9324, which added the PR-preparation skill under the old
`prepare-pr` name.

## What Changed

- Added `.agents/skills/prepare-paperclip-pr/SKILL.md` with the standard
worktree, commit, rebase, guardrail, review-loop, and handoff procedure
- Mapped company `skill_config:create`, `skill_config:update`, and
`skill_config:delete` actions to direct `skills:create`,
`skills:update`, and `skills:delete` responsible-user grants
- Preserved restrictive behavior for unsupported resources, missing
grants, and unrelated permissions
- Added authorization-service regression coverage for board actors and
responsible-user agent intersections
- Added route-level JWT regression coverage for company skill imports,
including allowed and denied cases

## Verification

- `pnpm exec vitest run
server/src/__tests__/authorization-service.test.ts
server/src/__tests__/company-skills-import-authz-routes.test.ts` — 42
tests passed
- `pnpm -r typecheck` — passed
- `pnpm build` — passed
- `pnpm test:run` — server and UI groups passed; one unrelated CLI
doctor assertion failed because the execution environment injects static
`AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, which intentionally changes
the result from `pass` to `warn`
- `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY
NODE_ENV=development pnpm exec vitest run
cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when
non-secret provider config is present'` — passed, confirming the
full-suite failure is environment-specific
- GitHub CI — all required checks passed on head `0758393c`; one
unrelated `packages/db/src/client.test.ts` 5-second timing timeout
passed on the single allowed failed-job rerun after three consecutive
local passes (42/42 tests)

## Risks

- Low-to-moderate authorization risk: the change expands accepted
responsible-user grants only for company-scoped skill configuration
actions and is protected by explicit allow/deny regression cases
- No database migrations, workflow changes, lockfile changes, or UI
changes
- The added skill is documentation consumed by agent tooling and does
not alter runtime application behavior

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex coding agent; exact runtime model ID and context-window
size were not exposed to the session. Used reasoning, terminal
execution, Git/GitHub tooling, and test/build execution.
- Earlier commits were assisted by Claude Fable 5 (`claude-fable-5`) and
an OpenAI Codex coding agent, as recorded in the branch history/task
workflow.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and documented the one
environment-specific full-suite failure
- GitHub CI — all required checks passed on head `0758393c`; one
unrelated `packages/db/src/client.test.ts` 5-second timing timeout
passed on the single allowed failed-job rerun after three consecutive
local passes (42/42 tests)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 07:52:54 -05:00
Nicky Leach 7947308276
fix(codex): classify refresh auth failures (#9598)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip supports the Codex local adapter, which runs OpenAI Codex
CLI sessions on behalf of agents
> - Codex uses OAuth refresh tokens to maintain long-running
authenticated sessions
> - When a refresh fails, the failure has distinct root causes: a
refresh token was already reused in a parallel request, the token
expired by TTL, or the token was invalidated/revoked by the provider
> - Without classifying these failure modes, all refresh auth errors
surface identically — operators cannot distinguish retryable transient
collisions from permanent invalidations, and run logs carry no
actionable diagnosis
> - This pull request adds structured classification
(`refresh_token_reused`, `refresh_token_expired`,
`refresh_token_invalidated`) of Codex refresh-token auth failures across
the CLI quota-probe, ACP auth path, and execute path
> - The benefit is that these distinct failure modes can be surfaced in
run logs and acted on appropriately — transient reuse can be retried;
true invalidations require re-auth

## Linked Issues or Issue Description

<!-- Path B: no public GitHub issue — describing inline as a bug fix -->

**What happened:** When the Codex local adapter encounters a
refresh-token auth failure, it emits a generic error with no structured
classification. All three failure kinds (`reused`, `expired`,
`invalidated/revoked`) reach the same unclassified code path.

**Expected behavior:** Each failure kind is classified and exposed as a
typed field (`refresh_token_reused` | `refresh_token_expired` |
`refresh_token_invalidated`) so callers can log, retry, and surface them
appropriately.

**Steps to reproduce:**
1. Run a Codex agent session with a reused or expired OAuth refresh
token.
2. Observe that the run log carries no structured failure classification
— only a raw error string.

**Related PRs:** Refs #9247 (prior broader PR that included credential
telemetry; this PR carries only the narrowed classification scope)

## What Changed

- Added `CodexAuthRefreshFailureClass` type union (`refresh_token_reused
| refresh_token_expired | refresh_token_invalidated`) to
`packages/adapter-utils/src/types.ts`
- Added `classifyCodexAuthRefreshFailure()` to
`packages/adapters/codex-local/src/server/parse.ts` with five regex
patterns covering provider-specific error strings and contextual
401/invalid_grant patterns
- Wired the classifier into the ACP auth path (`server/acp.ts`), execute
path (`server/execute.ts`), and CLI quota-probe (`cli/quota-probe.ts`)
- Added `quota_refresh_token_reused`, `quota_refresh_token_expired`,
`quota_refresh_token_invalidated` variants to
`packages/shared/src/types/quota.ts`
- Added classification unit tests (`parse.test.ts`,
`quota-spawn-error.test.ts`, `acp.test.ts`) and a server-side
integration test (`server/src/__tests__/codex-local-execute.test.ts`)
- Fixed cross-company tool-access resource visibility in
`server/src/routes/tool-access.ts`
- Stabilized `heartbeat-retry-scheduling.test.ts` (CASCADE cleanup),
`heartbeat-run-log.test.ts`, and `quota-windows.test.ts`

## Verification

- `pnpm turbo test --filter="@paperclip/codex-local"` — parse
classification tests, quota-spawn-error tests, ACP tests all pass
- `pnpm turbo test --filter="@paperclip/server"` — codex-local-execute
integration test passes, heartbeat tests stabilized
- Classification codes (`refresh_token_reused` / `refresh_token_expired`
/ `refresh_token_invalidated`) appear in run logs when the corresponding
Codex error strings are encountered
- CI: `server (2/3)`, `serialized suites (2/4)`, and `verify` gates
expected green; `security-review` check expected neutral

## Risks

Low risk. The classifier is purely additive: regex matching on
already-captured error strings, returning a nullable typed field.
Callers that do not inspect the classification field are unaffected. No
execution paths, retry logic, or existing error surfaces changed.

## Model Used

- **Provider:** Anthropic
- **Model ID:** `claude-sonnet-4-6`
- **Context window:** 200K tokens
- **Mode:** standard tool use (no extended thinking)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-14 22:40:38 -07:00
Jannes Stubbemann 7f2ed0ad90
security(server): close cross-tenant existence oracle (404 instead of 403) (#3967)
## Thinking Path

> - Paperclip orchestrates AI agents for zero-human companies
> - In a multi-tenant deployment, route handlers that take a resource id
(`issue`, `goal`, `project`, `approval`, etc.) look the resource up by
id and then call `assertCompanyAccess` on its `companyId` — 404 if it
doesn't exist, 403 if it exists in another tenant
> - The split status codes are a classic *existence oracle*: any
authenticated user can enumerate ids across tenants by probing for the
403/404 boundary, mapping out which issues, labels, approvals, etc.
exist in other customers' tenants even when they cannot read the
contents
> - The right fix is a single uniform 404 for both "not found" and
"found but cross-tenant", which collapses the oracle but still preserves
write-path checks (active membership, viewer-readonly) for *authorized*
tenants
> - This pull request adds a non-throwing `hasCompanyAccess(req,
companyId)` helper plus a `getAccessibleResource` wrapper that ~130
handlers across 14 route files now use, folding the access check into
the existence check while still running `assertCompanyAccess` for
authorized tenants so viewer-readonly / inactive-membership rejections
fire unchanged on write paths
> - The benefit is closing a multi-tenant information leak without
breaking write-path security or single-tenant local-first behavior

## Linked Issues or Issue Description

Refs #709 — asks for company-scope regression coverage across
approval/activity/access routes, because a subtle route refactor could
leak cross-tenant data; this PR hardens exactly those surfaces (uniform
404 across 14 route files including `approvals`, `activity`, `secrets`)
and updates cross-tenant expectations in test files. It does not add the
full coverage matrix #709 asks for — hence Refs, not Closes.

No existing issue covers the oracle itself — described in-PR:

- Route handlers returned 404 for "not found" but 403 for "exists in
another tenant", a classic *existence oracle*: any authenticated user
could enumerate ids across tenants by probing the 403/404 boundary.
- That maps out which issues, labels, approvals, etc. exist in other
customers' tenants even when their contents are unreadable.
- Fix: a uniform 404 for both cases, while keeping write-path checks
(active membership, viewer-readonly) for authorized tenants.

## What Changed

- **`server/src/routes/authz.ts`** — new `hasCompanyAccess(req,
companyId): boolean` helper alongside the existing
`assertCompanyAccess`. Docstring spells out the two-step pattern (404
gate, then `assertCompanyAccess` for write-path checks). The helper
mirrors `assertCompanyAccess`'s company-scope semantics exactly — in
particular, signed-in instance admins do **not** get blanket access to
companies they are not a member of (the repo's `authz-company-access`
tests pin that behavior for `assertCompanyAccess`; an earlier draft of
the helper accidentally widened it for reads).
- **`getAccessibleResource(req, res, lookup, notFoundMessage)`** — the
safe thing is now the easy thing. One helper wraps the whole pattern
(uniform 404 for missing/cross-tenant, then `assertCompanyAccess` for
write-path membership checks) and ~130 handlers across 14 route files
use it:
  ```ts
const goal = await getAccessibleResource(req, res, svc.getById(id),
"Goal not found");
  if (!goal) return;
  ```
Files: `activity`, `agents`, `approvals`, `assets`, `costs`,
`environments`, `execution-workspaces`, `file-resources`, `goals`,
`issue-tree-control`, `issues`, `projects`, `routines`, `secrets`.
Handlers with bespoke not-found behavior (the legacy `200 []` contract,
audit-logged denials in `file-resources`, null-returning authz helpers)
compose `hasCompanyAccess` directly using the documented two-step
pattern:
  ```ts
// step 1: close the oracle (uniform 404 for both not-found and
cross-tenant)
  if (!existing || !hasCompanyAccess(req, existing.companyId)) {
    res.status(404).json({ error: "Goal not found" });
    return;
  }
// step 2: enforce write-path membership checks for authorised tenants
(no-op on GET)
  assertCompanyAccess(req, existing.companyId);
  ```
Routes where `companyId` comes from *request input*
(`req.params.companyId`, `req.body.companyId`, e.g. in `companies.ts`
and `plugins.ts`) deliberately retain plain `assertCompanyAccess` —
there's no existence oracle to close because the companyId is an input,
not a discovered value.
- **Full-sweep coverage** — a scripted audit of every
`assertCompanyAccess(req, <resource>.companyId)` call site in
`server/src/routes/` found ~55 lookup-then-assert pairs the first pass
missed; all are now gated. Notable ones: the
`/secret-provider-configs/:id` CRUD routes, the agents
instructions-bundle/config-revision/skills-sync routes (which check
access via the `assertCanUpdateAgent` / `assertCanReadAgent` /
`assertCanManageInstructionsPath` helpers), `POST
/heartbeat-runs/:runId/watchdog-decisions`, `GET
/issues/:id/cost-summary`, the environment + environment-lease GET
routes, all six issue-tree-control routes, ~24 issue sub-resource routes
(document annotations, interactions, approvals links, recovery actions,
plan decompositions, lock/unlock), and the three workspace file-resource
routes (these throw `notFound` instead of `forbidden` inside their
audit-logging wrappers, so denied attempts are still activity-logged
server-side while the client sees a uniform 404).
- **Helpers made self-defending** — `assertCanUpdateAgent` /
`assertCanReadAgent` / `assertCanManageInstructionsPath` (agents) and
`assertCanManage{Project,Execution}WorkspaceRuntimeServices` throw
`notFound` for cross-tenant resources before their `assertCompanyAccess`
step, so a future caller that forgets the route-level gate still can't
reopen the oracle.
- **Pattern enforcement** — new `authz-existence-oracle-guard.test.ts`
statically scans `server/src/routes/*.ts` and fails CI on any
`assertCompanyAccess(req, <resource>.companyId)` call that is not
preceded by a `hasCompanyAccess` gate, with an explicit allowlist (plus
staleness check) for the request-input cases. New routes that regress to
the 403/404 split fail the suite with a message pointing at the
documented pattern.
- **Tests** — cross-tenant expectations updated from 403→404 where
routes are now gated; new `hasCompanyAccess` unit tests in
`authz-company-access.test.ts` pin the
instance-admin/local-implicit/agent/none semantics in lockstep with
`assertCompanyAccess`; `write-path-membership.test.ts` (added in an
earlier round) confirms viewer/inactive users are still rejected on
writes.
- **One legacy-contract preserve** — `GET /heartbeat-runs/:runId/issues`
still returns `200 []` for both "doesn't exist" and "cross-tenant" so
the legacy contract is preserved while the oracle stays closed.

## Verification

- `pnpm run typecheck` — PASS.
- `pnpm -F @paperclipai/server exec vitest run` — full server suite
green locally apart from 4 pre-existing local-environment failures
(`paperclip-skill-utils` ×2 and `workspace-runtime` ×1 are
cwd/git-environment dependent — verified identical on a clean checkout
of the base; `heartbeat-process-recovery` is the known macOS flake).
- The new `authz-existence-oracle-guard` test sweeps
`server/src/routes/*.ts` and confirms no remaining
`assertCompanyAccess(resource.companyId)` site without a
`hasCompanyAccess` gate; the only allowlisted holdouts take `companyId`
from request input.

## Risks

- **API contract narrowing.** Any client that specifically checked for
`403` on cross-tenant access now sees `404`. This is a strict narrowing
(one status instead of two for the same negative outcome) and matches
what a client should expect for any id it can't access.
- **Write-path checks preserved.** `assertCompanyAccess` still runs
after the 404 gate on write routes, so viewer-readonly /
inactive-membership rejections fire unchanged for legitimate users.
- **Instance-admin scope unchanged.** `hasCompanyAccess` denies
signed-in instance admins without an explicit membership, exactly like
`assertCompanyAccess` (pinned by unit tests) — so the gate introduces no
new read access for admins.
- **Single-tenant local-first deploys** behave identically — the helper
short-circuits to `true` for `local_implicit` sessions.
- No new env vars, no deployment-mode switch.

## Model Used

Claude Opus 4.7 (1M context), extended thinking mode; completeness sweep
+ instance-admin parity fix by Claude Fable 5 (1M context).

## Checklist

- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] Thinking path traces from project context to this change
- [x] Model used specified
- [x] Checked ROADMAP.md — part of the multi-tenant hardening initiative
- [x] Tests run locally and pass
- [x] Added/updated cross-tenant 404 expectations across test files
- [x] No UI changes
- [x] Documented risks above
- [x] Will address all Greptile and reviewer comments before merge

Part of the multi-tenant hardening initiative — see also #5864
(per-company JWT keys) and #5865 (plugin tables `company_id`).

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-07-14 15:53:09 -07:00
Nicky Leach b79f744a8d
Fix Codex auth merge host-unusable fail closed (#9276)
## Thinking Path

> - Paperclip is the open source platform people use to manage AI agents
for work
> - The Codex adapter runs agent tasks in isolated sandbox environments
on the user's machine
> - When a Codex sandbox is reused across agent runs, its home directory
(including `~/.codex/auth.json`) is restored from a prior snapshot
> - Both the host machine and the sandbox independently maintain
`auth.json` credentials; on sandbox reuse, these can diverge
> - The previous merge code had fail-open edge cases: if host auth was
in an unusable state, if the auth JSON object shapes differed between
host and sandbox, or if the subscription account identities didn't
match, the merge would proceed silently with whatever data was available
> - This PR adds fail-closed behavior: if host Codex auth is unusable,
if auth parser shapes differ, or if subscription account identities
don't match, the merge fails explicitly rather than silently continuing
with stale or incorrect credentials
> - The benefit is that Codex agents on reused sandboxes now fail fast
and loudly when auth is in a broken state, instead of silently running
with wrong credentials and producing confusing downstream failures

## Linked Issues or Issue Description

No pre-existing public GitHub issue. This is a targeted security
hardening fix for the Codex reused-sandbox auth merge path.

**Problem:** When a Codex sandbox is reused, the merge logic that
reconciles host and sandbox `auth.json` credentials failed open in
several cases:
- Host `auth.json` present but in an unusable state (missing required
keys, empty token material, malformed JSON) → merge would proceed with
whatever the sandbox had
- Host and sandbox auth payloads had different shapes (e.g., one uses
`OPENAI_API_KEY`, the other uses a `tokens` object) →
parser-differential case not detected
- Subscription account identities (`tokens.account_id`) differed between
host and sandbox → stale sandbox identity would be used silently

**Fix:** All three cases now fail closed. The merge returns an explicit
error rather than proceeding with potentially stale or mismatched
credentials.

Related PRs:
- Refs #9262 — sandbox Codex auth shadow warning (adjacent auth area)
- Refs #9259 — auth precedence exports (adjacent auth area)

## What Changed

- `packages/adapters/codex-local/src/server/codex-home.ts` — New file
with `hasUsableAuthPayload()`, `codexHomeHasUsableAuth()`, and full
Codex home setup/teardown. Includes fail-closed auth merge guards:
rejects unusable host auth, detects parser shape differentials, and
checks subscription account identity match before merging
- `packages/adapter-utils/src/workspace-restore-merge.ts` — New file
with directory snapshot diffing and restore-merge logic; the merge
operation fails closed when auth validation fails
- `packages/adapters/codex-local/src/server/codex-home.test.ts` — Unit
tests covering auth usability checks, symlink management, and
fail-closed merge paths
- `packages/adapter-utils/src/workspace-restore-merge.test.ts` — Unit
tests for snapshot/restore-merge behavior including fail-closed cases
- `packages/adapter-utils/src/sandbox-managed-runtime.ts` — Updated to
invoke the fail-closed auth merge during sandbox restore

## Verification

Tests run and passing:

```sh
corepack pnpm exec vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts packages/adapters/codex-local/src/server/codex-home.test.ts
corepack pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts
corepack pnpm --filter @paperclipai/adapter-utils typecheck
corepack pnpm --filter @paperclipai/adapter-codex-local typecheck
git diff --check origin/master HEAD
```

All passed locally before push.

## Risks

- **Intentional behavioral change (breaking for previously-silent
failures):** Reused sandboxes that previously completed auth merge with
unusable host auth, parser-differential auth shapes, or mismatched
account identities will now fail with an explicit error. This is the
correct behavior — the prior silent-proceed path was the bug. Users
affected will see a clear error message rather than a confusing
downstream auth failure.
- **Auth.json symlink migration:** `ensureSymlink()` detects stale
copied `auth.json` files (written by older Paperclip versions) and
replaces them with symlinks on first run. This is safe: the target is
always under the Paperclip-managed company home, never the user's real
`~/.codex`. Directories at the symlink path are left untouched (EISDIR
is not silently swallowed).
- **Low risk for non-reuse paths:** The fail-closed logic only activates
during sandbox restore/reuse. Fresh sandbox allocations are unaffected.

## Model Used

- **Provider:** Anthropic
- **Model ID:** claude-sonnet-4-6 (Claude Sonnet 4.6)
- **Context window:** 200k tokens
- **Mode:** Agentic coding with tool use; extended thinking not used
- **Role:** Code author (Priya Raman, BackendEngineer) with Harold Kim
(Git Expert) handling push and PR operations

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Priya Raman <priya.raman@paperclip.local>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Harold Kim <harold@paperclip.ing>
2026-07-14 15:49:36 -07:00
Jannes Stubbemann 1cfed0c0ff
security(invites): widen invite-token entropy and rate-limit public invite endpoints (#8979)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Companies onboard human members through shareable invite links; the
`/api/invites/:token` endpoints are deliberately public so a recipient
can view the invite and accept it without being logged in
> - That publicness makes the invite token itself the only secret
guarding company membership — and it was guessable: the token suffix
carried only ~41 bits of entropy, and the endpoints had no rate limiting
> - An attacker could therefore enumerate the token space online and
accept an invite into someone else's company, gaining member access to
its onboarding data, skills, and workspace
> - This pull request widens invite tokens to 256 bits of entropy and
puts a per-IP rate limit in front of every public `/invites/:token`
sub-route
> - The benefit is that invite links stop being brute-forceable while
their shape, storage scheme, and UX stay exactly the same — existing
links keep working

## Linked Issues or Issue Description

No public issue exists; describing the problem in-PR (security/bug):

**What happens:** Company invite tokens are **public**: anyone with the
link can `GET /api/invites/:token`, fetch onboarding/logo/skills, and
`POST /api/invites/:token/accept`. Two weaknesses combined to make them
brute-forceable:

1. **Token entropy ~41 bits.** The token suffix was 8 chars over a
36-char alphabet (`8 * log2(36) ≈ 41.4` bits). That is
online-enumerable.
2. **No rate limit on `/invites/:token*`.** The public endpoints had no
throttling, so the ~41-bit space could be enumerated online.

**Impact:** an attacker who guesses a live token can accept the invite
and join the company as a member — unauthenticated, from any IP.

**Expected:** invite tokens should be computationally infeasible to
guess, and the public endpoints should throttle guessing attempts anyway
(defense in depth).

## What Changed

**Entropy**

- `createInviteToken` now uses `crypto.randomBytes(32)` (256 bits)
base64url-encoded, keeping the human-readable `pcp_invite_` prefix so
link shape and UX are unchanged. The duplicate generator in
`plugin-host-services.ts` is updated to match.
- Tokens are stored **hashed** (sha256) in `invites.tokenHash`; the raw
value is only returned once on creation. Storage scheme is unchanged.
- **Backward compatible**: only newly minted tokens are affected; lookup
is by hash of the presented value, so existing invite links keep
working.

**Rate limit**

- New generic in-memory per-IP sliding-window limiter
(`server/src/services/invite-rate-limit.ts`, 20 req/min/IP), applied as
a router-level middleware on `/invites/:token` so every current and
future sub-route is covered (summary, logo, onboarding, onboarding.txt,
skills/index, skills/:name, test-resolution, and POST accept).
- Returns `429` with `Retry-After` and `X-RateLimit-*` headers.
In-memory ⇒ per-process, which bounds enumeration per replica. Mirrors
the existing `company-search-rate-limit` pattern; no new dependency.
- Adds a `tooManyRequests(429)` error helper in `server/src/errors.ts`.

## Verification

- `invite-token-entropy.test.ts`: prefix preserved, suffix ≥ 128 bits /
22 chars, charset, 1000 unique tokens.
- `invite-rate-limit.test.ts`: allows up to limit then 429s with
retry-after; per-IP isolation; forgets hits after the window.
- `invite-rate-limit-route.test.ts`: `GET /invites/:token` and `POST
/invites/:token/accept` return 429 once the per-IP threshold is
exceeded.
- Manual: create an invite, open the link (works once per token as
before), then hammer `GET /api/invites/<token>` >20 times within a
minute from one IP → `429` with `Retry-After`.
- Server package typechecks clean for all touched files.

## Risks

- Low risk. Token change affects only newly minted tokens; existing
links resolve via the same sha256-hash lookup.
- The limiter is in-memory and per-process: in multi-replica deployments
each replica enforces its own 20 req/min/IP budget. That still bounds
enumeration (per-replica) and matches the existing
`company-search-rate-limit` approach; a shared store can be layered
later if needed.
- Legitimate users behind a single NAT/proxy IP share the 20 req/min
budget for invite endpoints; the invite flow makes only a handful of
requests, so headroom is ample.
- No DB migration, no API shape change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended
thinking enabled, agentic tool use (code search, editing, local
typecheck) via Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Supersedes #8147.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 15:47:58 -07:00
Jannes Stubbemann b4e7ba5143
feat(run-logs): durable run-log store via object-storage mirror (#8984)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every agent run streams its stdout/stderr/system output into the
run-log store (`server/src/services/run-log-store.ts`), and the run-log
API serves those logs back for review and debugging
> - The only store implementation is `local_file`: logs live on the
server pod's filesystem under `PAPERCLIP_HOME`
> - In hardened / ephemeral deployments, `PAPERCLIP_HOME` is an
`emptyDir` with no persistent volume, so every pod restart wipes the log
files while the DB row still references them — the run-log API then
returns "Run log not found" for every completed run after any redeploy
> - Run logs are the primary audit/debugging trail for agent work;
losing them on routine redeploys undermines trust in the platform
> - This pull request adds transparent durability: when
`RUN_LOG_S3_BUCKET` is set, the store mirrors each completed log to
object storage on `finalize` (same `logRef` key) and falls back to it on
`read` when the local file is gone; live append/tail stays on the fast
pod-local file
> - The benefit is that completed run logs survive pod restarts and
redeploys with zero changes for existing deployments (unset bucket =
today's behaviour) and zero downstream changes (store id stays
`local_file`)

## Linked Issues or Issue Description

No existing public issue — inline description following the bug report
template:

**What happened?** After any server pod restart/redeploy, the run-log
API returns "Run log not found" for all previously completed runs. The
DB still references the log file, but the file is gone because run logs
are written only to the pod-local filesystem.

**Expected behavior:** Completed run logs remain readable across pod
restarts and redeploys.

**Steps to reproduce:**
1. Deploy the server with `PAPERCLIP_HOME` on an `emptyDir` (no
persistent volume — common in hardened/ephemeral Kubernetes
deployments).
2. Complete an agent run and confirm its log is readable via the run-log
API.
3. Restart or redeploy the server pod.
4. Request the same run's log — the API throws "Run log not found".

**Paperclip version or commit:** reproducible on current `master`.
**Deployment mode:** Kubernetes (server pod without persistent volume).
**Agent adapter(s) involved:** Not adapter-specific (core bug).

Supersedes #8795.

## What Changed

- `server/src/services/run-log-store.ts`: the local-file store becomes a
durable store with an optional object-storage mirror
- `finalize` mirrors the completed NDJSON log to S3-compatible object
storage (keyed by the same `logRef`), best-effort so a failed upload can
never break run finalization; upload failures are logged via
`console.warn` so operators can detect a persistently broken mirror
before a pod roll makes logs unreadable
- `read` serves the pod-local file when present and falls back to a
ranged object-storage read (with correct `nextOffset`) when the local
file is gone
- Live `append`/tail stays on the pod-local file — fast path unchanged,
no per-chunk PUT
- Store id stays `local_file`, so nothing downstream changes (feedback
pipeline, read casts, fixtures untouched)
- New optional config, all read at store construction:
`RUN_LOG_S3_BUCKET`, `RUN_LOG_S3_ENDPOINT`, `RUN_LOG_S3_REGION` (default
`us-east-1`), `RUN_LOG_S3_PREFIX` (default `run-logs`),
`RUN_LOG_S3_FORCE_PATH_STYLE` (default `true`); credentials via the
standard AWS env chain; works with any S3-compatible endpoint
- Reuses the existing `createS3StorageProvider`; deliberately
independent from `PAPERCLIP_STORAGE_PROVIDER` so enabling durable logs
does not redirect workspace/file storage
- `server/src/services/run-log-store.test.ts` (new): 7 tests with an
in-memory `StorageProvider` mock

## Verification

- `npx vitest run src/services/run-log-store.test.ts` in `server/` — 7/7
pass locally:
  - store id stays `local_file`
  - live read served from the local file (no S3 round-trip)
  - `finalize` uploads the completed log to the mirror
- read falls back to S3 after a simulated pod roll (local file deleted)
  - ranged S3 read returns correct slice + `nextOffset`
  - not-found when neither local nor mirror has the log
  - local-only safe degrade when no bucket is configured
- `npx tsc --noEmit -p server` — clean for the touched files
- Manual: set `RUN_LOG_S3_*` against any S3-compatible endpoint (e.g.
MinIO), complete a run, delete the local `.ndjson` file, and re-request
the log via the run-log API — it is served from the mirror

## Risks

- Low risk: with `RUN_LOG_S3_BUCKET` unset (the default), behaviour is
byte-for-byte today's local-only store
- Mirror upload is best-effort by design — a misconfigured bucket loses
durability (not correctness) for affected runs; failures are now
surfaced via a `console.warn` per failed upload
- No DB migration, no API shape change, no change to the persisted
`store`/`logRef` handle format

## Model Used

- Claude (Anthropic), model ID `claude-fable-5` (Fable 5), via Claude
Code with extended thinking and tool use (code execution, file editing).
Original implementation TDD-authored with the same tooling.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 15:45:39 -07:00